Skip to content
pols.so docs
Esc
↑↓navigate↵open⌘Jpreview
On this page

Computer use

Let an AI agent see and operate a sandbox's desktop with screenshots, mouse and keyboard, or drive its Chrome over the DevTools Protocol.

Every running sandbox is a small computer an agent can use: a 1920x1080 desktop with Google Chrome. There are two ways to control it:

  • The desktop: take screenshots and send clicks, drags, typing, key presses and scrolling, like a person at the screen. Works for any application.
  • The browser over CDP: connect Playwright or Puppeteer to the sandbox’s Chrome. Faster and more reliable for anything that happens in a web page.

Both are available in the CLI, the API, the MCP server and the TypeScript client.

The desktop

Coordinates are pixels from the top left corner: x from 0 to 1919, y from 0 to 1079.

pols computer screenshot web -o before.png      # -o - writes the PNG to stdout
pols computer click web 960 540                 # --button right or middle
pols computer double-click web 960 540
pols computer drag web 100 100 400 300
pols computer type web 'hello world'            # up to 10,000 characters, or --stdin
pols computer key web ctrl+l                    # X key names: Return, Escape, Tab, BackSpace, F5; + for chords
pols computer key web ctrl+a BackSpace          # several keys in a row
pols computer scroll web 960 540 down --amount 5
API call Does
GET /v1/sandboxes/{sandbox}/computer/screenshot a PNG of the whole desktop, mouse pointer included
POST /v1/sandboxes/{sandbox}/computer/actions one action: click, double_click, drag, type, key or scroll

An action returns once it has been sent; it does not wait for the screen to change. Typing runs at about 80 characters per second. Take a screenshot after each action to check what happened, and before the next click to find its target.

To watch what an agent does, open the desktop in your browser at the same time.

The browser over CDP

pols browser cdp web
# wss://sbx-...-cdp.on.pols.so/...

The command (POST /v1/sandboxes/{sandbox}/browser/cdp) starts Chrome on the desktop if it is not running and returns a WebSocket URL for Chrome’s DevTools Protocol. Pass it to your automation library:

import { chromium } from "playwright";
const browser = await chromium.connectOverCDP(url);
const page = browser.contexts()[0]?.pages()[0] ?? (await browser.newPage());
await page.goto("https://example.com");
import puppeteer from "puppeteer";
const browser = await puppeteer.connect({ browserWSEndpoint: url });

The URL carries a token that is valid only for this sandbox, for 5 minutes, and only while the API key that created it is not revoked. An open connection is not cut when the token expires. Whoever has the URL controls the browser: keep it secret.

Clients that can send headers can also connect to GET /v1/sandboxes/{sandbox}/browser/cdp directly with the API key, for example Playwright’s connectOverCDP(url, { headers }).

Chrome runs on the visible desktop, so what the automation does also shows in screenshots and in the browser view, and you can mix CDP with desktop actions. Chrome’s debugging port listens only inside the VM; the connection is relayed by the control plane.

Tips for agents

  • Prefer CDP for web pages: selectors are more robust than pixel coordinates.
  • Use the desktop for everything else: native applications, file dialogs, browser extensions, or checking what a page actually looks like.
  • Record a session for review with pols-record start and pols-record stop inside the sandbox; the MP4 lands in ~/recordings/.