- How you drive the browser — the control surface your code or model uses to act on the page.
- Where the loop runs — the machine your decision-making code runs on, relative to the browser.
1. How you drive the browser
KERNEL browsers accept several control surfaces. Pick by what’s driving the page, not by what you already know.
For agents, start with playwright execution with a computer use fallback: script the deterministic steps, and hand the page to a computer use model when a step doesn’t respond to a selector.
Why the choice matters on hardened sites
CDP is what Playwright and Puppeteer speak, and anti-bot vendors scan for its signatures. Computer controls carry no CDP connection, so there’s no protocol fingerprint to leak. That makes them the stronger option on sites with aggressive detection, and it’s why managed auth drives logins with coordinate-based input rather than CDP. How much this matters is site-specific, so test before you commit — see bot anti-detection.Control surface examples
- Computer Use
- Playwright Execution
- CDP
- WebDriver BiDi
Kernel’s Computer Controls API exposes OS-level mouse, keyboard, and screen primitives — the surface a computer-use model already knows how to drive (screenshot, click, type, key, scroll, drag). No CDP or WebDriver connection required, so there’s no protocol fingerprint to leak. Ideal for Claude, OpenAI, or Gemini computer-use loops.
2. Where the loop runs
Your loop is whatever decides the next action: a script, an agent, or a model. It can run in three places.- Your own infrastructure
- Playwright execution API
- Code execution platform
Connect to
cdp_ws_url or webdriver_ws_url from wherever your code already runs. Any CDP client works, and there’s no lock-in.Costs: a network round trip per action, disconnects to handle, screenshot and DOM bandwidth, and the CDP fingerprint above. It’s fine for low-frequency or deterministic work, and it hurts most in a vision loop.Where computer use fits
A computer use agent answers the first question, not the second — it still has to run its loop somewhere. Because every turn ships a screenshot instead of a small script, running that loop off-platform costs far more than it does for a Playwright-driven agent: you pay image bandwidth and a round trip on every step. That makes computer use the strongest case for running your loop next to the browser. Model inference stays with the model vendor either way.Putting it together
Computer use + playwright execution
Computer controls drive the browser the way a person would — they don’t speak the programmatic API surface. Anything you’d reach for the DOM or Playwright client for (reading text and attributes,page.goto, file uploads, cookie or storage access, switching tabs) belongs on the playwright execution side. When computer use is driving, expose playwright execution to the agent as a tool it can call for structured data or a programmatic action. For the full pattern in the other direction — playwright execution first, computer use when a step doesn’t respond to a selector — see playwright with computer use fallback.
Lower-level access
For work that isn’t driving the page, you can also reach the browser’s VM directly:- Browser curl: send HTTP requests through the browser’s network stack, with its cookies and proxy.
- SSH: open a shell in the browser’s VM, or forward a local port into it.
Give a model these surfaces as tools
If you’re building your own agent, Browser Loop packages these surfaces as tools for each model provider and runs every action against a KERNEL browser, so you don’t write the translation layer yourself.Going deeper
- Computer Controls reference — every mouse, keyboard, and screen primitive.
- Playwright Execution reference — the full execution surface, return values, and timeouts.
- Computer use integrations — drop-in examples for Anthropic, Gemini, OpenAI, and more.
- Firewall allowlist lists the domains and ports to allow for API, CDP, and WebDriver BiDi connections.