> ## Documentation Index
> Fetch the complete documentation index at: https://tbd-6fc993ce-hypeship-ia-how-it-works.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Control Overview

> Choose a control surface and where your agent loop runs

You make two choices before you write any automation. They're independent, but the first constrains the second:

1. **How you drive the browser** — the control surface your code or model uses to act on the page.
2. **Where the loop runs** — the machine your decision-making code runs on, relative to the browser.

## 1. How you drive the browser

KERNEL browsers accept several control surfaces. Pick by what's driving the page, not by what you already know.

| Surface | Use it when | Trade-off |
| - | - | - |
| [Playwright execution](/browsers/playwright-execution) | **Default.** You know what to do on the page — navigate, fill, extract, upload. | Needs a selector or DOM path that exists. |
| [Computer controls](/browsers/computer-controls) | **Recommended fallback.** A model is looking at pixels, or the page can't be driven programmatically. | Slower per step, and the model has to see the state to act. |
| [WebMCP](/browsers/webmcp) | The site exposes structured tools for the action you need. | Only works on sites that register tools. |
| [Browser REPL](/browsers/repl) | An agent writes its own helpers and reuses them across turns. | JavaScript only, and state lives until the REPL resets. |
| CDP | You have an existing Playwright, Puppeteer, or CDP codebase to point at Kernel. | Adds a protocol fingerprint and a network hop. |
| WebDriver BiDi | You need the W3C standard protocol. | Smaller client ecosystem. |

For agents, start with [playwright execution with a computer use fallback](/browsers/playwright-computer-use-fallback): script the deterministic steps, and hand the page to a computer use model when a step doesn't respond to a selector.

### Why the choice matters on hardened sites

CDP is what Playwright and Puppeteer speak, and anti-bot vendors scan for its signatures. Computer controls carry no CDP connection, so there's no protocol fingerprint to leak. That makes them the stronger option on sites with aggressive detection, and it's why [managed auth](/auth/managed-auth) drives logins with coordinate-based input rather than CDP. How much this matters is site-specific, so test before you commit — see [bot anti-detection](/browsers/bot-detection/overview).

### Control surface examples

<Tabs>
  <Tab title="Computer Use">
    Kernel's [Computer Controls](/browsers/computer-controls) API exposes OS-level mouse, keyboard, and screen primitives — the surface a computer-use model already knows how to drive (screenshot, click, type, key, scroll, drag). No CDP or WebDriver connection required, so there's no protocol fingerprint to leak. Ideal for [Claude](/integrations/computer-use/anthropic), [OpenAI](/integrations/computer-use/openai), or [Gemini](/integrations/computer-use/gemini) computer-use loops.

    <CodeGroup>
      ```typescript Typescript/Javascript theme={null}
      import Kernel from '@onkernel/sdk';

      const kernel = new Kernel();
      const kernelBrowser = await kernel.browsers.create();

      const screenshot = await kernel.browsers.computer.captureScreenshot(kernelBrowser.session_id);

      await kernel.browsers.computer.clickMouse(kernelBrowser.session_id, {
        x: 420,
        y: 280,
      });

      await kernel.browsers.computer.typeText(kernelBrowser.session_id, {
        text: 'kernel cloud browsers',
      });
      ```

      ```python Python theme={null}
      from kernel import Kernel

      kernel = Kernel()
      kernel_browser = kernel.browsers.create()

      screenshot = kernel.browsers.computer.capture_screenshot(id=kernel_browser.session_id)

      kernel.browsers.computer.click_mouse(
          id=kernel_browser.session_id,
          x=420,
          y=280,
      )

      kernel.browsers.computer.type_text(
          id=kernel_browser.session_id,
          text="kernel cloud browsers",
      )
      ```

      ```go Go theme={null}
      package main

      import (
      "context"

      "github.com/kernel/kernel-go-sdk"
      )

      func main() {
      ctx := context.Background()
      client := kernel.NewClient()

      kernelBrowser, err := client.Browsers.New(ctx, kernel.BrowserNewParams{})
      if err != nil {
      	panic(err)
      }

      screenshot, err := client.Browsers.Computer.CaptureScreenshot(
      	ctx,
      	kernelBrowser.SessionID,
      	kernel.BrowserComputerCaptureScreenshotParams{},
      )
      if err != nil {
      	panic(err)
      }
      defer screenshot.Body.Close()

      if err := client.Browsers.Computer.ClickMouse(
      	ctx,
      	kernelBrowser.SessionID,
      	kernel.BrowserComputerClickMouseParams{
      		X: 420,
      		Y: 280,
      	},
      ); err != nil {
      	panic(err)
      }

      if err := client.Browsers.Computer.TypeText(
      	ctx,
      	kernelBrowser.SessionID,
      	kernel.BrowserComputerTypeTextParams{
      		Text: "kernel cloud browsers",
      	},
      ); err != nil {
      	panic(err)
      }
      }
      ```
    </CodeGroup>
  </Tab>

  <Tab title="Playwright Execution">
    Run any Playwright code from anywhere — no local Playwright install, no Chromium download, no CDP connection to manage. Your code executes inside the browser's VM with the full Playwright API in scope and returns structured data back to your agent. Ships with [Patchright](/browsers/bot-detection/stealth) by default.

    <CodeGroup>
      ```typescript Typescript/Javascript theme={null}
      const response = await kernel.browsers.playwright.execute(
        kernelBrowser.session_id,
        {
          code: `
            await page.goto('https://example.com');
            return await page.title();
          `,
        },
      );

      console.log(response.result);
      ```

      ```python Python theme={null}
      response = kernel.browsers.playwright.execute(
          id=kernel_browser.session_id,
          code="""
            await page.goto('https://example.com')
            return await page.title()
          """,
      )

      print(response.result)
      ```

      ```go Go theme={null}
      response, err := client.Browsers.Playwright.Execute(
      ctx,
      kernelBrowser.SessionID,
      kernel.BrowserPlaywrightExecuteParams{
      	Code: `
            await page.goto('https://example.com');
            return await page.title();
          `,
      },
      )
      if err != nil {
      panic(err)
      }

      fmt.Println(response.Result)
      ```
    </CodeGroup>
  </Tab>

  <Tab title="CDP">
    Chrome DevTools Protocol — the wire format Playwright, Puppeteer, and most browser frameworks speak. Use `cdp_ws_url` from the created browser session for deterministic, scripted automation driven from your own infra.

    <CodeGroup>
      ```typescript Typescript/Javascript theme={null}
      import { chromium } from 'playwright';

      const browser = await chromium.connectOverCDP(kernelBrowser.cdp_ws_url);
      const context = browser.contexts()[0];
      const page = context.pages()[0];

      await page.goto('https://example.com');
      const title = await page.title();
      console.log(title);
      ```

      ```python Python theme={null}
      from playwright.async_api import async_playwright

      async with async_playwright() as playwright:
          browser = await playwright.chromium.connect_over_cdp(kernel_browser.cdp_ws_url)
          context = browser.contexts[0]
          page = context.pages[0]

          await page.goto('https://example.com')
          title = await page.title()
          print(title)
      ```
    </CodeGroup>
  </Tab>

  <Tab title="WebDriver BiDi">
    W3C-standard browser control. Use `webdriver_ws_url` with [Vibium](/integrations/vibium) or any other BiDi client.

    <CodeGroup>
      ```typescript Typescript/Javascript theme={null}
      import { browser } from 'vibium';

      const bro = await browser.start(kernelBrowser.webdriver_ws_url);
      const page = await bro.page();

      await page.goto('https://example.com');
      const title = await page.title();
      console.log(title);
      ```

      ```python Python theme={null}
      from vibium.sync_api import browser

      bro = browser.start(kernel_browser.webdriver_ws_url)
      page = bro.page()

      page.goto('https://example.com')
      title = page.title()
      print(title)
      ```
    </CodeGroup>
  </Tab>
</Tabs>

## 2. Where the loop runs

Your loop is whatever decides the next action: a script, an agent, or a model. It can run in three places.

<Tabs>
  <Tab title="Your own infrastructure">
    Connect to `cdp_ws_url` or `webdriver_ws_url` from wherever your code already runs. Any CDP client works, and there's no lock-in.

    **Costs:** a network round trip per action, disconnects to handle, screenshot and DOM bandwidth, and the CDP fingerprint above. It's fine for low-frequency or deterministic work, and it hurts most in a vision loop.
  </Tab>

  <Tab title="Playwright execution API">
    Send code, not commands. Each call runs in the browser's VM against the live session, so state carries across calls and an agent can drive the page turn by turn — one tool call per step, structured data back.

    **Costs:** the code you send has to be self-contained per call. There's nothing to install and no connection to manage.
  </Tab>

  <Tab title="Code execution platform">
    Deploy the whole agent next to the browser with the [code execution platform](/apps/develop), invoked on demand or on a schedule, with no infrastructure of your own.

    **Costs:** your agent has to be deployable as a Kernel app. It's worth it once the automation is long-running, stateful, or triggered by events rather than by a person.
  </Tab>
</Tabs>

### Where computer use fits

A computer use agent answers the first question, not the second — it still has to run its loop somewhere. Because every turn ships a screenshot instead of a small script, running that loop off-platform costs far more than it does for a Playwright-driven agent: you pay image bandwidth and a round trip on every step. That makes computer use the strongest case for running your loop next to the browser. Model inference stays with the model vendor either way.

### Putting it together

| Your automation | Control surface | Where the loop runs |
| - | - | - |
| Scheduled scrape of a known page | Playwright execution | Anywhere — one call, one result |
| Agent doing multi-step work on a normal site | Playwright execution, computer use fallback | Playwright execution API, or the code execution platform once it's long-running |
| Agent on a site with aggressive detection | Computer controls | Code execution platform |
| Existing Playwright suite you're migrating | CDP | Your own CI, then move hot paths to playwright execution |

## Computer use + playwright execution

Computer controls drive the browser the way a person would — they don't speak the programmatic API surface. Anything you'd reach for the DOM or Playwright client for (reading text and attributes, `page.goto`, file uploads, cookie or storage access, switching tabs) belongs on the [playwright execution](/browsers/playwright-execution) side. When computer use is driving, expose playwright execution to the agent as a tool it can call for structured data or a programmatic action. For the full pattern in the other direction — playwright execution first, computer use when a step doesn't respond to a selector — see [playwright with computer use fallback](/browsers/playwright-computer-use-fallback).

## Lower-level access

For work that isn't driving the page, you can also reach the browser's VM directly:

* **[Browser curl](/browsers/curl):** send HTTP requests through the browser's network stack, with its cookies and proxy.
* **[SSH](/browsers/ssh):** open a shell in the browser's VM, or forward a local port into it.

## Give a model these surfaces as tools

If you're building your own agent, [Browser Loop](/browsers/browser-loop) packages these surfaces as tools for each model provider and runs every action against a KERNEL browser, so you don't write the translation layer yourself.

## Going deeper

* [Computer Controls reference](/browsers/computer-controls) — every mouse, keyboard, and screen primitive.
* [Playwright Execution reference](/browsers/playwright-execution) — the full execution surface, return values, and timeouts.
* [Computer use integrations](/integrations/computer-use/anthropic) — drop-in examples for Anthropic, Gemini, OpenAI, and more.
* [Firewall allowlist](/info/network-access) lists the domains and ports to allow for API, CDP, and WebDriver BiDi connections.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.