Skip to main content
Scaling an agent on KERNEL comes down to three questions, in this order: how many browsers you can run and create, how fast each one starts, and whether your workload needs a browser pool. Most workloads scale on on-demand browsers alone. This guide assumes you’re comfortable creating and controlling browsers.

1. Know your limits

Three limits shape a workload at scale, and each is covered in concurrency and limits:
  • Concurrency: how many browsers can exist at once across your organization. Browsers in standby still count, so delete browsers when a task finishes to free the slot.
  • Create rate: how fast you can create new browsers. Exceeding it returns 429 Too Many Requests, which the SDKs retry automatically.
  • Per-browser resources: how much memory each browser has, which caps how many tabs and how heavy a page one browser can handle.
Each plan has set limits, and they go up when you upgrade your plan. Enterprise limits are custom. Use project concurrency limits to split one organization’s limit across teams or environments.

2. Keep browsers fast

A browser is created in about 30ms at P50 and 105ms at P99 (performance). Most slow starts come from configuration rather than load: custom viewports, extensions, and kiosk mode restart Chromium on creation and add seconds. Check troubleshooting latency before reaching for anything else, and run your code next to the browser with Playwright execution or the code execution platform to cut the time each action takes.

3. Decide between on-demand browsers and a browser pool

We recommend defaulting to on-demand browsers, both when you’re getting started and as you scale. Browser pools fit a specific type of workload, described below. Stick with on-demand browsers.create() when:
  • you’re still building
  • your configuration changes per user (a pool is one fixed config)
  • you need a GPU browser (not available in pools)
Reach for a browser pool when:
  • you’ve built and scaled your workload, and every run uses the same workload attributes
  • you’re hitting the browsers.create() rate limit at volume
  • you need the lowest possible acquisition latency (for example, a heavily customized browser config)
  • traffic is steady or high-frequency enough to keep the browser pool utilized
If you’re on an Enterprise plan, speak with your account manager about applicable rate limits for browsers.create() and what’s best for your workloads.

What a browser pool gives you

A browser pool keeps a set of identically-configured browsers ready for immediate use. Compared to creating browsers on demand, it gives you:
  • Lowest-latency acquisition — the browser is already booted with your configuration applied (including settings like custom viewports, extensions, and kiosk-mode live view that otherwise restart Chromium on a fresh browser), so acquire hands you one that’s ready to drive.
  • Reserved, pre-configured capacity — a fixed set of browsers on your exact configuration, ready before traffic arrives.
  • Higher creation throughput — acquiring from a pool isn’t subject to the rate limit on browsers.create() that high-volume workloads hit.
The tradeoff: a browser pool counts against your concurrency limit whether or not its browsers are currently acquired — a pool sized to 40 holds 40 of your limit. Idle pooled browsers aren’t billed, but they hold the slot.

Sizing a pool

Watch available_count and target 10–20% available under normal load, resizing before traffic peaks rather than during them. See Sizing a browser pool for the full guidance.

Architecture patterns

On-demand creation

Creating a browser per task is the simplest approach. It’s the right fit while you’re building and when each task needs its own configuration. When to use:
  • Early development and testing
  • Configuration that changes per user or per task
  • GPU browsers

Single browser pool

For production workloads that run on the same configuration every time, a browser pool hands you ready-to-drive browsers, and acquiring from it isn’t subject to the browsers.create() rate limit. When to use:
  • Consistent, high-frequency workloads on a fixed configuration
  • Steady request patterns, or latency-sensitive acquisition
Key considerations:
  • Pool size should match your typical concurrency
  • Always release browsers in a finally block to prevent browser pool exhaustion
  • Set acquire_timeout_seconds based on your SLA requirements

Queue-based processing

When request volume exceeds your concurrency or traffic arrives in unpredictable bursts, put a task queue in front of your browsers. The example below acquires from a browser pool; the same pattern works on demand, with browsers.create() in place of acquire and deleteByID in place of release. When to use:
  • Request volume exceeds your available concurrency
  • Highly variable traffic patterns
  • Need to prioritize certain tasks
  • Want to decouple request ingestion from processing
Queue-specific considerations:
  • Set worker concurrency to match or slightly exceed browser pool size
  • Implement proper retry logic for transient failures
  • Monitor queue depth to scale browser pools dynamically
  • Use priority queues for different SLAs