Browser automation API for AI agents

The cloud browser API that finishes the task, not just opens the page

Send a task in plain English. A real Chrome on a residential IP does it. You get the result back as data, plus a live viewer URL to watch or take over.

$0.10 per browser-hour plus $0.02 per agent step, AI included. No subscription. First top-up can be $1.

How it works

  1. 1

    Send a goal

    POST a message/send JSON-RPC call to the A2A endpoint with your token. The task is plain English. Optional metadata: country, profile, mobile_ua, model.

  2. 2

    Watch it run

    The response carries metadata.session_id and metadata.viewer_url. Open the viewer to see the browser live. If the agent needs an OTP, the task moves to input-required and you can answer or take over.

  3. 3

    Get the result

    Poll with tasks/get and wait_seconds, or use message/stream for server-sent events. States: working, input-required, then completed, failed or canceled.

1. Send a task
curl -X POST https://agent.humanbrowser.cloud/a2a \
  -H 'Authorization: Bearer hb_live_xxxx' \
  -H 'Content-Type: application/json' \
  -d '{
    "jsonrpc": "2.0", "id": 1, "method": "message/send",
    "params": {
      "message": { "role": "user",
        "parts": [{"kind":"text","text":"Open news.ycombinator.com and return the top 5 story titles as JSON."}] },
      "metadata": { "country": "us" }
    }
  }'
2. Poll for the result
curl -X POST https://agent.humanbrowser.cloud/a2a \
  -H 'Authorization: Bearer hb_live_xxxx' -H 'Content-Type: application/json' \
  -d '{"jsonrpc":"2.0","id":2,"method":"tasks/get","params":{"id":"<task id from result.id>","wait_seconds":30}}'

Or plug it into your agent over MCP

Add the remote MCP server to Claude, Cursor or Codex. Tools: humanbrowser_run (starts a task, returns task_id and viewer_url without blocking), humanbrowser_status (long-polls up to 60 seconds) and humanbrowser_viewer_url. Running locally instead: npx -y @virixlabs/humanbrowser mcp with HB_TOKEN set.

MCP config
{ "mcpServers": { "humanbrowser": {
    "url": "https://agent.humanbrowser.cloud/mcp",
    "headers": { "Authorization": "Bearer hb_live_..." } } } }

Agent card: agent.humanbrowser.cloud/.well-known/agent-card.json · OpenAPI: /openapi.json · Full reference: /llms-full.txt

What is in one API call

A headless browser API hands you a browser. This one hands you a browser, the agent that drives it, and the parts that usually break a web task in production.

The agent is included

An LLM plans and drives the browser for you. Default model gpt-6-luna; pick gpt-6-sol, the gpt-5.x family, Claude Sonnet 4.6, Haiku 4.5, Opus 4.8, Kimi K2 or MiniMax per task.

Real headful Chrome

A full Chrome window running under Xvfb, not a stripped headless build. Pages render the way they do for a person.

Residential IP, 75 countries

Residential exits are on by default and sticky for the whole session. Pass metadata.country with an ISO code to choose the exit.

Captcha handling, 12+ types

reCAPTCHA v2/v3/Enterprise, hCaptcha, Cloudflare Turnstile, GeeTest v3/v4, FunCaptcha, Amazon WAF, Yandex SmartCaptcha, image captchas and more. $0.005 each.

Persistent profiles

Pass metadata.profile with a name and cookies and logins carry over between runs. Log in once, reuse the session tomorrow.

Live viewer and human takeover

Every task returns a viewer_url. Watch it run, or take over for an OTP or a judgement call. Takeover time is not billed.

File upload

Attach files as A2A FileParts, by URI or as base64 bytes (up to 25 MB, up to 8 files), and the agent uploads them into the page.

A2A, MCP and OpenAPI

A2A (v0.3) JSON-RPC with SSE streaming, a remote MCP server for Claude, Cursor and Codex, and an OpenAPI spec. One token for all of them.

What teams send it

Built for developers and companies shipping AI agents, data extraction, QA and monitoring. Each example below is the literal text you would put in the task.

Data behind a login

“Log in to our supplier portal with the saved profile, open Invoices, and return every invoice from September as JSON with number, date and total.”

QA regression flows

“Go to staging.example.com, sign up with a new test email, complete onboarding, and report any step that errors or takes longer than 10 seconds.”

Price, stock and page monitoring

“Open these five product pages from a UK IP and return price, currency and in-stock status for each.”

Forms and back-office work

“Open the vendor registration form, fill it from the attached PDF, upload the certificate, submit, and return the confirmation number.”

Research agents

“Search for the pricing pages of the ten tools in this list and return plan names and monthly prices, with the source URL for each.”

Pricing

Prepaid balance, no subscription, no monthly minimum. Your first top-up can be $1, later top-ups from $5; balance never expires.

ItemPriceNotes
Browser session$0.10/hrBilled per second. Under 30 s not billed. Human takeover time not billed.
Agent step (AI included)$0.02 / stepThe model driving the browser is included (default gpt-6-luna). If you pin a model whose own cost for a step is higher, that step is billed at its cost +30%.
Captcha$0.005 each$5 per 1,000. 12+ types.
Residential traffic$4 / GBOn by default, sticky per session, 75 countries.
Live viewer and takeoverIncludedNo extra charge.
$0.24
median cost per successful task, all-in (steps + browser time + traffic + captcha)
8 steps
median agent steps per task
$0.12–$0.39
middle half of tasks (p25–p75)

Measured on our production: 132 successful tasks run by 35 customer accounts in September 2026, our own and benchmark traffic excluded, re-priced at the current rates. Cost per task includes failed attempts in the same session. About 13% met a captcha.

Pay by card through Stripe (Apple Pay and Google Pay included) or in crypto (USDT, USDC, BTC, ETH, SOL). Refund terms are on the refund policy page.

When a raw headless browser API is the better fit

Per raw browser-hour we now charge about what Browserbase, Steel and Hyperbrowser charge: $0.10 an hour. Browser Use Cloud is cheaper at $0.02 an hour. The difference is what sits on top: on Human Browser the agent is $0.02 per step with the AI included, and the residential IP and captcha handling are part of the service. If you already run your own Playwright or Puppeteer script and just need a hosted Chrome to connect it to, a raw browser service with no agent costs less. Prices below are from their own pricing pages, as of Sep 2026.

ServiceBrowser timeProxyPlan
Browserbase$0.12 / hr over 100 hrs (Developer), $0.10 / hr over 500 hrs (Startup)$12 / GB, $10 / GB$20 or $99 / month; free tier 1 hr
Steel$0.10 / hr (Launch), $0.08 / hr (Scale)$10 / GB, $6 / GB$0 + usage, or $250 / month
Browser Use Cloud$0.02 / hr$5 / GBPay as you go; agents are model cost + 20%
Human Browser$0.10 / hr; agent $0.02 / step, AI included$4 / GBPay as you go, no monthly fee

They are also stronger on compliance (SOC 2, and HIPAA and SSO on Browserbase Scale), on concurrency (25 to 100 and up on paid plans, against our 10 per token), and on raw CDP access. Where we come out ahead is residential traffic ($4 per GB against $10–12 on Browserbase and $6–10 on Steel), no monthly fee, and the work you do not have to write: the agent, captcha handling, human takeover and persistent logins on one bill. Steel Launch is cheaper than us on captchas ($3 per 1,000 against our $5).

So compare on cost per finished task and on engineering time, not per hour. Detailed comparisons: Human Browser vs Browserbase, vs Steel, Browser Use alternative, Browser Use Cloud pricing, Skyvern alternative.

Limits

  • Concurrency: up to 10 sessions at once per paid token. Need more? Write to [email protected].
  • Goal-level API: you send tasks, not browser commands. There is no raw CDP or WebSocket endpoint to connect your own Playwright or Puppeteer script to.
  • Files: up to 8 files per task, base64 attachments up to 25 MB.
  • Captchas: if automatic handling fails, the task asks for a person in the live viewer rather than guessing.

FAQ

What is a cloud browser API?

A cloud browser API gives your code or your AI agent a browser that runs on someone else's servers, so you do not have to host Chrome, proxies and captcha handling yourself. Human Browser is a goal-level cloud browser API: you send a task in plain English, our agent drives a real Chrome on a residential IP, and you get the result plus a live viewer URL.

How much does the Human Browser API cost?

Browser time is $0.10 per hour, billed per second. The agent is $0.02 per step with the AI that drives the browser included. Captchas are $0.005 each and residential traffic is $4 per GB. Measured on our production over 132 successful customer tasks in September 2026, the median task took about 8 agent steps and cost $0.24 all-in. There is no subscription and no monthly minimum, and your first top-up can be $1; later top-ups start at $5.

Can I connect my own Playwright or Puppeteer script to it?

No. Human Browser is a goal-level API, not a raw CDP or WebSocket endpoint. You describe the task and our agent drives the browser. If you already have a Playwright or Puppeteer script and only need a hosted Chrome to run it on, a raw headless browser service such as Browserbase, Steel or Browser Use is the better and cheaper fit.

How do I get results back?

Call message/send on the A2A endpoint, or message/stream for server-sent events. The response includes a task id, a session_id and a viewer_url. Poll with tasks/get and wait_seconds until the task is completed, failed or canceled. If the agent needs something from a person, such as an OTP, the task moves to input-required.

Does it handle captchas and logins?

Yes. The agent handles more than 12 captcha types, including reCAPTCHA, hCaptcha, Cloudflare Turnstile, GeeTest, FunCaptcha and Amazon WAF, at $0.005 each. If automatic handling fails, the session can hand over to a person in the live viewer. Logins persist across runs when you pass a named profile.

How many sessions can I run at once?

Up to 10 concurrent sessions per paid token. If you need more, write to [email protected].

Which AI models can drive the browser?

The default is gpt-6-luna. You can choose gpt-6-sol, the gpt-5.x family, Claude Sonnet 4.6, Claude Haiku 4.5, Claude Opus 4.8, Kimi K2 or MiniMax per task with metadata.model. Model usage is included in the $0.02 per agent step; a step on a model you pin whose own cost is higher is billed at that cost +30%.

Try it on a real task

A $1 first top-up covers about four median tasks. Later top-ups from $5.

Try it for $1
Top up — $20

Loading secure checkout…

More coins →