Skip to content
Browse docs

The built-in agent harness (betterwright exec)

This page covers the standalone shape: BetterWright supplies a browser-tuned agent loop, you plug a model into it, and you hand it a natural-language task. For how it compares to the integrated shape, see Pick your shape first.

The two nest: a coding agent can shell out betterwright exec "<task>" as a browser sub-agent — one command in, one JSON answer out, with the entire browsing transcript (snapshots, retries, verification) kept out of the caller's context.

This shape exists because the browser runtime is rarely the slow part. The end-to-end gap is usually the agent scaffold — a browser-specialized loop takes fewer, tighter steps than a general coding agent. betterwright exec gives BetterWright that scaffold.

CLI

betterwright exec "find the top Hacker News story and give me its title and points" --model gpt-5.6-sol

Shells expand dollar signs inside double quotes, so use single quotes for tasks that contain prices or other literal $ text:

betterwright exec 'find options under $4000' --model gpt-5.6-sol
printf '%s\n' 'find options under $4000' | betterwright exec --stdin --model gpt-5.6-sol

The --stdin form is also useful for generated or multiline tasks because the task does not go through another round of shell parsing.

Progress notes stream to stderr as the loop runs — the last one summarizes the run's cost (done in 6 steps, 7 tool calls, 11.4s, 6,880 in / 1,330 out · 40,000 cache read · context 20,000) — and the final result is one JSON object on stdout:

{
  "ok": true,
  "answer": "…",
  "steps": 6,
  "reason": "done",
  "toolCalls": 7,
  "usage": {
    "inputTokens": 6880,
    "outputTokens": 1330,
    "cacheReadTokens": 40000,
    "cacheWriteTokens": 0,
    "context": 20000
  },
  "durationMs": 11400,
  "proof": "/…/proof-….png"
}

toolCalls counts the browser/login/ask/done calls the model issued (it can exceed steps when a turn batches several). usage sums the token counts the model adapter reported across turns (a field is 0 when the provider returned no usage block); inputTokens is fresh input only: each turn's provider input total minus the portion served from cache. cacheReadTokens and cacheWriteTokens come straight from the provider's usage block — the Responses API's input_tokens_details.cached_tokens / cache_write_tokens, the Chat Completions prompt_tokens_details equivalents, or Anthropic's cache_read_input_tokens / cache_creation_input_tokens. The CLI always shows cache reads; it shows cache writes only when the run has a positive, provider-reported count. It never derives writes from fresh input. context is the full prompt size at the end of the task — the last turn's provider input total, i.e. how much context the model was holding when it finished. durationMs is the task wall-clock (it excludes tearing down a browser the loop created for itself). The loop has no fixed step cap, but it does have a 30-minute wall-clock budget and a 1,000,000-character transcript bound so a stalled or repetitive provider cannot run forever or grow context without limit. Expiry aborts model requests, and BetterWright's worker timeout terminates in-flight browser work.

A third bound catches the loop that is running but not progressing: when a browser step fails the same way three times in a row, the observation carries a warning telling the model to change approach; at five, the run ends with reason: "no_progress" rather than spending the rest of the budget on a step that cannot succeed. Any successful browser call — or new human steering through the live view — clears the streak.

reason is one of done, answered (the model gave its answer in prose instead of calling done), stopped, interrupted (see sessions.md), timeout, context_limit, no_progress, max_tokens (the provider truncated the final response at the output-token limit), refusal (the model declined the task), or model_error. Only done and answered yield ok: true; every other reason still returns the partial answer and the full transcript.

Flags:

FlagDefaultMeaning
--model <id>whatever backend you have configured (or BETTERWRIGHT_MODEL)Real model id. Bare ids auto-select a source when unique; use source/id only to pin a collision. See Choosing a model
--base-url <url>nonePin --model to a custom OpenAI-compatible /v1 base URL (--endpoint is an alias)
--api-key-env <name>source defaultRead the API key from a named environment variable; raw keys are never accepted as CLI values
--protocol <name>chatchat for Chat Completions (widest compatibility), or responses when the server implements it
--allow-insecure-model-endpointoffAllow a key over non-loopback plain HTTP; HTTPS and loopback HTTP work without it
--effort <level> (alias --reasoning)lowReasoning effort: low/medium/high/xhigh/max where the model supports it
--session <name>defaultBrowser session name
--headedoffShow the managed browser

Network flags (--block-private-network, --allow-host, …) work the same as on run/repl.

Interactive console (betterwright)

betterwright exec runs one task and exits. Run betterwright with no subcommand for the interactive counterpart — a console where you type tasks and watch the agent work:

$ betterwright --model gpt-5.6-sol
BetterWright — interactive agent console
model gpt-5.6-sol · reasoning model default · session default · headless
Type a task and press Enter. /help for commands, /exit or Ctrl-D to quit.

▸ what is the page title of example.com
  · [1] browser: opening the page and reading its title

The page title is "Example Domain."
proof: /…/proof-….png
done · 2 steps · 2 tool calls · 2.1s · 1,889 in / 120 out · 3,072 cache read · context 4,961

Each step the agent takes streams as it happens, then the answer, the proof screenshot path, and the same cost summary exec prints. The session carries across tasks: both the browser (you stay signed in, tabs stay open) and the conversation — a follow-up task remembers what earlier ones did and can refer back to them without repeating the work (it's fed the running transcript). /new clears both the memory and the browser to start fresh. The same --model, endpoint, --effort/--reasoning, --session, --headed, and network flags apply.

With --live-view, the console starts and prints one viewer before the first prompt. That viewer remains open across follow-up tasks. Commands that replace the browser (/new, /headed, and /headless) also replace the viewer and print its new URL without restarting the console.

While a task is running, type a plain-text message and press Enter to steer it. The message is queued safely and applied at the next model turn boundary, just like chat sent from the live viewer. Slash commands typed during a task wait until that task finishes, so /new cannot tear down a browser mid-step. The active prompt changes to steer ▸; progress output redraws that prompt without discarding partially typed guidance and wraps with aligned continuation lines.

Because the transcript accumulates, a long session grows the context each task sends (largely served from cache — watch cache read in the summary); /new when you switch to unrelated work.

Meta-commands (a line starting with /):

CommandEffect
/helplist the commands
/endpoint <url>switch to a custom OpenAI-compatible base URL
/models [source]list available ids, optionally limited to openrouter, ollama, or vllm
/model <id>switch model id; use source/id only to resolve a collision
/reasoning <level>change reasoning effort (/effort also works)
/headedshow the browser window (/headless to hide it again)
/newclear conversation memory and close open tabs (fresh session)
/clearclear the screen
/exitquit (or Ctrl-D)

The ask tool

Because a user is present, the interactive console gives the agent an ask tool: when it genuinely needs input — a code it cannot obtain, a consequential choice with no reasonable default, or a task ambiguous enough that guessing risks the wrong thing — it asks you a question (offering short concrete options when the answer is a choice) and waits for your typed reply before continuing. It still acts on its own for ordinary reversible steps; it does not ask permission to proceed. betterwright exec has no user watching, so it runs without the ask tool and never stalls on a question.

Choosing a model

--model always takes a real model id (or a source-qualified id). There are no adapter nicknames: claude, codex, and grok by themselves are rejected. Pick the id you want to run; BetterWright figures out where it comes from.

When you name none

Omitting --model uses whichever backend you have actually configured, in this order: ANTHROPIC_API_KEY (with the @anthropic-ai/sdk peer installed) → claude-opus-4-8; a betterwright auth --login codex session → gpt-5.6-sol; a betterwright auth --login grok session → grok-4.3. BETTERWRIGHT_MODEL overrides all of it, and --model overrides that.

With nothing configured, exec and the console say so before doing any work and print the ways to fix it, rather than failing inside a model adapter. betterwright doctor reports the same thing under Built-in agent.

How selection works

  1. Source-qualified (ollama/qwen3:8b, openrouter/anthropic/claude-sonnet-5, codex/gpt-5.6-sol) — used immediately; no discovery.
  2. Custom base URL (--base-url / --endpoint) — pins the source to that URL; the bare id is the model name on that server.
  3. Bare id — BetterWright probes reachable catalogs and native family routes. If exactly one source exposes the id, that source is selected. If several do, the error lists the shortest unambiguous choices (for example codex/gpt-5.6-sol and ollama/gpt-5.6-sol). If none do, it tells you to run betterwright models or pin a source.

What bare-id discovery probes:

SourceWhen it is probed
OllamaAlways (default http://127.0.0.1:11434/v1; short timeout if down)
vLLMAlways (default http://127.0.0.1:8000/v1)
OpenRouterOnly when OPENROUTER_API_KEY is set
Native Claude / Codex / GrokWhen the id's family prefix matches (claude*, gpt* / o*, grok*, …)

Listing is separate from selection:

betterwright models                 # native defaults + reachable endpoints
betterwright models ollama          # only Ollama
betterwright models openrouter      # only OpenRouter
betterwright models --json          # machine-readable

Suggested lines use a bare id when it is unique, and source/id only when the same name appears in more than one place.

Quick starts by source

# Codex (ChatGPT subscription) — sign in once, then use a real GPT/Codex id
betterwright auth --login codex
betterwright exec "inspect example.com" --model gpt-5.6-sol

# Claude (Anthropic API) — optional peer dep + API key
npm install @anthropic-ai/sdk
ANTHROPIC_API_KEY=… betterwright exec "inspect example.com" \
  --model claude-opus-4-8

# Grok (xAI OAuth or API key)
betterwright auth --login grok
betterwright exec "inspect example.com" --model grok-4.3
# or: XAI_API_KEY=… betterwright exec "…" --model grok-4.3

# Ollama (local, no key) — model must support tool/function calling
ollama pull qwen3:8b
betterwright models ollama
betterwright exec "inspect example.com" --model qwen3:8b
# pin if another source also exposes that id:
betterwright exec "inspect example.com" --model ollama/qwen3:8b

# vLLM (local OpenAI-compatible server with tool calling enabled)
betterwright exec "inspect example.com" --model vllm/<served-model-id>

# OpenRouter — canonical author/model id; pin only on collisions
OPENROUTER_API_KEY=… betterwright exec "inspect example.com" \
  --model anthropic/claude-sonnet-5
OPENROUTER_API_KEY=… betterwright exec "inspect example.com" \
  --model openrouter/anthropic/claude-sonnet-5

# Any other OpenAI-compatible /v1 base URL
BETTERWRIGHT_MODEL_API_KEY=… betterwright exec "inspect example.com" \
  --base-url https://models.example/v1 --model <model-id>

Preset endpoints and environment

SourceDefault base URLKey env varNotes
OpenRouterhttps://openrouter.ai/api/v1OPENROUTER_API_KEY (required for runs)Listing can work without a key; execution needs one
Ollamahttp://127.0.0.1:11434/v1OLLAMA_API_KEY (optional)No key for local defaults
vLLMhttp://127.0.0.1:8000/v1VLLM_API_KEY (optional)Start the server with tool-calling flags (below)
Customfrom --base-url or BETTERWRIGHT_MODEL_BASE_URLBETTERWRIGHT_MODEL_API_KEY (optional)--base-url alone pins the source

Override a preset URL with OPENROUTER_BASE_URL, OLLAMA_BASE_URL, or VLLM_BASE_URL. Use --api-key-env MY_KEY when the key lives under another name (CLI flags never accept raw key values). BetterWright refuses to send a key to a non-loopback http:// URL unless --allow-insecure-model-endpoint is set; HTTPS and loopback HTTP are fine without it.

Environment-driven defaults for scripts and hosts:

VariableRole
BETTERWRIGHT_MODELDefault --model when the flag is omitted
BETTERWRIGHT_MODEL_BASE_URLDefault custom endpoint base URL
BETTERWRIGHT_MODEL_API_KEYKey for that custom endpoint
BETTERWRIGHT_MODEL_PROTOCOLchat (default) or responses
BETTERWRIGHT_CLAUDE_MODEL / BETTERWRIGHT_CODEX_MODEL / BETTERWRIGHT_GROK_MODELDefaults shown by betterwright models for each native source

Protocol and tool calling

Chat Completions (--protocol chat, the default) is the widest common surface. Use --protocol responses only when the server implements the Responses API.

Tool / function calling is required. The loop drives the browser through tools (browser, done, optional login / ask / handoff). A text-only model will not work, even if chat completions succeed.

  • Ollama — use a model that advertises tool support (and a recent Ollama build). Prefer ids you have verified with ollama run tool demos, or pin explicitly with --model ollama/<id>.
  • vLLM — start the server with --enable-auto-tool-choice and the --tool-call-parser required by the served model.
  • OpenRouter / custom — pick models known to support tools; partial OpenAI compatibility without tools is not enough.

Native Claude, Codex, and Grok

Native sources are chosen from the model id family, or by an explicit source/ prefix:

Family patternSourceAuth
starts with claudeClaudeANTHROPIC_API_KEY + optional peer dep @anthropic-ai/sdk
starts with gpt, o+digit, chatgpt, or codexCodexbetterwright auth --login codex, or OPENAI_API_KEY / CODEX_BASE_URL
starts with grokGrokbetterwright auth --login grok, or GROK_API_KEY / XAI_API_KEY

Defaults: Claude claude-opus-4-8, Codex gpt-5.6-sol (or the model stored in ~/.codex/config.toml), Grok grok-4.3. Qualify when needed (codex/gpt-5.6-sol). Base-URL overrides: CODEX_BASE_URL, GROK_BASE_URL / XAI_BASE_URL. xAI may gate OAuth API access by subscription tier; a 403 usually means switch to an API key.

Signing in (betterwright auth)

betterwright auth --login codex     # opens "Sign in to Codex with ChatGPT"
betterwright auth --login grok      # opens the xAI sign-in
betterwright auth --status          # who am I signed in as?

auth --login runs the provider's OAuth 2.0 PKCE flow: it starts a loopback callback server, opens the consent page in your default browser, and stores tokens (Codex in ~/.codex/auth.json, shared with the codex CLI; Grok in ~/.grok/auth.json). No API key is pasted and no router is required — BetterWright refreshes the access token per request.

Migration notes

Older docs and scripts used adapter nicknames and a separate id flag. Current CLI behavior:

OldNew
--model claude / codex / grokReal id: --model claude-opus-4-8, gpt-5.6-sol, grok-4.3
--model-id <id>Merged into --model <id> (the old flag errors with a hint)
--provider ollama (or similar)Put the source in the id only when needed: --model ollama/<id>

Troubleshooting model selection

SymptomWhat to try
Unknown model "codex" (or claude / grok)Pass a real id, not the old adapter name
No available model source exposes "…"Run betterwright models; start Ollama/vLLM; set OPENROUTER_API_KEY; or pin source/id / --base-url
available from multiple sources: …Pick one of the listed selectors, e.g. ollama/qwen3:8b
Ollama listed empty / connection errorsConfirm ollama serve is up; override with OLLAMA_BASE_URL if not on the default port
Model answers in prose but never calls toolsSwitch to a tool-calling model; for vLLM enable auto tool choice + parser
Key over plain HTTP refusedUse HTTPS, loopback, or pass --allow-insecure-model-endpoint deliberately

Programmatic use

runAgentTask is the same loop the CLI uses:

import { runAgentTask } from "betterwright/agent";

const result = await runAgentTask({
  task: "log in to example.com and download this month's invoice",
  model: "claude-opus-4-8",       // actual model id, or your own model object
  maxDurationMs: 30 * 60 * 1000,   // configurable; there is no fixed step cap
  maxTranscriptChars: 1_000_000,   // bound accumulated model/tool context
  guardrails: { confirmBeforePurchase: true },
  onStep: ({ step, tool, note }) => console.error(`[${step}] ${tool}: ${note}`),
});
console.log(result.answer, result.proof);
console.log(result.toolCalls, result.usage, result.durationMs);
// 7, { inputTokens, outputTokens, cacheReadTokens, cacheWriteTokens, context }, 11400

Pass an askUser handler to let the loop ask the human mid-task — this is how the interactive console wires the ask tool. Given a handler, runAgentTask exposes an ask tool to the model; without one it runs fully autonomously.

const result = await runAgentTask({
  task: "book me a table for dinner tonight",
  model: "gpt-5.6-sol",
  askUser: async ({ question, options }) => {
    // Route to your UI/host; return the user's answer as a string.
    return await promptTheUser(question, options);
  },
});

runAgentTask constructs and closes its own browser unless you pass one:

import { BetterWright } from "betterwright";
import { runAgentTask } from "betterwright/agent";

const browser = new BetterWright(); // encrypted local vault + `login` tool by default
await runAgentTask({ task: "…", browser, model: "claude-opus-4-8" });
await browser.close();

The loop exposes a selector-free login tool backed by the same URL-matched worker fill as MCP/Pi's browser_login. Generated signup/rotation passwords remain pending until a later browser step verifies success and commits them, so failed forms do not leave stale saved items. Pass vault: false to disable the tool or a custom adapter to replace the local store.

Bring your own model or agent

The model is a small pluggable interface — this is how you plug in anything the three built-in adapters don't cover:

const myModel = {
  name: "my-model",
  async complete({ system, messages, tools }) {
    // messages is a neutral transcript (see types/agent.d.ts); tools is
    // [{ name, description, parameters }]. Return the model's reply:
    return { text: "…", toolCalls: [{ id, name, input }], stopReason };
  },
};

await runAgentTask({ task: "…", model: myModel });

The harness handles the browser tools, the observe/act/verify discipline, proof capture, and the operator guidance; your adapter only has to translate the neutral transcript to and from your provider's wire format.

For a preset-compatible endpoint, prefer endpointModel:

import { endpointModel, runAgentTask } from "betterwright/agent";

const model = endpointModel({
  source: "ollama",          // openrouter | ollama | vllm | custom
  model: "qwen3:8b",
  // baseURL: "http://127.0.0.1:11434/v1",  // optional override
});
await runAgentTask({ task: "…", model });

Or pass a string the same way the CLI does ("ollama/qwen3:8b", "gpt-5.6-sol", …). The lower-level openaiModel({ baseURL, model, apiKey }) remains available when you need full control of the request shape.

What the loop does

Each turn: the model sees the task, the operator guidance from agentSystemPrompt, and the tools (browser, done, login when a vault is present, and ask when an askUser handler is present). It calls browser with async Playwright JavaScript; BetterWright runs it and feeds back a compact JSON observation (ok, result, console, pages, challenges, skills, warnings, screenshots). The loop ends when the model calls done (or answers in prose) — there is no step cap. Screenshots are captured as artifacts and their paths surfaced; the last proof screenshot is returned on the result.

A browser call can also end the task in the same turn: when the model's code returns { finalAnswer: "…" } (from an ok run), the harness records that as the answer and finishes without spending another model round-trip. The preamble teaches the model to use this single-call shape for read-only tasks — navigate, extract, compute, capture proof, and return finalAnswer in one call — which makes a simple lookup cost one model turn instead of two.