Backend integration
Browser Skills: taught web actions under human approval
Teach an agent one business action on a web system you own — sign in as a step, bind the values that change to typed parameters, mark the commit step, add checks, pick the result — then let agents perform it through browser_skill:<slug>.run / dry_run under a per-skill approval policy (approve before the run or at the commit point on a pre-commit screenshot, four-eyes, dry-run only), with idempotency, verification and honest failures.
A Browser Skill is one business action on a web system your organisation owns or is authorised to operate — record a payment, create a ticket, approve a leave request, look up a shipment — taught once in the console's Recorder and performed many times by agents with different parameter values, always on the sandboxed browser and always under a human-approval policy you set per skill. Masar is the data layer: it reads any website into a structured source. Browser Skills are the action layer: they act on a system. Replay is deterministic — the recorded steps run with scored fallback selectors and text anchors; a model only supplies typed parameter values and a reason. The model never sees a credential, nothing is committed without the policy's human decision, and a failure is reported as what it is.
1. What a skill is
A skill is a declarative definition the Recorder produces: the start URL and the hosts the recording touched, the steps (an auth phase for the sign-in and a main phase for the action), typed parameters bound to the values that change, exactly one commit step (the click or key press that submits), the checks the confirmation page must pass, the result fields to read back, optional compensation steps, and the authentication metadata. Every name below belongs to the administrator who taught it — the example records a supplier payment behind a login.
{
"start_url": "https://finance.example.internal/portal/login",
"target_hosts": ["finance.example.internal"],
"steps": [
{ "id": "s1", "phase": "auth", "kind": "navigate", "url": "https://finance.example.internal/portal/login" },
{ "id": "s2", "phase": "auth", "kind": "type", "target": { "selector": "#email", "candidates": [] },
"value": "ops@example.com", "secret": false },
{ "id": "s3", "phase": "auth", "kind": "type", "target": { "selector": "#password", "candidates": [] },
"value": null, "secret": true, "secret_ref": "password" },
{ "id": "s4", "phase": "auth", "kind": "click", "target": { "selector": "button[type=submit]", "candidates": [], "text": "Sign in" } },
{ "id": "s5", "kind": "navigate", "url": "https://finance.example.internal/portal/payments/new" },
{ "id": "s6", "kind": "type", "target": { "selector": "#amount", "candidates": ["[name=amount]"] }, "value": null, "param": "amount" },
{ "id": "s7", "kind": "type", "target": { "selector": "#note", "candidates": ["[name=note]"] }, "value": null, "param": "note" },
{ "id": "s8", "kind": "click", "target": { "selector": "#submit-payment", "candidates": ["button[type=submit]"], "text": "Submit payment" } }
],
"params": [
{ "name": "amount", "type": "number", "label": "Amount (AED)", "required": true, "min": 0.01, "max": 100000, "example": "1250" },
{ "name": "note", "type": "string", "label": "Note", "required": false, "max_length": 200, "example": "Invoice 4471" }
],
"commit_step_id": "s8",
"assertions": [
{ "kind": "url_contains", "value": "/portal/payments/" },
{ "kind": "text_present", "value": "Payment recorded" }
],
"result_fields": [
{ "name": "reference", "selector": "#ref", "candidates": ["[data-testid=reference]"], "attr": "text", "type": "string" }
],
"compensation": [],
"auth": { "kind": "credentials", "secret_refs": ["password"] }
}| Part | Rules enforced at save (and re-checked before any run) |
|---|---|
steps[] | Up to 60 of navigate · click · type · select · check · press · wait · scroll. A secret `type` step (a password field) carries `value: null` and a `secret_ref` and can never be bound to a parameter. A `navigate` URL may carry `{{param}}` only after a literal host and only for parameters with `allow_in_url`. No selector may contain `{{`. |
params[] | Up to 20, lower-case identifiers, typed string · number · integer · boolean · enum · date with min / max, length, pattern, enum values and an example (what the admin typed while teaching — a hint, never a value). Each becomes one JSON-schema property of the generated tools; a declared but unbound parameter is never asked of the model. |
commit_step_id | The one main-phase click or press that changes state. Required by `at_commit_point` and `dry_run_only`; a skill with state-changing steps and no commit marker can only be saved as `none_for_idempotent_reads` (a lookup form). |
assertions[] | Up to 10 checks run after the commit: url_contains · text_present · element_present · element_text (equals / contains). At least one is recommended. |
result_fields[] | Up to 20 fields picked on the confirmation page (selector + fallbacks, text / href / value, string / number / integer) — the skill’s output schema. |
compensation[] | Up to 10 steps that cancel or undo (Cancel, Delete draft). They run only after a commit whose checks failed, in a fresh session, never bind parameters or type secrets, and are reported as what ran — never as “undone”. |
auth | none · credentials · cookies — names only (`secret_refs`). Signing in while teaching is just a step; on save the admin chooses credentials (write-only values per `secret_ref`) or “keep me signed in” (the recorder session’s cookies pinned to the skill’s hosts). Both are AES-256-GCM at rest and decrypted only inside the runtime executor. |
2. The approval policy — per skill
Human-in-the-loop for an action on a website is a policy on the skill, not a prompt. It is a floor: an agent manifest can never relax it, and a tenant approval policy matching browser_skill:<slug>.run can only tighten it.
| approval | What the agent may do | When the human decides | The card shows |
|---|---|---|---|
always_before_run | nothing before approval | before any step runs | skill, values, start URL, risk, the agent’s reason |
at_commit_point (default) | fill the form up to the commit step | after the form is filled, before the commit click | the pre-commit screenshot, the values read back from the page, the landed URL, risk, reason, the steps so far, “what will be pressed next” |
dry_run_only | only dry_run (the run tool is not exposed at all) | never — nothing is ever committed by an agent | — (a dry run returns the screenshot and values to the model) |
none_for_idempotent_reads | run a skill that has no commit step | none from the policy (the agent’s own HITL rules and tenant approval policies still apply) | — |
risk_level (low · medium · high · critical), approver_role (owner · admin · manager · member, default admin), sla_minutes (5 min to one week, default 24 h), required_approvals (1–3; 2 is four-eyes — distinct approvers, counted by the platform's approvals machinery) and allow_edit_params (may an approver change the values before approving; edited values are re-validated against the parameter schema) complete the policy. none_for_idempotent_reads is accepted only for a skill with no commit step and risk_level: low.
3. Teach it, test it, save it
Teaching happens in the console (Build → Integrations → Automation → "Browser action — taught", or "Teach a browser action" on an agent's page): the Recorder opens the system in a sandboxed browser, the administrator signs in (the password is detected and masked), does the action once, binds the values that change to parameters, marks the commit step, adds checks, picks the result fields and saves with the attestation I own or am authorised to operate this system. The save is one call you can also script once you hold a recording.
curl -sS -X POST "https://<api-host>/v1/browser-skills" \
-H "Authorization: Bearer hive_…" -H "X-Tenant-Id: <tenant>" \
-H "Content-Type: application/json" \
-d '{
"name": "Record a payment", "slug": "record-payment",
"description": "Records a supplier payment in the finance portal and returns its reference.",
"definition": { …the definition above… },
"policy": { "approval": "at_commit_point", "risk_level": "high", "approver_role": "admin",
"sla_minutes": 1440, "required_approvals": 2, "allow_edit_params": false },
"auth": { "mode": "credentials", "secrets": { "password": "<write-only, encrypted, never returned>" } },
"recorder_session_id": "<the Recorder session the skill was taught in>",
"owner_attested": true,
"enabled": true
}'
# → 201 { skill } (the row never carries encrypted_credentials; credential_meta = { kind, secret_refs })
# → 400 BROWSER_SKILL_INVALID a cross-field rule failed — the message names it
# → 400 BROWSER_SKILL_ATTESTATION_REQUIREDTest without committing
A dry run prepares the skill in a fresh browser context — session restored or the sign-in replayed if a login wall shows, the main steps replayed up to (never through) the commit step, every bound field read back, a screenshot taken — and returns what it saw. Nothing is submitted, no approval is opened, no receipt is written.
curl -sS -X POST "https://<api-host>/v1/browser-skills/<skillId>/dry-run" \
-H "Authorization: Bearer hive_…" -H "X-Tenant-Id: <tenant>" \
-H "Content-Type: application/json" \
-d '{ "params": { "amount": 1250, "note": "Invoice 4471" } }'
# → 200 { run: { mode: "dry_run", status: "succeeded",
# precommit: { url, title, screenshot_key, values: { amount: "1250", note: "Invoice 4471" } },
# steps: [{ id, kind, label, status, screenshot_key }] } } — nothing was submitted
# → 503 BROWSER_SKILL_UNAVAILABLE when no browser sandbox is configured4. What happens when an agent runs it
Under at_commit_point the run's approval gate first performs a side-effect-free prepare (the same steps as a dry run) and opens the approval with a frozen pre-commit preview; the run pauses on the platform's checkpointed interrupt — any worker may resume it, minutes or a day later. A prepare that cannot land (a selector not found, an expired session, invalid parameters) fails the tool call in the same turn; nothing is parked in the inbox for an action that cannot happen.
{
"requested_action": {
"browser_skill": {
"skill_id": "<id>", "slug": "record-payment", "name": "Record a payment",
"mode": "run", "approval": "at_commit_point", "risk_level": "high",
"target_url": "https://finance.example.internal/portal/payments/new",
"title": "New payment", "hosts": ["finance.example.internal"],
"screenshot_key": "browser-skills/<tenant>/<run>/7.jpeg", "run_id": "<run>",
"values": [
{ "name": "amount", "label": "Amount (AED)", "type": "number", "value": "1250" },
{ "name": "note", "label": "Note", "type": "string", "value": "Invoice 4471" }
],
"steps": [{ "id": "s5", "kind": "navigate", "label": "/portal/payments/new", "status": "succeeded", "matched": -1,
"screenshot_key": "browser-skills/<tenant>/<run>/5.jpeg", "at": "…" }, "…"],
"commit_label": "Submit payment", "checks": 2,
"reason": "Supplier invoice 4471 approved by finance on 1 Oct"
}
}
}After the decision the worker does not click on a tab it held open. It re-prepares in a fresh context with the approved (or edited) values, compares the read-back to what the approver saw — same URL path, every bound field equal to the value being committed; a mismatch fails precommit_mismatch without clicking — writes a committing receipt before the click, performs the commit step and the post-commit steps, runs the checks in order, reads the result fields and answers the model:
{
"ok": true, "run_id": "<run>", "status": "succeeded", "committed": true,
"verification": [
{ "kind": "url_contains", "ok": true, "detail": null },
{ "kind": "text_present", "ok": true, "detail": null }
],
"result": { "reference": "PAY-20261002-0147" },
"steps": [{ "id": "s8", "kind": "click", "screenshot_key": "browser-skills/<tenant>/<run>/8.jpeg" }, "…"],
"cite": "browser_skill:record-payment#<run>"
}| Outcome | Run status | Compensation | The model hears |
|---|---|---|---|
| parameters invalid | — (precheck) | — | INVALID_ARGS with the field and rule |
| selector not found before the commit | failed | — (nothing committed) | not_found + the step and the selectors tried |
| login wall still showing after the sign-in steps | failed | — | session_expired |
| page differs from the approved preview | failed | — (nothing committed) | precommit_mismatch + which field or URL |
| rejected / responded / expired | rejected | — (nothing committed) | { rejected: true, feedback } |
| committed, a check failed | verification_failed | recorded compensation steps run in a fresh session; outcome recorded on the run | verification_failed + screenshot reference + whether compensation ran |
| crash between the receipt and the verification | committing / committed | none automatic | unknown_outcome — check the system; do not retry blindly |
Idempotency. Every run row carries a key derived from the tenant, the agent run, the tool key and the canonical arguments, unique per tenant. A model that calls the same skill with the same values twice in one conversation gets the stored result the second time (duplicate: true); a resumed or retried turn finds its row; a row found in committing or committed is never re-driven. Different values are a different key — and a second approval.
5. Grant it to agents
Declare the skill in the agent manifest and grant it — both default-deny, like every connector. allow defaults to both ops; under a dry_run_only policy run is not resolved even when granted.
{
"tools": {
"browser_skills": [
{ "skill": "record-payment", "allow": ["run", "dry_run"] },
{ "skill": "lookup-shipment", "allow": ["run"] }
]
}
}POST /v1/browser-skills/:id/grants with agent_id runs the same attach service as the hub (internal connector.kind = "browser_skill"): tool grants (source: browser_skill, browser_skill_id), the draft-manifest merge and publish in one transaction (Connectors, tools & grants). The public POST /v1/integrations/attach does not accept browser skills — they are never catalog cards. GET · POST · DELETE /v1/browser-skills/:id/grants manage grants per skill.
curl -sS -X POST "https://<api-host>/v1/integrations/attach" \
-H "Authorization: Bearer hive_…" -H "X-Tenant-Id: <tenant>" \
-H "Content-Type: application/json" \
-d '{
"connector": { "kind": "browser_skill", "id": "<skillId>" },
"tools": ["run", "dry_run"],
"targets": [{ "agent_id": "<agentId>" }],
"publish": false
}'
# the skill's own policy already gates run; a tenant approval policy can only tighten it| Tool | What it does | Approval |
|---|---|---|
browser_skill:<slug>.run | Performs the taught action with the parameter values and a `reason`. Arguments are generated from the definition: one typed property per bound parameter, validation mirrored. Privileged. | the skill’s policy — always_before_run or at_commit_point are forced, non-relaxable |
browser_skill:<slug>.dry_run | Fills everything up to the commit point and returns the pre-commit screenshot reference, the read-back values and the step list — without submitting. | the agent’s normal policy |
Tool descriptions are generated from the skill — its name and description, the hosts, the parameters, the result fields and when a human decides — so the model is told the truth about the gate before it calls. The result carries a cite (browser_skill:<slug>#<run>) the agent quotes; the console resolves it to the run's step-by-step evidence.
6. Routes, posture and errors
| Route | Role | Purpose |
|---|---|---|
GET /v1/browser-skills | member | List skills (definition, policy, auth kind, credential names/counts, grants count, last run) — never a credential. |
POST /v1/browser-skills | admin | Create from a recording; attestation required; credentials write-only. |
GET · PATCH · DELETE /v1/browser-skills/:id | member · admin · admin | Detail with grants and last runs; update name / description / policy / enabled / node / credentials (definition edits = teach again); delete (runs and grants cascade). |
POST /v1/browser-skills/:id/dry-run | admin | Prepare without committing; 503 BROWSER_SKILL_UNAVAILABLE without a sandbox. |
GET /v1/browser-skills/:id/runs · GET /v1/browser-skills/runs?agent_run_id= · GET /v1/browser-skills/runs/:runId | member | Runs of a skill; runs an agent run made (the Workspace timeline); one run with steps, pre-commit state, verification, result, compensation. |
GET · POST · DELETE /v1/browser-skills/:id/grants[/:grantId] | admin | Grants per skill (agent or org node; run · dry_run · *). |
GET /v1/browser-skills/artifacts/* | member | Step screenshots (JPEG); the key must start with browser-skills/<tenant>/ — another tenant’s key is a 404. |
- Ownership, not robots.txt. A skill performs an action the owner would perform by hand on their own system, so robots.txt is not consulted; the ownership attestation is mandatory and audited with who and when.
- Every URL — the start URL, each
navigatestep, every landing and redirect — passes the platform SSRF policy and the tenant's residency guard; private hosts are refused. A landing outside the skill's hosts ends the runout_of_scope. Downloads are denied and the file chooser intercepted. - Parameters land only in step values and in declared URL path / query components (URL-encoded; the host is literal). Values are capped at 2 000 characters, stripped of control characters and bidirectional overrides, coerced to their declared type and checked against enum / pattern / min / max before anything runs. Selectors are never templated.
- Credentials are decrypted only inside the executor for the sign-in phase, zeroed after use and never present in run rows, screenshots' metadata, the approval preview, tool results, events, audit rows or logs; every decryption is audited. The executor refuses to take a screenshot while a password field is on the page.
- A result field is data, scanned by the same tool-output guardrail as every tool result; the approval card renders values as text, never as HTML.
| Code | Status | When |
|---|---|---|
BROWSER_SKILL_NOT_FOUND | 404 | Unknown skill in this tenant (another tenant’s skill is a 404, never a 403). |
BROWSER_SKILL_INVALID | 400 | A cross-field rule of the definition or the policy failed; the message names it. |
BROWSER_SKILL_ATTESTATION_REQUIRED | 400 | owner_attested was not true. |
BROWSER_SKILL_UNAVAILABLE | 503 | No browser sandbox is configured (teaching, dry runs and runs all need it). |
BROWSER_SKILL_RUN_FAILED | 502 | A run did not land; extensions.code is one of not_found · step_failed · session_expired · precommit_mismatch · verification_failed · out_of_scope · unknown_outcome. |
Evidence: audit actions browser_skill.create | update | delete | grant | revoke | dry_run, browser_skill.run.prepared | committed | succeeded | failed | rejected | compensated and browser_skill.credential_access (ids, hosts, step counts, statuses and reasons — never a credential, a page's content or a screenshot); usage kind browser_skill_run with mode, status, steps and screenshot count, beside the ordinary tool-call record of the agent step.