Agentic Workforce ME Developer PortalDocs 1.10 · Widget 0.5.0

Backend integration

Masar: any website as a structured source

Turn any website into a structured, queryable source for your agents without an API: define a source (URLs, extractors, your JSON Schema, schedule), discover and test it, refresh into versioned snapshots with per-field provenance and diffs, and grant the generated masar:<slug>.query / latest / refresh tools — refresh always needs a human approval.

Masar turns any website into a structured, queryable source for your agents — no API, no database access, no export from the other side. You point it at a page, say what the information looks like (a JSON Schema you own) and how to read it (CSS selectors over the rendered DOM, the JSON the page itself loads, the text layer of a PDF, a vision model over a screenshot, or that same model as OCR for scanned images), and Agentic Workforce ME keeps a versioned snapshot with per-field provenance and a diff between refreshes. Agents consume it as three generated tools — masar:<slug>.query, .latest and .refresh — under the same grants and approval policies as every other tool. A public tender board, a supplier price list, a ministry notice board, a regulator's PDF circular and an exchange-rate widget are all just sources; nothing in the engine knows which is which.

1. What a source is

A source is a declarative definition: entry URLs and a crawl scope, an ordered list of extractors (deterministic ones first), the output schema, the key fields a diff aligns rows on, a refresh schedule and politeness limits. Every field name below belongs to you — the example is a supplier price list whose table is filled from an XHR endpoint, with a DOM extractor as the fallback.

JSON
{
  "slug": "supplier-prices",
  "name": "Acme supplier price list",
  "entry_urls": ["https://shop.example.com/catalogue"],
  "scope": { "allowed_hosts": ["shop.example.com"], "follow_links": "none", "max_pages": 1 },
  "renderer": "auto",
  "extractors": [
    { "kind": "network", "match": { "url_includes": "/api/products" }, "pick": "$.items",
      "target": "items", "map": { "sku": "code", "name": "title", "price": "price.amount",
                                  "currency": "price.currency", "in_stock": "available" } },
    { "kind": "dom", "fields": {}, "list": { "target": "items", "item": "table tbody tr",
      "fields": { "sku": { "selector": "td:nth-child(1)" }, "name": { "selector": "td:nth-child(2)" },
                  "price": { "selector": "td:nth-child(3)" } } } }
  ],
  "output_schema": {
    "type": "object",
    "properties": {
      "items": { "type": "array", "items": { "type": "object",
        "properties": { "sku": { "type": "string" }, "name": { "type": "string" },
                        "price": { "type": "number" }, "currency": { "type": "string" },
                        "in_stock": { "type": "boolean" } },
        "required": ["sku", "name", "price"] } }
    },
    "required": ["items"]
  },
  "key_fields": ["items.sku"],
  "schedule": "0 6 * * 1-5", "timezone": "Asia/Dubai",
  "politeness": { "min_delay_ms": 2000, "respect_robots": true },
  "limits": { "timeout_ms": 30000, "vlm_max_pages": 0, "ocr_max_images": 0, "max_age_seconds": 172800 },
  "contains_pii": false,
  "terms_attested": true
}
ExtractorReadsNeeds the browser renderer
domCSS selectors over the (rendered) HTML — scalar fields and a list (`item` rows inside `target`); `attr` picks text, html, href, src or any attribute.no (fetch works for static pages; browser for JS-rendered ones)
networkXHR / fetch JSON responses the page itself received — `match` the URL, `pick` a JSON path, `map` schema fields to item keys. Bodies are captured, never replayed.yes
readabilityThe main article text (boilerplate removed) into one field, optionally the title into another.no
pdfLinked or entry PDFs: text layer first; pages without one are rasterised and read by the vision model.no
vlmA vision model over the page screenshot, restricted to the schema fields still missing (`mode: fallback`) or to listed fields (`always`); every value comes back with a confidence and a quote.yes
ocrTranscription of images matching a selector / URL filter into a text field (the vision model as OCR; `engine: vlm`).no (images are fetched directly)

renderer: auto fetches first and switches to the browser when the page is a JavaScript shell or an extractor needs network capture or a screenshot. Vision and OCR are opt-in per source, capped per run (vlm_max_pages, ocr_max_images, default 5) and metered. Starter templates (GET /v1/masar/templates) are deliberately domain-neutral — listing / table, article / notice, document / PDF, dashboard / KPIs — and discovery renames their fields from the page's own headers and keys.

2. Create it: discover, test, save

The console wizard (Build → Integrations → Masar → "Add a source": URL → discover → extractors → schema → test extraction → schedule and attestation) is a thin client over three calls you can also script.

Discover what the page really loads

Shell
curl -sS -X POST "https://<api-host>/v1/masar/discover" \
  -H "Authorization: Bearer hive_…" -H "X-Tenant-Id: <tenant>" \
  -H "Content-Type: application/json" \
  -d '{ "url": "https://shop.example.com/catalogue" }'
# → { tables[], lists[], json_endpoints[{url, pick, item_keys}], images[], pdfs[],
#     suggested_extractors[], suggested_schema, suggested_template, warnings[] }

Test a draft without persisting anything

Shell
curl -sS -X POST "https://<api-host>/v1/masar/test-extract" \
  -H "Authorization: Bearer hive_…" -H "X-Tenant-Id: <tenant>" \
  -H "Content-Type: application/json" \
  -d '{ "definition": { …entry_urls, scope, renderer, extractors, output_schema, key_fields… },
        "max_pages": 1 }'
# → { data, provenance, status, record_count, coverage: {found, expected}, pages[], warnings[],
#     screenshot_keys[], vlm_calls, tokens_in, tokens_out }   (nothing persisted)

Each value in provenance names its extractor, URL and fetch time plus the selector path, the captured request and JSON path, the PDF page or the screenshot region it came from, and a confidence (1 for deterministic extractors). A field the page does not have is null with a reason (not_found, invalid, extractor_unavailable, cap_reached, low_confidence) — never a guess.

Save

POST /v1/masar/sources with the definition plus slug, name, schedule, timezone and terms_attested: true — the creator confirms the site's content may be used this way, and the attestation (who, when) is stored on the source. The schema is compiled before anything is persisted (400 MASAR_INVALID_SCHEMA); entry URLs that fail the SSRF policy, the scope or robots.txt are refused at definition time (400 MASAR_URL_REFUSED); a cron that fires more often than every 15 minutes is a validation error.

3. Refresh, runs and snapshots

Shell
curl -sS -X POST "https://<api-host>/v1/masar/sources/<sourceId>/refresh" \
  -H "Authorization: Bearer hive_…" -H "X-Tenant-Id: <tenant>" \
  -H "Content-Type: application/json" -d '{ "reason": "weekly price check" }'
# → 202 { run: { id, status: "queued", trigger: "manual", … } }
# → 409 MASAR_RUN_IN_PROGRESS while a run of this source is still running

A refresh is a worker job: it crawls the pages in scope (one at a time, with the politeness delay, robots.txt checked per host), runs the extractors, validates and coerces the result against your schema, and — only when the content hash changed — writes snapshot version + 1 with the diff against the previous version (added / removed / changed rows keyed on key_fields, changed scalar fields). Scheduled sources refresh on their cron; an agent may ask for one through masar:<slug>.refresh, which always pauses for a human decision.

RouteRolePurpose
POST /v1/masar/discoveradminProbe one URL; report + suggested extractors and schema draft.
POST /v1/masar/test-extractadminDry run of a draft definition (up to 3 pages); nothing persisted but preview screenshots.
GET /v1/masar/sourcesmemberList (counts, last run, latest version, grants).
POST /v1/masar/sourcesadminCreate; schema compiled first; terms attestation required.
GET · PATCH · DELETE /v1/masar/sources/:idmember · admin · adminDetail with grants and latest snapshot summary; update (scheduler re-synced); delete (runs, snapshots and grants cascade).
POST /v1/masar/sources/:id/testadminTest extraction on the saved definition.
POST /v1/masar/sources/:id/refreshadmin202 with the queued run; 409 MASAR_RUN_IN_PROGRESS.
GET /v1/masar/sources/:id/runs · GET /v1/masar/runs/:idmemberRuns newest first; run detail (pages crawled / blocked, requests captured, vision calls, tokens, warnings).
GET /v1/masar/sources/:id/snapshots · GET /v1/masar/snapshots/:idmemberSnapshot list with diff summaries; one snapshot with data, provenance and diff.
GET /v1/masar/sources/:id/snapshots/:version/diff?against=memberDiff between any two versions.
GET /v1/masar/artifacts/*memberScreenshot bytes; keys are tenant-prefixed and checked.
GET /v1/masar/templatesmemberThe domain-neutral starter templates.

4. Grant it to agents

Declare the source in the agent manifest and grant it — both default-deny, like every connector. allow defaults to query and latest; list refresh only for agents that may ask for fresher data.

JSON
{
  "tools": {
    "masar": [
      { "source": "supplier-prices", "allow": ["query", "latest"] },
      { "source": "regulator-circulars", "allow": ["query", "latest", "refresh"] }
    ]
  }
}

The composite POST /v1/integrations/attach step accepts connector.kind = "masar" and creates the tool grants (source: masar, masar_source_id), the approval policies and the draft-manifest merge in one transaction (Connectors, tools & grants).

Shell
curl -sS -X POST "https://<api-host>/v1/integrations/attach" \
  -H "Authorization: Bearer hive_…" -H "X-Tenant-Id: <tenant>" \
  -H "Content-Type: application/json" \
  -d '{
    "connector": { "kind": "masar", "id": "<sourceId>" },
    "tools": ["query", "latest", "refresh"],
    "targets": [{ "agent_id": "<agentId>" }],
    "approvals": [{ "tool": "refresh", "approver_role": "admin" }],
    "publish": false
  }'
# refresh is force-gated in the runtime anyway; the policy row makes the card explicit
ToolWhat it doesApproval
masar:<slug>.queryFilters the latest snapshot: `filters[{field, op: eq | neq | contains | gte | lte | in, value}]` over the schema’s scalar fields (up to 10), `fields` projection, `list` when the schema has several lists, `limit` ≤ 100, `offset`. Never crawls.the agent’s normal policy
masar:<slug>.latestThe latest snapshot’s version, fetch time, age, `stale` flag and diff summary — "what changed since last time". Never crawls.the agent’s normal policy
masar:<slug>.refreshEnqueues a crawl and answers `{ queued: true, run_id }` — the next `query` reads the new snapshot.always required (forced, not relaxable)
JSON
{
  "source": "supplier-prices",
  "snapshot": { "id": "<id>", "version": 12, "fetched_at": "2026-10-01T06:00:04Z",
                "age_seconds": 5400, "stale": false },
  "records": [
    { "cite": "masar:supplier-prices#12/0", "sku": "A-100", "name": "Hex bolt M8",
      "price": 1.25, "currency": "AED", "in_stock": true }
  ],
  "total": 1, "truncated": false, "provenance_available": true
}

Tool descriptions are generated from the source — its name, description and the schema's field names and descriptions, plus the latest version and fetch time — so the model knows the vocabulary without any hard-coded domain. Every record carries a cite (masar:<slug>#<version>/<row>) the agent quotes; a snapshot older than the source's max_age_seconds is returned with stale: true and the agent is told to ask for a refresh (which needs a person) rather than guess. Field-level provenance stays in the console's snapshot browser; snapshot values reach the model as data, scanned by the same tool-output guardrail as every other tool result.

5. Posture, limits and errors

  • robots.txt is honoured and cannot be switched off; a disallowed URL is recorded as blocked (and audited), the run continues with what is allowed. Crawl-delay is respected up to 30 s.
  • The crawler identifies itself (HiveMasar/1.0 (+https://agenticworkforce.me/masar)), fetches one page at a time per source with at least one second between pages (default 1.5 s), caps pages (default 20, hard cap 200), depth (3) and run time (10 min), and never solves a bot challenge, captcha or paywall — such pages are recorded with blocked_reason: challenge.
  • Every URL — entry, followed link, image, PDF — passes the platform SSRF policy and the tenant's data-residency guard; captured XHR bodies are read, never re-requested.
  • Vision / OCR calls are opt-in, capped and metered (masar_vlm_call, tokens); a run that hits the cap stops calling the model and says so. A snapshot is capped at 2 MB / 2 000 records (422 MASAR_BUDGET_EXCEEDED, truncated and flagged).
  • A source may be flagged contains_pii; the wizard warns on PII-looking field names and the snapshot browser redacts those columns for non-admins. Snapshots follow the tenant retention purge.
CodeStatusWhen
MASAR_SOURCE_NOT_FOUND404Unknown source in this tenant (another tenant’s source is a 404, never a 403).
MASAR_INVALID_SCHEMA400The output schema does not compile or has no object root.
MASAR_URL_REFUSED400An entry URL fails the SSRF policy, the scope or robots.txt at definition time.
MASAR_RUN_IN_PROGRESS409A refresh was requested while a run of the source is running.
MASAR_RENDERER_UNAVAILABLE503An extractor needs the browser renderer and the browser sandbox is not configured.
MASAR_VLM_UNAVAILABLE503A vision extractor is configured but no vision-capable model is.
MASAR_BUDGET_EXCEEDED422The snapshot exceeded its size or record cap.

Evidence: audit actions masar.source.create | update | delete, masar.discover, masar.test_extract, masar.refresh.requested, masar.run.completed | failed and masar.navigation.blocked (ids, counts, public URLs and reasons — never page content); usage kinds masar_page and masar_vlm_call plus the model tokens of vision calls.

Three sources, one engine

SourceShapeExtractorsTemplate
Supplier price listJavaScript shell; the table is filled from a JSON endpoint; pagination via a query parameternetwork ($.items) with a dom fallback on the table rowsListing / table — sku, name, price (number), currency, in_stock (boolean); diff keyed on sku
Notice boardServer-rendered list with a "next" link; each notice has a date; one notice body is an image of textdom list + ocr on the scanned image + readabilityArticle / notice — items[] of ref, title, published (date), body
PDF circularA PDF with one text-layer page and one scanned pagepdf (text layer first, the vision model reads the scanned page)Document / PDF — title, reference, issued (date), summary, body
Public tender boardA JavaScript-rendered notice table on a government or international procurement portaldom over the rendered rows (or network when the board exposes JSON)Listing / table — tender_id, title, buyer, category, closing_date (date), link; diff keyed on tender_id

The first three are the fixture sites the platform's own test suite crawls; the fourth is the tender tenant's showcase source. Same wizard, same preview, same provenance, same diff.