Backend integration
Masar: any website as a structured source
Turn any website into a structured, queryable source for your agents without an API: define a source (URLs, extractors, your JSON Schema, schedule), discover and test it, refresh into versioned snapshots with per-field provenance and diffs, and grant the generated masar:<slug>.query / latest / refresh tools — refresh always needs a human approval.
Masar turns any website into a structured, queryable source for your agents — no API, no database access, no export from the other side. You point it at a page, say what the information looks like (a JSON Schema you own) and how to read it (CSS selectors over the rendered DOM, the JSON the page itself loads, the text layer of a PDF, a vision model over a screenshot, or that same model as OCR for scanned images), and Agentic Workforce ME keeps a versioned snapshot with per-field provenance and a diff between refreshes. Agents consume it as three generated tools — masar:<slug>.query, .latest and .refresh — under the same grants and approval policies as every other tool. A public tender board, a supplier price list, a ministry notice board, a regulator's PDF circular and an exchange-rate widget are all just sources; nothing in the engine knows which is which.
1. What a source is
A source is a declarative definition: entry URLs and a crawl scope, an ordered list of extractors (deterministic ones first), the output schema, the key fields a diff aligns rows on, a refresh schedule and politeness limits. Every field name below belongs to you — the example is a supplier price list whose table is filled from an XHR endpoint, with a DOM extractor as the fallback.
{
"slug": "supplier-prices",
"name": "Acme supplier price list",
"entry_urls": ["https://shop.example.com/catalogue"],
"scope": { "allowed_hosts": ["shop.example.com"], "follow_links": "none", "max_pages": 1 },
"renderer": "auto",
"extractors": [
{ "kind": "network", "match": { "url_includes": "/api/products" }, "pick": "$.items",
"target": "items", "map": { "sku": "code", "name": "title", "price": "price.amount",
"currency": "price.currency", "in_stock": "available" } },
{ "kind": "dom", "fields": {}, "list": { "target": "items", "item": "table tbody tr",
"fields": { "sku": { "selector": "td:nth-child(1)" }, "name": { "selector": "td:nth-child(2)" },
"price": { "selector": "td:nth-child(3)" } } } }
],
"output_schema": {
"type": "object",
"properties": {
"items": { "type": "array", "items": { "type": "object",
"properties": { "sku": { "type": "string" }, "name": { "type": "string" },
"price": { "type": "number" }, "currency": { "type": "string" },
"in_stock": { "type": "boolean" } },
"required": ["sku", "name", "price"] } }
},
"required": ["items"]
},
"key_fields": ["items.sku"],
"schedule": "0 6 * * 1-5", "timezone": "Asia/Dubai",
"politeness": { "min_delay_ms": 2000, "respect_robots": true },
"limits": { "timeout_ms": 30000, "vlm_max_pages": 0, "ocr_max_images": 0, "max_age_seconds": 172800 },
"contains_pii": false,
"terms_attested": true
}| Extractor | Reads | Needs the browser renderer |
|---|---|---|
dom | CSS selectors over the (rendered) HTML — scalar fields and a list (`item` rows inside `target`); `attr` picks text, html, href, src or any attribute. | no (fetch works for static pages; browser for JS-rendered ones) |
network | XHR / fetch JSON responses the page itself received — `match` the URL, `pick` a JSON path, `map` schema fields to item keys. Bodies are captured, never replayed. | yes |
readability | The main article text (boilerplate removed) into one field, optionally the title into another. | no |
pdf | Linked or entry PDFs: text layer first; pages without one are rasterised and read by the vision model. | no |
vlm | A vision model over the page screenshot, restricted to the schema fields still missing (`mode: fallback`) or to listed fields (`always`); every value comes back with a confidence and a quote. | yes |
ocr | Transcription of images matching a selector / URL filter into a text field (the vision model as OCR; `engine: vlm`). | no (images are fetched directly) |
renderer: auto fetches first and switches to the browser when the page is a JavaScript shell or an extractor needs network capture or a screenshot. Vision and OCR are opt-in per source, capped per run (vlm_max_pages, ocr_max_images, default 5) and metered. Starter templates (GET /v1/masar/templates) are deliberately domain-neutral — listing / table, article / notice, document / PDF, dashboard / KPIs — and discovery renames their fields from the page's own headers and keys.
2. Create it: discover, test, save
The console wizard (Build → Integrations → Masar → "Add a source": URL → discover → extractors → schema → test extraction → schedule and attestation) is a thin client over three calls you can also script.
Discover what the page really loads
curl -sS -X POST "https://<api-host>/v1/masar/discover" \
-H "Authorization: Bearer hive_…" -H "X-Tenant-Id: <tenant>" \
-H "Content-Type: application/json" \
-d '{ "url": "https://shop.example.com/catalogue" }'
# → { tables[], lists[], json_endpoints[{url, pick, item_keys}], images[], pdfs[],
# suggested_extractors[], suggested_schema, suggested_template, warnings[] }Test a draft without persisting anything
curl -sS -X POST "https://<api-host>/v1/masar/test-extract" \
-H "Authorization: Bearer hive_…" -H "X-Tenant-Id: <tenant>" \
-H "Content-Type: application/json" \
-d '{ "definition": { …entry_urls, scope, renderer, extractors, output_schema, key_fields… },
"max_pages": 1 }'
# → { data, provenance, status, record_count, coverage: {found, expected}, pages[], warnings[],
# screenshot_keys[], vlm_calls, tokens_in, tokens_out } (nothing persisted)Each value in provenance names its extractor, URL and fetch time plus the selector path, the captured request and JSON path, the PDF page or the screenshot region it came from, and a confidence (1 for deterministic extractors). A field the page does not have is null with a reason (not_found, invalid, extractor_unavailable, cap_reached, low_confidence) — never a guess.
Save
POST /v1/masar/sources with the definition plus slug, name, schedule, timezone and terms_attested: true — the creator confirms the site's content may be used this way, and the attestation (who, when) is stored on the source. The schema is compiled before anything is persisted (400 MASAR_INVALID_SCHEMA); entry URLs that fail the SSRF policy, the scope or robots.txt are refused at definition time (400 MASAR_URL_REFUSED); a cron that fires more often than every 15 minutes is a validation error.
3. Refresh, runs and snapshots
curl -sS -X POST "https://<api-host>/v1/masar/sources/<sourceId>/refresh" \
-H "Authorization: Bearer hive_…" -H "X-Tenant-Id: <tenant>" \
-H "Content-Type: application/json" -d '{ "reason": "weekly price check" }'
# → 202 { run: { id, status: "queued", trigger: "manual", … } }
# → 409 MASAR_RUN_IN_PROGRESS while a run of this source is still runningA refresh is a worker job: it crawls the pages in scope (one at a time, with the politeness delay, robots.txt checked per host), runs the extractors, validates and coerces the result against your schema, and — only when the content hash changed — writes snapshot version + 1 with the diff against the previous version (added / removed / changed rows keyed on key_fields, changed scalar fields). Scheduled sources refresh on their cron; an agent may ask for one through masar:<slug>.refresh, which always pauses for a human decision.
| Route | Role | Purpose |
|---|---|---|
POST /v1/masar/discover | admin | Probe one URL; report + suggested extractors and schema draft. |
POST /v1/masar/test-extract | admin | Dry run of a draft definition (up to 3 pages); nothing persisted but preview screenshots. |
GET /v1/masar/sources | member | List (counts, last run, latest version, grants). |
POST /v1/masar/sources | admin | Create; schema compiled first; terms attestation required. |
GET · PATCH · DELETE /v1/masar/sources/:id | member · admin · admin | Detail with grants and latest snapshot summary; update (scheduler re-synced); delete (runs, snapshots and grants cascade). |
POST /v1/masar/sources/:id/test | admin | Test extraction on the saved definition. |
POST /v1/masar/sources/:id/refresh | admin | 202 with the queued run; 409 MASAR_RUN_IN_PROGRESS. |
GET /v1/masar/sources/:id/runs · GET /v1/masar/runs/:id | member | Runs newest first; run detail (pages crawled / blocked, requests captured, vision calls, tokens, warnings). |
GET /v1/masar/sources/:id/snapshots · GET /v1/masar/snapshots/:id | member | Snapshot list with diff summaries; one snapshot with data, provenance and diff. |
GET /v1/masar/sources/:id/snapshots/:version/diff?against= | member | Diff between any two versions. |
GET /v1/masar/artifacts/* | member | Screenshot bytes; keys are tenant-prefixed and checked. |
GET /v1/masar/templates | member | The domain-neutral starter templates. |
4. Grant it to agents
Declare the source in the agent manifest and grant it — both default-deny, like every connector. allow defaults to query and latest; list refresh only for agents that may ask for fresher data.
{
"tools": {
"masar": [
{ "source": "supplier-prices", "allow": ["query", "latest"] },
{ "source": "regulator-circulars", "allow": ["query", "latest", "refresh"] }
]
}
}The composite POST /v1/integrations/attach step accepts connector.kind = "masar" and creates the tool grants (source: masar, masar_source_id), the approval policies and the draft-manifest merge in one transaction (Connectors, tools & grants).
curl -sS -X POST "https://<api-host>/v1/integrations/attach" \
-H "Authorization: Bearer hive_…" -H "X-Tenant-Id: <tenant>" \
-H "Content-Type: application/json" \
-d '{
"connector": { "kind": "masar", "id": "<sourceId>" },
"tools": ["query", "latest", "refresh"],
"targets": [{ "agent_id": "<agentId>" }],
"approvals": [{ "tool": "refresh", "approver_role": "admin" }],
"publish": false
}'
# refresh is force-gated in the runtime anyway; the policy row makes the card explicit| Tool | What it does | Approval |
|---|---|---|
masar:<slug>.query | Filters the latest snapshot: `filters[{field, op: eq | neq | contains | gte | lte | in, value}]` over the schema’s scalar fields (up to 10), `fields` projection, `list` when the schema has several lists, `limit` ≤ 100, `offset`. Never crawls. | the agent’s normal policy |
masar:<slug>.latest | The latest snapshot’s version, fetch time, age, `stale` flag and diff summary — "what changed since last time". Never crawls. | the agent’s normal policy |
masar:<slug>.refresh | Enqueues a crawl and answers `{ queued: true, run_id }` — the next `query` reads the new snapshot. | always required (forced, not relaxable) |
{
"source": "supplier-prices",
"snapshot": { "id": "<id>", "version": 12, "fetched_at": "2026-10-01T06:00:04Z",
"age_seconds": 5400, "stale": false },
"records": [
{ "cite": "masar:supplier-prices#12/0", "sku": "A-100", "name": "Hex bolt M8",
"price": 1.25, "currency": "AED", "in_stock": true }
],
"total": 1, "truncated": false, "provenance_available": true
}Tool descriptions are generated from the source — its name, description and the schema's field names and descriptions, plus the latest version and fetch time — so the model knows the vocabulary without any hard-coded domain. Every record carries a cite (masar:<slug>#<version>/<row>) the agent quotes; a snapshot older than the source's max_age_seconds is returned with stale: true and the agent is told to ask for a refresh (which needs a person) rather than guess. Field-level provenance stays in the console's snapshot browser; snapshot values reach the model as data, scanned by the same tool-output guardrail as every other tool result.
5. Posture, limits and errors
- robots.txt is honoured and cannot be switched off; a disallowed URL is recorded as blocked (and audited), the run continues with what is allowed.
Crawl-delayis respected up to 30 s. - The crawler identifies itself (
HiveMasar/1.0 (+https://agenticworkforce.me/masar)), fetches one page at a time per source with at least one second between pages (default 1.5 s), caps pages (default 20, hard cap 200), depth (3) and run time (10 min), and never solves a bot challenge, captcha or paywall — such pages are recorded withblocked_reason: challenge. - Every URL — entry, followed link, image, PDF — passes the platform SSRF policy and the tenant's data-residency guard; captured XHR bodies are read, never re-requested.
- Vision / OCR calls are opt-in, capped and metered (
masar_vlm_call, tokens); a run that hits the cap stops calling the model and says so. A snapshot is capped at 2 MB / 2 000 records (422 MASAR_BUDGET_EXCEEDED, truncated and flagged). - A source may be flagged
contains_pii; the wizard warns on PII-looking field names and the snapshot browser redacts those columns for non-admins. Snapshots follow the tenant retention purge.
| Code | Status | When |
|---|---|---|
MASAR_SOURCE_NOT_FOUND | 404 | Unknown source in this tenant (another tenant’s source is a 404, never a 403). |
MASAR_INVALID_SCHEMA | 400 | The output schema does not compile or has no object root. |
MASAR_URL_REFUSED | 400 | An entry URL fails the SSRF policy, the scope or robots.txt at definition time. |
MASAR_RUN_IN_PROGRESS | 409 | A refresh was requested while a run of the source is running. |
MASAR_RENDERER_UNAVAILABLE | 503 | An extractor needs the browser renderer and the browser sandbox is not configured. |
MASAR_VLM_UNAVAILABLE | 503 | A vision extractor is configured but no vision-capable model is. |
MASAR_BUDGET_EXCEEDED | 422 | The snapshot exceeded its size or record cap. |
Evidence: audit actions masar.source.create | update | delete, masar.discover, masar.test_extract, masar.refresh.requested, masar.run.completed | failed and masar.navigation.blocked (ids, counts, public URLs and reasons — never page content); usage kinds masar_page and masar_vlm_call plus the model tokens of vision calls.
Three sources, one engine
| Source | Shape | Extractors | Template |
|---|---|---|---|
| Supplier price list | JavaScript shell; the table is filled from a JSON endpoint; pagination via a query parameter | network ($.items) with a dom fallback on the table rows | Listing / table — sku, name, price (number), currency, in_stock (boolean); diff keyed on sku |
| Notice board | Server-rendered list with a "next" link; each notice has a date; one notice body is an image of text | dom list + ocr on the scanned image + readability | Article / notice — items[] of ref, title, published (date), body |
| PDF circular | A PDF with one text-layer page and one scanned page | pdf (text layer first, the vision model reads the scanned page) | Document / PDF — title, reference, issued (date), summary, body |
| Public tender board | A JavaScript-rendered notice table on a government or international procurement portal | dom over the rendered rows (or network when the board exposes JSON) | Listing / table — tender_id, title, buyer, category, closing_date (date), link; diff keyed on tender_id |
The first three are the fixture sites the platform's own test suite crawls; the fourth is the tender tenant's showcase source. Same wizard, same preview, same provenance, same diff.