Skip to content

API & MCP reference

A governance harness with an API.

A default-deny harness around any model: propose → approve → dispatch, a write-only secrets vault, everything on the record. Three ways in — REST, a WebSocket and MCP — over one runtime, so the scoping and the gate cannot differ between them.

REST

/v1

Chat, memory, knowledge, proposals, goals, secrets, personas, delegation, projects, documents.

WebSocket

/v1/ws

The chain of thought as it happens — and steerable and stoppable mid-run.

MCP

/mcp · /mcp/:key

The brain itself on a workspace key, or one tool library on a scoped, revocable token.

Base URL https://monopea-runtime.fly.dev

Quickstart

Auth, then one call

Everything is a bearer token and a JSON body. This is a complete, working turn: the agent retrieves, plans, calls tools, and queues anything outward as a proposal.

POST /v1/chatcurl
curl https://monopea-runtime.fly.dev/v1/chat \
  -H "Authorization: Bearer brain_live_..." \
  -H "Content-Type: application/json" \
  -d '{"message":"Reconcile this week'\''s invoices and propose the transfers"}'

# -> {"ok":true,"result":{"conversationId":"...","answer":"...",
#     "stopReason":"blocked_on_user","steps":7,"proposals":[...],
#     "citations":[...],"toolsExecuted":[...]}}

The response carries the answer and the run's stopReason — note the camelCase; the WebSocket's done frame spells the same value stop_reason. If the agent wanted to do something outward, that reason is blocked_on_user and the queued calls are in result.proposals — see the approval gate. A tenant out of credit gets a 402 with code: insufficient_credits rather than a partial run.

Authentication

Two key kinds, two scopes

Pass your key as Authorization: Bearer …. Which key you use decides what the call can see — the boundary is enforced in Postgres with row-level security, not in application code that could forget.

brain_live_…

Tenant key

Full tenant scope: chat, memory, knowledge, proposals, goals, secrets, personas, delegation, projects and documents. The hard security boundary — a tenant never sees another tenant’s anything.

customer_live_…

Sub-user key

A private scope within your tenant for a downstream customer: their conversations, memory, documents and personas are invisible to other sub-users, and their usage is metered separately. Answers 403 on the whole secrets vault and on persona authoring.

REST

The /v1 surface

Everything the agent knows and does, over plain HTTP. The same routes back the dashboard, the WebSocket and the MCP server, so tenant scoping and the gate live in exactly one place. Rows marked gated can produce a proposal instead of an immediate effect.

Run the agent

  • POST/v1/chatOne full agent turn: retrieval, tool loop, proposals. Body: {message, conversation_id?, agent_key?, project_id?, all_projects?, incognito?, document_ids?, document_id?, toolsets?, delegate?, model?, temperature?, web_search?, force_tool?, phi?, region?} — every option the product has, read by one parser for all four transports. Send Accept: text/event-stream to get the same turn back as an NDJSON stream of frames instead. A second call on a conversation whose turn is still running is folded INTO that turn and answers {steered:true} with no result; watch /live for the answer.gated
  • GET/v1/conversationsEpisodic memory — the conversation list. ?limit= (max 200), ?project_id=.
  • GET/v1/conversations/:id/messagesThe messages of one conversation, each assistant turn carrying what it left behind: citations (where its claims came from), artifacts (the files it WROTE — never the ones it only read), trace and ui_elements. ?limit= (max 500).
  • GET/v1/toolsetsThe tool catalogs a turn can be given — built-in (word, excel, powerpoint, pdf, calendar, mail, meetings, browser, media, …) plus this workspace's connected products as mcp:<key>. These are the strings POST /v1/chat accepts in `toolsets`. weight:"heavy" means 20–25k tokens in the prompt.
  • GET/v1/jobs/:id/liveOne subagent, as something to watch: {state, live, step, partial, tools[], result}. GET /v1/jobs/:id returns the whole row and every related event, which is the right answer once and the wrong one thirty times a minute. live:false ends the wait.
  • GET/v1/conversations/:id/liveWhat the turn on this thread is doing right now: {state, step, thinking[], status[], answer, answers}. A turn outlives the connection that started it, so this is how a caller whose stream died picks it back up. state is running | stalled (its process died; it is re-run within a couple of minutes — keep waiting) | idle. Poll every ~2s; `answers` rising means the answer has landed.
  • POST/v1/conversations/:id/steerSend {message} INTO the turn already running on this thread. It arrives as a user message at the turn's next step boundary, so a correction lands inside the work rather than behind it. {steered:false} when nothing is running there.gated
  • POST/v1/conversations/:id/stopStop the running turn. Cooperative and durable: it halts at its next step boundary and KEEPS what it has done — the partial answer, the files it wrote, the trace — all readable from /messages.gated
  • DELETE/v1/conversations/:idDelete a thread and its messages. This one is a hard delete.

What it knows

  • GET/v1/memoryTyped durable memory. Filter with ?kind=, ?limit= (max 500), ?project_id=.
  • POST/v1/memoryRemember something: {kind, summary, body}. kind is one of user_profile, feedback_rule, project_fact, external_reference — anything else is a 400.
  • POST/v1/memory/archiveSoft-archive one memory by {id}. There is no hard delete on this route.
  • POST/v1/knowledge/searchHybrid semantic search {query, top_k?} → {entities, documents, claims}. top_k defaults to 6 and caps at 24.
  • POST/v1/knowledge/traverseGraph walk {entity_ids?|query?, max_hops?, max_facts?} → {facts}, each graded. max_hops caps at 3.
  • GET/v1/evidence/claimsGraded claims from the literature layer.
  • POST/v1/evidence/searchSemantic search over those claims: {query, top_k?}.

Govern it

The approval queue and the vault — the two surfaces that make the API a harness rather than a proxy.

  • GET/v1/proposals?status=The approval queue. Defaults to pending; status=all lifts the filter.
  • POST/v1/proposals/:id/decideDecide one: {decision: "approve"|"reject", note?}. Approving dispatches inline; a proposal that is no longer pending answers 409.gated
  • GET/v1/goalsLong-running goals the agent plans against. ?status=, ?project_id=.
  • POST/v1/goalsSet a new goal: {title, description?, domain?, priority?}.gated
  • GET/v1/secretsVault index — name, description and last-4 only. A value can never be read back.
  • PUT/v1/secrets/:nameWrite a value: {value, description?}.
  • DELETE/v1/secrets/:nameRemove a vault entry.

Hand work to subagents

The orchestration surface: spawn background agents, watch the tree they make, and stop it. The gate applies inside a delegated run exactly as it does inside yours.

  • POST/v1/delegateSpawn background subagents: {tasks: [{task, title?, persona?, model?, scope?, success_criteria?}], goal_id?, conversation_id?, project_id?}. Up to 8 tasks per call.gated
  • GET/v1/jobsBackground jobs for the tenant. ?project_id=.
  • GET/v1/jobs/:idOne job and its event stream.
  • GET/v1/jobs/:id/treeThe whole delegated run tree beneath a job.
  • POST/v1/jobs/:id/stopCooperative cancel — cascades to the subtree.
  • GET/v1/runsWhich agents are live right now; :id inspects one run’s subtree.

Shape it

  • GET/v1/personasThe caller’s persona set.
  • GET/v1/personas/:keyOne persona’s prompt sections plus the assembled prompt.
  • POST/v1/personasAuthor a tenant-owned persona: {name, role, slug?, display_name?, category?, visibility?} → {agent_key, visibility}.gated
  • POST/v1/personas/:key/promptsWrite one prompt section: {section, content_md, change_summary?}. section ∈ role, objective, tone, reasoning, tools, intents, rules, output_format, closing. Versions are immutable.gated
  • DELETE/v1/personas/:keyDelete a persona you own.gated

The workspace around it

  • GET/v1/projectsProjects in the tenant; POST creates one from {name}.
  • PATCH/v1/projects/:idRename or archive a project.
  • GET/v1/documentsThe library. ?project_id=; :id/content downloads the bytes.
  • POST/v1/documentsUpload a document as base64 bytes.gated
  • DELETE/v1/documents/:idDelete one document you own.
  • GET/v1/skillsSkills available to the agent. ?discover=1 widens the search.
  • GET/v1/routinesScheduled routines and when they next run.
  • GET/v1/eventsThe brain’s event feed.
  • GET/v1/security/scanAudit the MCP connection registry — configuration, not liveness.

A handful of older routes sit at the root rather than under /v1 and take the same tenant key: POST /etl (schema-driven extraction, with /etl/schemas and /etl/:id beside it), GET /usage and GET /billing. Anything under /internal is shared-secret territory and is not part of the public API.

The gate, as an API

The proposals lifecycle

Out of the box, nothing outward is dispatched inline. The agent drafts the exact call as a pending proposal and the run blocks — stopReason blocked_on_user — until someone with the authority decides. Approval dispatches it; rejection records why. Unknown tools fail closed to review, and a risk level the platform does not recognise is ranked as the most dangerous one rather than the least.

  1. 01 / pending

    An outward mutation queues as a proposal — tool name, arguments, target — and the run reports blocked_on_user. Nothing has executed.

  2. 02 / approve / reject

    POST /v1/proposals/:id/decide with {decision, note?} — from the dashboard, your own UI, or a script. Either decision is logged, and deciding the same proposal twice is a 409 rather than a second dispatch.

  3. 03 / dispatch

    Approval dispatches the call inline and the run resumes with the real result — and the decision, the dispatcher and the arguments that were actually sent are all on the record.

Widening the gate

Every tool the platform can call carries a risk level — NONE, LOW, MEDIUM, HIGH, CRITICAL — and a workspace sets one threshold against that scale. A call at or below the threshold dispatches inline; everything above it still becomes a proposal. The threshold starts at NONE, which is why nothing runs unattended until you say so, and a project can carry its own override. Moving it is itself a recorded event — autonomy_changed, carrying the level it moved from, the level it moved to, the scope and who moved it. It is a workspace setting rather than an API parameter — no /v1 route reads or writes it — so an integration key cannot widen the gate it is running behind.

Decide a proposalcurl
# the run came back stopReason: "blocked_on_user".
# list what is waiting, then decide.
curl "https://monopea-runtime.fly.dev/v1/proposals?status=pending" \
  -H "Authorization: Bearer brain_live_..."

curl -X POST https://monopea-runtime.fly.dev/v1/proposals/PROPOSAL_ID/decide \
  -H "Authorization: Bearer brain_live_..." \
  -H "Content-Type: application/json" \
  -d '{"decision":"approve"}'
# approve dispatches the call inline; or {"decision":"reject","note":"..."}
# a proposal that is no longer pending answers 409, never a silent re-dispatch
The Monopea chat, showing an answer whose figures each carry a citation to the board update it read, and beneath it a standing objective the agent has drafted as a proposal — its exact arguments in view, waiting on Reject or Approve.
Chats · the same proposal the API returns, waiting on a decision

Vault

Write-only secrets

Credentials go in and never come back out. GET returns the name, the description and the last four characters; there is no route that returns a value, so exfiltrating one would require a path that does not exist.

Write-only secretscurl
# write-only: PUT a value; GET returns name/description/last-4 only
curl -X PUT https://monopea-runtime.fly.dev/v1/secrets/HUBSPOT_KEY \
  -H "Authorization: Bearer brain_live_..." \
  -H "Content-Type: application/json" \
  -d '{"value":"pat-...","description":"HubSpot private app token"}'

# the agent references it as {{secret:HUBSPOT_KEY}} — plaintext is substituted
# only at dispatch and scrubbed from tool results; the model never sees it.
# all three vault routes answer 403 to a customer_live_ sub-user key.

The model sees only the reference {{secret:NAME}}. Substitution happens at dispatch, after the approval decision, and tool results are scrubbed of secret echoes before they re-enter context. Rotate a value and every workflow keeps working.

WebSocket

Watch the chain of thought, live

Connect to wss://monopea-runtime.fly.dev/v1/ws?key=… — the key is a query parameter because browsers cannot set headers on an upgrade — and stream every step of a turn as it happens: status, thinking, tokens, tool calls, tool results, proposals. Then steer the run mid-flight with a steer frame, folded in at the next step boundary, or stop it with cancel.

wss:// — live chain of thoughtjavascript
const ws = new WebSocket('wss://monopea-runtime.fly.dev/v1/ws?key=brain_live_...')

ws.onopen = () =>
  ws.send(JSON.stringify({ type: 'chat', message: 'Audit the inbound funnel' }))

// connection frames:  ready{tenant_id} -> accepted{conversation_id, request_id}
//                     -> ... -> result{request_id, ...} | error{error}
//
// turn frames, forwarded verbatim as the turn runs:
//   status{phase}         prepare | think | tool | answer
//   turn_start{tools_offered} / plan / step{thinking, tool_calls}
//   token / token_reset / reasoning        (the answer and the trace, streaming)
//   tool_call{...} / tool_result{...}
//   proposal{id, ...}     an outward mutation queued for approval
//   awaiting_choice{...}  the run is asking you to pick, not to approve
//   citation / artifact / ui_element / suggested_action / agent_mail
//   done{answer, stop_reason, steps}
//
// NOTE the case: the done frame carries snake_case stop_reason, while the REST
// response and the result frame carry the same value as stopReason.

// steer a run that is already in flight — folded in at the next step boundary.
// both fields are required; without conversation_id this is treated as a new turn.
ws.send(JSON.stringify({ type: 'steer', message: 'Skip the EU leads', conversation_id }))

// and stop one outright:
ws.send(JSON.stringify({ type: 'cancel', conversation_id }))

A socket opened without a key is closed with a bare HTTP/1.1 401 before the upgrade completes, and any path other than /v1/ws is dropped without a response. Client frames are chat, steer, cancel and ping; anything else comes back as an error frame naming what it did not recognise.

MCP

Two remote MCP endpoints

Claude Desktop, Cursor or any MCP client connects straight to monopea over Streamable HTTP — nothing of ours to install. Two URLs, and the difference between them is the credential rather than the protocol. Every tool library in a workspace is served at https://monopea-runtime.fly.dev/mcp/<service-key>, authenticated by a library token (mlib_…) you mint and revoke per client, so what leaves the workspace reaches that one library and nothing else. The brain itself is served at https://monopea-runtime.fly.dev/mcp with a workspace key — the agent, not one library’s tools.

MCP client config — one libraryclaude_desktop_config.json
{
  "mcpServers": {
    "monopea-library": {
      "command": "npx",
      "args": [
        "-y", "mcp-remote",
        "https://monopea-runtime.fly.dev/mcp/<service-key>",
        "--header", "Authorization: Bearer mlib_..."
      ]
    }
  }
}

Scopes

A token is read-only by default. Listing shows every enabled tool; calling one whose declared risk is not NONE requires the write scope. The in-workspace path routes those calls through the proposal gate, and an outside caller has no such gate — so the honest answer is to refuse rather than to silently mutate. A proxied code library, which publishes no per-tool risk, falls back to treating get_, list_, search_, read_, find_, fetch_, query_, check_ and describe_ as the reads.

Mint and revoke tokens from the library’s Connect panel in the Catalog. Issue one per customer and their calls are attributed to them. The endpoint takes mlib_ tokens and nothing else — a tenant key does not open it.

The Catalog, showing a grid of integrations in which only a handful read Ready and the rest await a key or a connection.
Catalog · where a library’s endpoint and its client tokens are minted

The brain’s own tools — 72

monopea-runtime.fly.dev/mcp
MCP client config — the brainclaude_desktop_config.json
{
  "mcpServers": {
    "monopea": {
      "command": "npx",
      "args": [
        "-y", "mcp-remote",
        "https://monopea-runtime.fly.dev/mcp",
        "--header", "Authorization: Bearer brain_live_..."
      ]
    }
  }
}

What the key reaches

A brain_live_… key authenticates a workspace, so this endpoint reaches everything /v1 reaches — which is the point, and the reason to think before pasting it into a client on a laptop. Rotate it if that laptop walks.

A key is not a person. A mailbox, a calendar and a meeting space belong to a seat, so they are deliberately absent here and from /v1 — membership is not consent to impersonate. A reseller’s downstream customer key opens the endpoint too, minus the four tools that would let it administer the reseller above it.

Start with ask_brain: it runs a full agent turn and is the right tool for almost anything phrased as a task rather than a lookup. A turn can run for minutes and ask_brain blocks until it ends, so when that is too long, delegate_work starts subagents and tail_agent reads one back.

These are the tools it registers:

Ask and remember

  • ask_brain
  • search_knowledge
  • traverse_knowledge
  • list_memory
  • remember
  • archive_memory

Govern

  • list_proposals
  • approve_proposal
  • reject_proposal
  • list_choices
  • answer_choice
  • list_goals
  • create_goal

Delegate

  • delegate_work
  • stop_agent
  • list_active_runs
  • list_jobs
  • get_run_tree
  • tail_agent

Watch a running turn

  • get_conversation_status
  • read_conversation
  • list_conversations
  • steer_brain
  • stop_brain
  • compact_conversation

What the workspace holds

  • list_projects
  • create_project
  • list_toolsets
  • list_models
  • list_secrets
  • list_events
  • list_routines

Personas and skills

  • list_personas
  • get_persona
  • create_persona
  • set_persona_prompt
  • list_skills
  • start_skill

Documents

  • list_documents
  • read_document
  • create_document
  • update_document
  • patch_document
  • list_office_tools

Dashboards

  • search_widgets
  • get_widget
  • list_widget_categories
  • list_dashboards
  • create_dashboard
  • update_dashboard

Calendar

  • list_calendars
  • list_calendar_events
  • find_free_time
  • check_calendar_conflicts
  • create_calendar_event
  • move_calendar_event
  • cancel_calendar_event

Mail

  • search_mail
  • read_mail
  • update_mail
  • create_mail_draft

Evidence

  • list_experiments
  • list_evidence

Schedule

  • get_live_schedule
  • get_digest

Metering and resale

  • get_usage
  • get_balance
  • get_pricing
  • list_customers
  • mint_customer_key
  • revoke_customer_key
  • create_topup_link

Writes (ask_brain, remember, approve_proposal, delegate_work, create_persona) proxy the runtime and its gate rather than re-implementing it; archive_memory is a soft archive, never a hard delete.

Isolation

Incognito turns & project scoping

Two isolation controls ride on the chat call itself: a no-retention flag for turns that must leave nothing behind, and a project lens that keeps one tenant's parallel workstreams from bleeding into each other.

incognito: true

The turn runs the full loop and still reads existing knowledge, but persists nothing durable — no conversation row, messages, memory, embeddings, summary or trace. It is still metered. Works on POST /v1/chat and on WebSocket chat frames.

project_id

One brain, many projects: a turn scoped to a project recalls and writes within that project only, with no cross-project awareness. Sub-agents inherit it. The tenant boundary stays the hard one; projects are a working lens inside it.

Stored in Switzerland. Processed in the EU.

Send the first turn in one call.

Get a key, POST a message, and watch the run stop at the gate with the drafted calls waiting for you.