Open source, Apache-2.0

Tools for agents that do real work.

Eight projects. Three control what an agent can touch. Three measure what it does. Two put agents to work on invoices and on daily check-ins for older adults.

Control

Control what an agent can touch.

agent-auth TypeScript

Scoped, expiring, auditable credentials for agents.

A user grants an agent a short-lived token with explicit scopes and spend limits. Tool servers ask agent-auth whether a request fits. A sub-agent only ever gets a narrower token, revoking a token revokes everything derived from it, and every decision lands in a hash-chained audit log.

Install, npm, from the GitHub release
npm install https://github.com/superintelligenceco/agent-auth/releases/download/v0.2.0/agent-auth-0.2.0.tgz
$ agent-auth check --token "$TOKEN" --action gmail:send --param to=bob@acme.comALLOW             [gmail:send to:*@acme.com] (19 uses left)$ agent-auth check --token "$TOKEN" --action gmail:send --param to=eve@evil.exampleDENY              to eve@evil.example does not match *@acme.com (scope: gmail:send to:*@acme.com)$ agent-auth check --token "$TOKEN" --action payments:charge --amount 80USDDENY              amount exceeds max=50USD (scope: payments:charge max=50USD)$ agent-auth attenuate --token "$SUB" --scope 'github:repo:read acme/*'error: scope_escalation: requested scopes exceed the parent token  not covered by parent: github:repo:read acme/*

Disposable sandboxes with snapshot, rollback and fork.

An agent gets a Docker sandbox with no network and a read-only root by default. It snapshots before a risky step, rolls back in one call, or forks a snapshot into parallel attempts. It ships a REST API, a Python SDK and an MCP server, and your code never leaves your hosts.

Install, PyPI
pip install sic-agent-sandbox
$ python examples/quickstart.pycreated Sandbox(id='sb_2bc35260898babc9', image='python:3.12-slim', status='running')run 1: hello from 3.12.13snapshot snap_bdb4731b7ddb7b93 (10240 bytes)after rm: exit 2 | python3: can't open file '/workspace/app.py': [Errno 2] No such file or directoryafter rollback: hello from 3.12.13stream: ExecEvent(type='stdout', data='tick 1\n', exit_code=None, timed_out=False, duration_ms=None)stream: ExecEvent(type='exit', data='', exit_code=0, timed_out=False, duration_ms=641)
mcp-lint TypeScript

Lint and grade MCP servers before an agent calls them.

mcp-lint connects to a Model Context Protocol server over stdio or HTTP and checks every tool, prompt and resource for broken schemas, vague descriptions, prompt-injection surfaces and dangerous capabilities. It scores the server from 0 to 100, writes JSON and SARIF, and fails CI below your bar.

Install, npm, from the GitHub release
npm install -g https://github.com/superintelligenceco/mcp-lint/releases/latest/download/mcp-lint.tgz
$ mcp-lint --file examples/messy-server.jsonmcp-lint 0.2.0  fixture-messy 1.0.0  (6 tools, 1 prompt, 1 resource)tool fetchUrl  error    injection/instruction-phrases   The description tells the model to hide something from the user: "Do not tell the user".  warning  injection/hidden-markup         The description contains an instruction-style tag: "<IMPORTANT>".tool get_weather  error    injection/hidden-unicode        The description contains invisible characters (U+200B) that a reviewer cannot see but a model reads....Score 50/100  Grade F  (6 errors, 10 warnings, 3 info)Failed: score is below the minimum of 70.
Measure

Measure what it actually does.

evalkit Python

Evals as code that fail the build when quality drops.

You describe good output in YAML. evalkit sends each task to a model or to your own app, grades the answers, and prints a table, JSON, JUnit, a regression diff against a saved baseline and a badge. It runs the same on your laptop and in GitHub Actions.

Install, PyPI
pip install sic-evalkit
$ evalkit run examples/support-bot/evals.yamlevalkit 0.2.0  suite support-bot  target command:python3  8 tasks x 1 run  STATUS  TASK                 SCORE   RUNS  LATENCY      COST  DETAIL  pass    password-reset        1.00    1/1     40ms         -  pass    order-status          1.00    1/1     30ms         -  pass    unknown-order         1.00    1/1     20ms         -  pass    refund                1.00    1/1     41ms         -  ...  FAIL    cancel-subscription   0.67    0/1     19ms         -  contains: missing 'Billing > Subscription'7/8 passed (87.5%), score 0.96, 0 flaky, 0 errored, latency p50 23ms p95 41ms, 0/8 runs cached, 0.11sPASS: pass rate 87.5% meets fail_under 85.0%

Latency, barge-in and false barge-ins for voice agents.

voicebench plays a scripted conversation into your agent, records both sides as audio, and derives every metric from that audio with a voice activity detector. It talks to agents over a plain PCM WebSocket, a LiveKit room, a Pipecat bot or your own adapter.

Install, PyPI
pip install voicebench
$ voicebench run examples/mock-conversation.yaml --no-writevoicebench 0.2.0  scenario=mock-conversation  adapter=mock  clock=virtual  sessions=1Turn         Expect   Response  TTFA    Stop    WER  Resultopening      respond  760 ms    760 ms  -       0    passanswer-date  respond  700 ms    700 ms  -       0    passbarge-in     stop     660 ms    660 ms  380 ms  0    passcough        ignore   -         -       -       -    passconfirm      respond  660 ms    660 ms  -       0    passResponse latency p50 / p90     700 ms / 748 msBarge-in stop time p50 / max   380 ms / 380 msPASS  p90_response_latency_ms <= 1200 (actual 748)PASS  max_barge_in_stop_ms <= 400 (actual 380)

Seeded, reproducible success rates for robot policies.

You register a manipulation policy, as a Python callable or a model behind HTTP or WebSocket, and robo-evals runs it through seeded MuJoCo scenes. You get per-task success rates with Wilson confidence intervals, JSON and Markdown reports, and episode videos. The same command gives the same numbers on any machine with the same MuJoCo build.

Install, pip, from GitHub
pip install "robo-evals @ git+https://github.com/superintelligenceco/robo-evals"
$ robo-evals run --policy scripted --suite core --episodes 20...  stack       ep  19  ok    steps=30Wrote results/scripted/report.json and results/scripted/report.md$ robo-evals compare results/random/report.json results/scripted/report.json| Task | `random` | `scripted` || --- | --- | --- || reach | 0.0% [0.0, 16.1] (0/20) | 100.0% [83.9, 100.0] (20/20) || push | 0.0% [0.0, 16.1] (0/20) | 100.0% [83.9, 100.0] (20/20) || pick_place | 0.0% [0.0, 16.1] (0/20) | 100.0% [83.9, 100.0] (20/20) || drawer | 0.0% [0.0, 16.1] (0/20) | 100.0% [83.9, 100.0] (20/20) || stack | 0.0% [0.0, 16.1] (0/20) | 100.0% [83.9, 100.0] (20/20) || **overall** | **0.0% [0.0, 3.7] (0/100)** | **100.0% [96.3, 100.0] (100/100)** |
Apply

Put it to work.

Invoice PDFs in, approve, review or reject out.

invoice-agent pulls structured data out of an invoice PDF, checks the arithmetic, aligns every line with a purchase order and goods receipt, and catches duplicates and over-billing. Each decision carries readable reasons and stable reason codes. It runs offline and ships a labeled set of 50 synthetic invoices to score itself on.

Install, pip, from the GitHub release
pip install https://github.com/superintelligenceco/invoice-agent/releases/download/v0.2.1/invoice_agent-0.2.1-py3-none-any.whl
$ invoice-agent process dataset/invoices/016_b04.pdf dataset/invoices/005_q05.pdf \    --pos dataset/purchase_orders.json --receipts dataset/receipts.json016_b04.pdf: NEEDS REVIEW  vendor  Brightforge Maschinenteile GmbH  total   EUR 409.05  po      PO-2026-0114 (3-way)  lines   2/2 matched to PO lines  [!] PRICE_VARIANCE: Line 1 ('Zahnriemen HTD 8M / Timing belt HTD 8M', PO line 1): unit price 61.56 is 8.0% above the PO price 57.00 (tolerance 2.0% or 0.05).005_q05.pdf: NEEDS REVIEW  vendor  Quillfeather Office Supply Co.  total   USD 256.99  [!] PO_MISSING: The invoice shows no PO number. PO PO-2026-0105 fits the invoice lines (score 1.00). Review the inferred PO before approving.

A daily voice check-in for older adults who live alone.

care-voice asks a short, warm check-in every day: sleep, medication, food, pain, falls, the day of the week, mood. It turns the replies into structured answers, compares them with recent days, and sends caregivers plain-language alerts. It is not a medical device and not for emergencies.

Install, PyPI
pip install care-voice
$ care-voice simulate --name Margaret --replies examples/replies/concerning-day.txt --date 2026-09-30...agent: Have you had a fall or a stumble since we last spoke?  you: I slipped in the bathroom last nightagent: Just so I have it right, can you tell me what day of the week it is today?  you: Is it Sunday?...alerts:  [HIGH] FALL_REPORTED: Margaret reported a fall.  [HIGH] MISSED_MEDS: Margaret has not taken their morning medication.  [HIGH] PAIN_REPORTED: Margaret reported severe pain.  [MEDIUM] POSSIBLE_CONFUSION: Margaret showed 1 sign(s) of possible confusion.  [LOW] LOW_MOOD: Margaret reported low mood.

Studio

cutroom makes short motion videos for launches and social.

Kinetic type and motion graphics, built in code frame by frame. English and Arabic. From $300 per video.

Pricing and details