Eight projects. Three control what an agent can touch. Three measure what it does. Two put agents to work on invoices and on daily check-ins for older adults.
Scoped, expiring, auditable credentials for agents.
A user grants an agent a short-lived token with explicit scopes and spend limits. Tool servers ask agent-auth whether a request fits. A sub-agent only ever gets a narrower token, revoking a token revokes everything derived from it, and every decision lands in a hash-chained audit log.
Disposable sandboxes with snapshot, rollback and fork.
An agent gets a Docker sandbox with no network and a read-only root by default. It snapshots before a risky step, rolls back in one call, or forks a snapshot into parallel attempts. It ships a REST API, a Python SDK and an MCP server, and your code never leaves your hosts.
Lint and grade MCP servers before an agent calls them.
mcp-lint connects to a Model Context Protocol server over stdio or HTTP and checks every tool, prompt and resource for broken schemas, vague descriptions, prompt-injection surfaces and dangerous capabilities. It scores the server from 0 to 100, writes JSON and SARIF, and fails CI below your bar.
$ mcp-lint --file examples/messy-server.jsonmcp-lint 0.2.0 fixture-messy 1.0.0 (6 tools, 1 prompt, 1 resource)tool fetchUrlerror injection/instruction-phrases The description tells the model to hide something from the user: "Do not tell the user".warning injection/hidden-markup The description contains an instruction-style tag: "<IMPORTANT>".tool get_weathererror injection/hidden-unicode The description contains invisible characters (U+200B) that a reviewer cannot see but a model reads....Score 50/100 Grade F (6 errors, 10 warnings, 3 info)Failed: score is below the minimum of 70.
Evals as code that fail the build when quality drops.
You describe good output in YAML. evalkit sends each task to a model or to your own app, grades the answers, and prints a table, JSON, JUnit, a regression diff against a saved baseline and a badge. It runs the same on your laptop and in GitHub Actions.
Latency, barge-in and false barge-ins for voice agents.
voicebench plays a scripted conversation into your agent, records both sides as audio, and derives every metric from that audio with a voice activity detector. It talks to agents over a plain PCM WebSocket, a LiveKit room, a Pipecat bot or your own adapter.
Seeded, reproducible success rates for robot policies.
You register a manipulation policy, as a Python callable or a model behind HTTP or WebSocket, and robo-evals runs it through seeded MuJoCo scenes. You get per-task success rates with Wilson confidence intervals, JSON and Markdown reports, and episode videos. The same command gives the same numbers on any machine with the same MuJoCo build.
invoice-agent pulls structured data out of an invoice PDF, checks the arithmetic, aligns every line with a purchase order and goods receipt, and catches duplicates and over-billing. Each decision carries readable reasons and stable reason codes. It runs offline and ships a labeled set of 50 synthetic invoices to score itself on.
$ invoice-agent process dataset/invoices/016_b04.pdf dataset/invoices/005_q05.pdf \ --pos dataset/purchase_orders.json --receipts dataset/receipts.json016_b04.pdf: NEEDS REVIEW vendor Brightforge Maschinenteile GmbH total EUR 409.05 po PO-2026-0114 (3-way) lines 2/2 matched to PO lines[!] PRICE_VARIANCE: Line 1 ('Zahnriemen HTD 8M / Timing belt HTD 8M', PO line 1): unit price 61.56 is 8.0% above the PO price 57.00 (tolerance 2.0% or 0.05).005_q05.pdf: NEEDS REVIEW vendor Quillfeather Office Supply Co. total USD 256.99[!] PO_MISSING: The invoice shows no PO number. PO PO-2026-0105 fits the invoice lines (score 1.00). Review the inferred PO before approving.
A daily voice check-in for older adults who live alone.
care-voice asks a short, warm check-in every day: sleep, medication, food, pain, falls, the day of the week, mood. It turns the replies into structured answers, compares them with recent days, and sends caregivers plain-language alerts. It is not a medical device and not for emergencies.
$ care-voice simulate --name Margaret --replies examples/replies/concerning-day.txt --date 2026-09-30...agent: Have you had a fall or a stumble since we last spoke? you: I slipped in the bathroom last nightagent: Just so I have it right, can you tell me what day of the week it is today? you: Is it Sunday?...alerts: [HIGH] FALL_REPORTED: Margaret reported a fall. [HIGH] MISSED_MEDS: Margaret has not taken their morning medication. [HIGH] PAIN_REPORTED: Margaret reported severe pain. [MEDIUM] POSSIBLE_CONFUSION: Margaret showed 1 sign(s) of possible confusion. [LOW] LOW_MOOD: Margaret reported low mood.
Studio
cutroom makes short motion videos for launches and social.
Kinetic type and motion graphics, built in code frame by frame. English and Arabic. From $300 per video.