Moyai catches AI agents that go wrong in production. It reads the traces a team already collects in Langfuse, LangSmith, or OpenTelemetry, finds the runs where an agent behaved abnormally, checks which of those are real failures, and tells the team what broke before a customer does. To surface the ways agents fail, and to measure how many of them its platform catches, Moyai runs coding agents against benchmark tasks with Harbor, at a large scale. Every task needs a clean computer of its own. Since August 2026, Harbor has created 279,577 boxd machines for Moyai, each one a full Linux VM that booted in a median of 5 milliseconds.
Before: Modal was expensive and hard to debug
Moyai runs hundreds of benchmark tasks, many times each, and every run needs a fresh environment. Before boxd, Moyai ran them in sandboxes on Modal. Three things got in the way.
- Cost. At this volume the sandbox bill became a real line item. On one day Moyai spent 1,000 USD on sandboxes alone.
- Sizing. Each sandbox needed a CPU and memory size chosen up front, and it was never clear what a given task needed. Too small and the task failed for reasons that had nothing to do with the agent. Too large and every one of thousands of runs cost more than it should.
- No way in. When a run behaved strangely, nobody at Moyai could go into the machine and look. They had the logs the sandbox gave them and nothing else.
After: one machine per environment, created by Harbor
Harbor evaluates agents such as Claude Code, Codex, and OpenHands against benchmarks, across many sandboxes at once. We wrote its boxd environment provider together with Robert Hommes, Moyai's founder.
- Every environment is its own boxd machine: a KVM virtual machine with a full Docker daemon and a 100 GB disk. The task's Dockerfile or compose stack runs unmodified.
- Harbor creates the machines through the boxd Python SDK, runs the agent and the tests inside, copies the results out, and destroys the machine. Nobody at Moyai starts or cleans up a machine by hand.
- Every machine is the same full-size VM. There is no size to pick per task, and no task fails because its sandbox was too small.
- Any machine can be opened while it runs.
sshinto it, look at the containers, read the files, the same as on a server of your own. - The benchmark run itself lives on a boxd machine. Harbor runs on one main machine and creates the environment machines from there, so a run does not depend on anyone's laptop staying open.
- Machines boot from boxd's pre-warmed template. The median time from creation to running was 5 milliseconds over all 279,577 machines, 12 milliseconds at the 95th percentile.
- The whole run is one flag on the Harbor command line:
export BOXD_API_KEY=<your-key>
harbor run --dataset <benchmark> --agent claude-code --model anthropic/claude-opus-4-1 -e boxdMoyai started with 300 machines in the week of 10 August. Six weeks later the weekly figure was 131,903.
The tasks came from 877 different benchmark problems in open source projects, among them Django, SymPy, Astropy, Matplotlib, scikit-learn, xarray, Sphinx, pytest, and Requests. On the busiest day, 21 September, Harbor created 36,335 machines, and 3,064 of them inside a single hour. In total it ran more than 5 million commands inside those machines and uploaded 1.9 million files to them.
“One thing that is super powerful about boxd is that you get the main VM to run the benchmark, and the sandboxes spawn from that. You start treating benchmark runs like atomic functional flows, instead of having to monitor progress and success rates. Or even worse, not being able to close your laptop because that is what is spawning all these sandboxes.”
Next, Moyai plans to run more complex and more realistic agent tasks, to surface failure patterns that today's benchmarks hide. The goal is higher recall and precision for its platform, which is what keeps its customers' agents in check in production.
See it on your stack
Moyai did not build a benchmark cluster. It pointed Harbor at boxd and got a clean computer every time Harbor asked for one, 279,577 times. That is the part of the job we care about at boxd: a machine that is there the moment the work needs it and gone when the work is done.
The provider is awaiting merge upstream. You do not have to wait for it. Install Harbor from the pull request's branch and the command above works today:
uv tool install "harbor[boxd] @ git+https://github.com/MichielMAnalytics/harbor.git@add-boxd-environment"If you run agent benchmarks and your sandboxes are the slow part, we would like to show you Moyai's setup and what it would take on yours. Please shoot me an email at michiel@boxd.sh.



