How Moyai ran 279,577 benchmark machines

Moyai finds failing AI agents in production traces. Its Harbor benchmark runs get a boxd machine per environment, booted in a median of 5 milliseconds.

Michiel VoortmanMichiel Voortman4 min readCustomer stories
A card on near-black: the kicker CUSTOMER STORY in a thin box, and the boxd logo and the Moyai logo side by side, joined by a small ×, over a faint field of rose wireframe cubes.

Moyai catches AI agents that go wrong in production. It reads the traces a team already collects in Langfuse, LangSmith, or OpenTelemetry, finds the runs where an agent behaved abnormally, checks which of those are real failures, and tells the team what broke before a customer does. To surface the ways agents fail, and to measure how many of them its platform catches, Moyai runs coding agents against benchmark tasks with Harbor, at a large scale. Every task needs a clean computer of its own. Since August 2026, Harbor has created 279,577 boxd machines for Moyai, each one a full Linux VM that booted in a median of 5 milliseconds.

279,577
machines created by Harbor, each a full Linux VM
5 ms
median time from creation to running, and 12 ms at the 95th percentile
36,335
machines created on the busiest day, 21 September
877
different benchmark tasks, run unmodified

Before: Modal was expensive and hard to debug

Moyai runs hundreds of benchmark tasks, many times each, and every run needs a fresh environment. Before boxd, Moyai ran them in sandboxes on Modal. Three things got in the way.

  • Cost. At this volume the sandbox bill became a real line item. On one day Moyai spent 1,000 USD on sandboxes alone.
  • Sizing. Each sandbox needed a CPU and memory size chosen up front, and it was never clear what a given task needed. Too small and the task failed for reasons that had nothing to do with the agent. Too large and every one of thousands of runs cost more than it should.
  • No way in. When a run behaved strangely, nobody at Moyai could go into the machine and look. They had the logs the sandbox gave them and nothing else.

After: one machine per environment, created by Harbor

Harbor evaluates agents such as Claude Code, Codex, and OpenHands against benchmarks, across many sandboxes at once. We wrote its boxd environment provider together with Robert Hommes, Moyai's founder.

  • Every environment is its own boxd machine: a KVM virtual machine with a full Docker daemon and a 100 GB disk. The task's Dockerfile or compose stack runs unmodified.
  • Harbor creates the machines through the boxd Python SDK, runs the agent and the tests inside, copies the results out, and destroys the machine. Nobody at Moyai starts or cleans up a machine by hand.
  • Every machine is the same full-size VM. There is no size to pick per task, and no task fails because its sandbox was too small.
  • Any machine can be opened while it runs. ssh into it, look at the containers, read the files, the same as on a server of your own.
  • The benchmark run itself lives on a boxd machine. Harbor runs on one main machine and creates the environment machines from there, so a run does not depend on anyone's laptop staying open.
  • Machines boot from boxd's pre-warmed template. The median time from creation to running was 5 milliseconds over all 279,577 machines, 12 milliseconds at the 95th percentile.
  • The whole run is one flag on the Harbor command line:
Terminal
export BOXD_API_KEY=<your-key>
harbor run --dataset <benchmark> --agent claude-code --model anthropic/claude-opus-4-1 -e boxd

Moyai started with 300 machines in the week of 10 August. Six weeks later the weekly figure was 131,903.

MACHINES CREATED PER WEEK, 202630010 Aug6.3k17 Aug18k24 Aug21k31 Aug21k7 Sep81k14 Sep132k21 Sep
Machines Harbor created on boxd for Moyai, by week starting Monday, 11 August to 27 September 2026. 279,577 in total. Counted from boxd's own event log.

The tasks came from 877 different benchmark problems in open source projects, among them Django, SymPy, Astropy, Matplotlib, scikit-learn, xarray, Sphinx, pytest, and Requests. On the busiest day, 21 September, Harbor created 36,335 machines, and 3,064 of them inside a single hour. In total it ran more than 5 million commands inside those machines and uploaded 1.9 million files to them.

“One thing that is super powerful about boxd is that you get the main VM to run the benchmark, and the sandboxes spawn from that. You start treating benchmark runs like atomic functional flows, instead of having to monitor progress and success rates. Or even worse, not being able to close your laptop because that is what is spawning all these sandboxes.”

Robert Hommes, Founder, Moyai

Next, Moyai plans to run more complex and more realistic agent tasks, to surface failure patterns that today's benchmarks hide. The goal is higher recall and precision for its platform, which is what keeps its customers' agents in check in production.

See it on your stack

Moyai did not build a benchmark cluster. It pointed Harbor at boxd and got a clean computer every time Harbor asked for one, 279,577 times. That is the part of the job we care about at boxd: a machine that is there the moment the work needs it and gone when the work is done.

The provider is awaiting merge upstream. You do not have to wait for it. Install Harbor from the pull request's branch and the command above works today:

Terminal
uv tool install "harbor[boxd] @ git+https://github.com/MichielMAnalytics/harbor.git@add-boxd-environment"

If you run agent benchmarks and your sandboxes are the slow part, we would like to show you Moyai's setup and what it would take on yours. Please shoot me an email at michiel@boxd.sh.

Michiel VoortmanMichiel Voortman
PostShare
Published
Oct 7, 2026
Reading time
4 min
Words
830
Topic
Customer stories

Read next

Field notes

Subscribe for release notes and architecture write-ups

No spam, ever. Unsubscribe anytime.

Your inbox