Quickstart & demo
Run a free eval in three steps
Run a real coding-agent and harness eval without setting up eval infrastructure. Connect mo eval to Claude Code, choose a task and model from our shared catalog, and we’ll run and grade the benchmark for you.
It takes three steps to go from an empty terminal to a graded result.
In a hurry? Skip to the results explorer to browse real runs with no sign-up required.
*Momento covers the model tokens used by the hosted eval. Your normal Claude Code usage is subject to your existing Anthropic plan.
-
Connect Claude Code
Link your environment
$ claude
Connected to Claude Code -
Choose task + model
Pick from our catalog
TaskFix a failing testModelmo · GLM-5.3mo · Claude Opus 5.5 -
Get a graded result
See score, pass/fail, and evidence
92/ 100Eval resultPassedCorrectness100Tests pass100Code quality85Evidence
Quick start
Connect, sign in, and run
-
Add the server
Registers mo eval against your user configuration, so it's available from Claude Code.
claude mcp add --transport http -s user mo-eval https://mcp.evals.gomomento.ai/mcp -
Sign in
Log in using GitHub, Google, or email and password. Approve both permissions on the consent screen: one to read tasks and your own results, one to submit runs and share them.
claude mcp login mo-evalUse the same sign-in every time so all your runs stay in one account.
-
Start the guided quick start
Type this in Claude Code. It checks your account, shows you the tasks you can choose from, and proposes the smallest run there is: one task, one configuration, one attempt. Use the starter model route below when it asks, then review the exact arguments and approve the run.
/mcp__mo-eval__quickstartChoose a task from the catalog, use the
moharness, and enteranthropic/claude-opus-5-5as the starter model route. The quickstart does not list model routes yet.
Check your run
Ask for an update in your own words, like how’s the run?
Good to know: Trial runs use a shared queue, so a single attempt may take up to an hour to complete.
When it’s done
You’ll get the result and supporting evidence, including test results, a patch, and run logs.
From there, you can:
analyze-runGo into the numbers and explore the results.
design-runConfigure a customized comparison.
Works with Claude Code today. Support for more MCP clients is coming.
What happens
From an empty terminal to a graded run
The guided prompt does the driving. You make three decisions and approve the eval.
Connect
Two commands, one browser sign-in
mo eval appears in Claude Code alongside your other tools. The consent screen names exactly what it is asking for.
Choose
Pick a task and a model
Real tasks mined from real repositories, each with its language, size, category and how many tests grade it, and the model routes you can run them on.
Confirm
See the run before it costs anything
One task, one configuration, one attempt, with its turn and time caps stated. Nothing is submitted until you say so.
Inspect
Read the result and the evidence
What passed, what it cost, how long it took — plus the patch, the logs and the trajectory behind every attempt.
Demo
Pick a repository. Compare harnesses and models.
Three real mo eval runs, shown as mo eval reported them: 194 attempts across Go, Rust and C. Looking costs nothing and starts no run. Every configuration's cost, solve rate and time are shown together, allowing you to see metrics like cost vs. speed trade-offs.
The cheaper model wins on Go and loses on Rust — same two models, same harness, different language. A single run on either one would have told you the opposite of the other.
Good to know
What the free demo does and doesn't cover
Momento pays for the eval
The model tokens a hosted run spends are on us. Your Claude Code usage is billed by your own plan, as usual.
Tasks come from a shared catalog
Every trial user sees the same curated tasks. We'd love to work with you on testing tasks from your own repositories and using your full AI workflows: clients, harnesses, models, skills, etc.
Runs are bounded
At most 1,000 cells in one run, tasks × configurations × repeats. At most 20 runs per account per day. Both are guards against runaway runs. Not production limits.
Results are private until you share them
A run is readable by whoever submitted it. Share it with a colleague by the email address they sign in with, or mint a link that expires.
Get started
Put your own repositories to the test
The demo shows you mo eval working on our selected tasks. To decide which model and harness to standardize on, you need results from your own code. Send us a few real repositories and see the benchmark data, cost comparison, and gateway configuration before you commit to anything.
Request a workload audit →