Quickstart & demo

Run a free eval in three steps

Run a real coding-agent and harness eval without setting up eval infrastructure. Connect mo eval to Claude Code, choose a task and model from our shared catalog, and we’ll run and grade the benchmark for you.

It takes three steps to go from an empty terminal to a graded result.

In a hurry? Skip to the results explorer to browse real runs with no sign-up required.

*Momento covers the model tokens used by the hosted eval. Your normal Claude Code usage is subject to your existing Anthropic plan.

  1. Connect Claude Code

    Link your environment

    $ claude

    Connected to Claude Code
  2. Choose task + model

    Pick from our catalog

    Task
    Fix a failing test
    Model
    mo · GLM-5.3
    mo · Claude Opus 5.5
  3. Get a graded result

    See score, pass/fail, and evidence

    92/ 100
    Eval resultPassed
    Correctness100
    Tests pass100
    Code quality85
    Evidence

Quick start

Connect, sign in, and run

  1. Add the server

    Registers mo eval against your user configuration, so it's available from Claude Code.

    claude mcp add --transport http -s user mo-eval https://mcp.evals.gomomento.ai/mcp
  2. Sign in

    Log in using GitHub, Google, or email and password. Approve both permissions on the consent screen: one to read tasks and your own results, one to submit runs and share them.

    claude mcp login mo-eval

    Use the same sign-in every time so all your runs stay in one account.

  3. Start the guided quick start

    Type this in Claude Code. It checks your account, shows you the tasks you can choose from, and proposes the smallest run there is: one task, one configuration, one attempt. Use the starter model route below when it asks, then review the exact arguments and approve the run.

    /mcp__mo-eval__quickstart

    Choose a task from the catalog, use the mo harness, and enter anthropic/claude-opus-5-5 as the starter model route. The quickstart does not list model routes yet.

Check your run

Ask for an update in your own words, like how’s the run?

Good to know: Trial runs use a shared queue, so a single attempt may take up to an hour to complete.

When it’s done

You’ll get the result and supporting evidence, including test results, a patch, and run logs.

From there, you can:

analyze-run

Go into the numbers and explore the results.

design-run

Configure a customized comparison.

Works with Claude Code today. Support for more MCP clients is coming.


What happens

From an empty terminal to a graded run

The guided prompt does the driving. You make three decisions and approve the eval.

Connect

Two commands, one browser sign-in

mo eval appears in Claude Code alongside your other tools. The consent screen names exactly what it is asking for.

Choose

Pick a task and a model

Real tasks mined from real repositories, each with its language, size, category and how many tests grade it, and the model routes you can run them on.

Confirm

See the run before it costs anything

One task, one configuration, one attempt, with its turn and time caps stated. Nothing is submitted until you say so.

Inspect

Read the result and the evidence

What passed, what it cost, how long it took — plus the patch, the logs and the trajectory behind every attempt.


Demo

Pick a repository. Compare harnesses and models.

Three real mo eval runs, shown as mo eval reported them: 194 attempts across Go, Rust and C. Looking costs nothing and starts no run. Every configuration's cost, solve rate and time are shown together, allowing you to see metrics like cost vs. speed trade-offs.

The cheaper model wins on Go and loses on Rust — same two models, same harness, different language. A single run on either one would have told you the opposite of the other.

Sort and chart by


Good to know

What the free demo does and doesn't cover

Momento pays for the eval

The model tokens a hosted run spends are on us. Your Claude Code usage is billed by your own plan, as usual.

Tasks come from a shared catalog

Every trial user sees the same curated tasks. We'd love to work with you on testing tasks from your own repositories and using your full AI workflows: clients, harnesses, models, skills, etc.

Runs are bounded

At most 1,000 cells in one run, tasks × configurations × repeats. At most 20 runs per account per day. Both are guards against runaway runs. Not production limits.

Results are private until you share them

A run is readable by whoever submitted it. Share it with a colleague by the email address they sign in with, or mint a link that expires.


Get started

Put your own repositories to the test

The demo shows you mo eval working on our selected tasks. To decide which model and harness to standardize on, you need results from your own code. Send us a few real repositories and see the benchmark data, cost comparison, and gateway configuration before you commit to anything.

Request a workload audit →