nutmeg 0.5: analysis that shows its work

nutmeg 0.5 turns analysis you will publish or decide on into a research project. Every number ties to the run, source or ID behind it, and the person who publishes can defend it.

nutmeg

An AI agent can write a recruitment shortlist or an opposition report in minutes. The numbers look right. But when someone asks “where does 0.31 come from?”, nobody can say which data, which filter, which definition of the metric, or whether the input changed after the run.

And “looks right” is not enough. In a 2026 study, AI agents given different personas reached opposing conclusions from the same data and the same question, and most of those analyses still passed review: 86% passed independent AI review and 78% passed majority human expert review (Miao, Pritchard and Zou). Another study found that flaws in AI research systems, such as data leakage and post-hoc selection, are far easier to catch from the trace logs and code than from the final paper (Luo, Kasirzadeh and Shah).

When defensible analyses are cheap, the question stops being “is this analysis plausible?” and becomes “which choices produced this number, and would another reasonable choice change it?”

nutmeg 0.5 is built for that question.

What a research project is

For analysis you will publish or decide on, /nutmeg:research starts a research project. It has three parts:

The skills run a nutmeg command for you, and it does the checking:

Metric definitions come from the football-docs metric cards, and nutmeg names the exact version it uses, for example ppda.statsbomb-hudl. So “PPDA 7.42” in a report says which PPDA.

Robustness you can’t drop

The plan locks at the first run, and any change after that is listed. Before you publish, each headline result shows how it holds up under two or more defensible alternatives (a different filter, a different metric version, a different model). Once recorded, none of them can be dropped, and the publish card shows them all, so a result that only holds under one choice says so.

Understanding before sharing

A checkable number is only half of it. The person publishing also has to be able to defend it.

What you can do with it

  1. A recruitment shortlist a head of recruitment can check line by line. Every number on the list traces back to its run and its metric definition.
  2. A match or opposition report for coaching staff, where the claims that matter show how they hold up under other reasonable choices.
  3. A public thread whose numbers survive the replies. When someone asks where a number comes from, nutmeg why answers with the run, the data and the definition.
  4. Handing a project to a colleague. nutmeg bundle gives them the question, the plan, the ledger and the runs, not just the final chart.

For a team there are autonomy levels (suggest, draft, or execute with checkpoints), a team config with limits and sign-off rules, and secrets are redacted in every record.

Why not just use…

What it doesn’t do

It also costs less to run

Outside research projects, nutmeg now uses about 1,850 tokens of context per session, down from about 2,650 during development. The research rules load only in projects with a research/ folder, and publishing guidance loads only when you publish.

The public eval suite has 48 cases, including point-in-time leakage, understanding, provenance and robustness. On the final changes (Sonnet 5.5 as agent and judge) it scores 0.93 to 0.95 at about $0.05 per case, and a private held-out set scores 0.95.

Try it

/plugin marketplace add withqwerty/plugins
/plugin install nutmeg@withqwerty

Already installed? Run /plugin update nutmeg. Then start a project with /nutmeg:research.

What would you want a research project to check that it doesn’t yet? I’m taking requests.