nutmeg 0.5: analysis that shows its work
nutmeg 0.5 turns analysis you will publish or decide on into a research project. Every number ties to the run, source or ID behind it, and the person who publishes can defend it.
An AI agent can write a recruitment shortlist or an opposition report in minutes. The numbers look right. But when someone asks “where does 0.31 come from?”, nobody can say which data, which filter, which definition of the metric, or whether the input changed after the run.
And “looks right” is not enough. In a 2026 study, AI agents given different personas reached opposing conclusions from the same data and the same question, and most of those analyses still passed review: 86% passed independent AI review and 78% passed majority human expert review (Miao, Pritchard and Zou). Another study found that flaws in AI research systems, such as data leakage and post-hoc selection, are far easier to catch from the trace logs and code than from the final paper (Luo, Kasirzadeh and Shah).
When defensible analyses are cheap, the question stops being “is this analysis plausible?” and becomes “which choices produced this number, and would another reasonable choice change it?”
nutmeg 0.5 is built for that question.
What a research project is
For analysis you will publish or decide on, /nutmeg:research starts a research project. It has three parts:
- A question card. What you are asking, written down before any data is touched.
- A plan. Every choice (the data, the filters, the metric, the model) has a one-sentence reason and the source it rests on.
- A claim ledger. Each number, provider fact, ID and citation in the output is tied to its evidence: the recorded run, the football-docs page, the Reep ID or the paper behind it.
The skills run a nutmeg command for you, and it does the checking:
nutmeg runshows a gate card before anything runs: the code, the inputs and any data sent out. Then it records the run.nutmeg checklists every number in the outputs that has no claim, every claim whose input changed after the run, and every run that changed its own input.nutmeg whyandnutmeg traceexplain where a number comes from.contest,resolveandsignofflet a teammate check claims.nutmeg data addrecords where each input came from, with a profile that flags duplicated rows, repeated IDs and empty columns.nutmeg publishwaits until open problems are fixed or accepted with a reason, andnutmeg bundlehands the project to someone else.
Metric definitions come from the football-docs metric cards, and nutmeg names the exact version it uses, for example ppda.statsbomb-hudl. So “PPDA 7.42” in a report says which PPDA.
Robustness you can’t drop
The plan locks at the first run, and any change after that is listed. Before you publish, each headline result shows how it holds up under two or more defensible alternatives (a different filter, a different metric version, a different model). Once recorded, none of them can be dropped, and the publish card shows them all, so a result that only holds under one choice says so.
Understanding before sharing
A checkable number is only half of it. The person publishing also has to be able to defend it.
- nutmeg explains with your own numbers, and offers once, at a useful moment, to write an explainer page you can edit and share.
- Before you publish, it checks that you can defend the work. For data scientists and researchers this is a short reviewer-style challenge on the weakest points. For newcomers it is a guided talk-through. If your messages already show you understand the work, it doesn’t ask. Your answers stay in your own nutmeg folder and are never published.
- If you ask to change a method or drop a caveat after seeing the results, you get a plain answer, and the change is listed on the publish card.
What you can do with it
- A recruitment shortlist a head of recruitment can check line by line. Every number on the list traces back to its run and its metric definition.
- A match or opposition report for coaching staff, where the claims that matter show how they hold up under other reasonable choices.
- A public thread whose numbers survive the replies. When someone asks where a number comes from,
nutmeg whyanswers with the run, the data and the definition. - Handing a project to a colleague.
nutmeg bundlegives them the question, the plan, the ledger and the runs, not just the final chart.
For a team there are autonomy levels (suggest, draft, or execute with checkpoints), a team config with limits and sign-off rules, and secrets are redacted in every record.
Why not just use…
- A general AI assistant? It can write the code and the report, but the report doesn’t carry where each number came from, and the trace is where the flaws show. Nothing checks that the inputs are unchanged since the run, that the result holds under other choices, or that the person publishing understands it.
- A notebook? A notebook shows the code, but not why each choice was made, which definition a metric follows, or which numbers in the write-up came from which cell.
- Doing it by hand? You can keep a claim log yourself, and good analysts do. nutmeg does the bookkeeping and the checks, so it happens every time, not only when there is time.
What it doesn’t do
- It doesn’t make a number right, only checkable. A wrong input or a bad choice is still wrong. It is just visible.
- It doesn’t replace your judgement. It is an assistant, not an autopilot, and the autonomy levels decide how much it does on its own.
- It doesn’t get data you don’t have. Paid provider data still needs your own access.
- ID joins across providers need Reep. Download the free register and set
REEP_DUCKDB_PATH, or setREEP_API_KEY. Without either, it explains the set-up instead of matching IDs. - Research projects need Python 3.10 or newer.
It also costs less to run
Outside research projects, nutmeg now uses about 1,850 tokens of context per session, down from about 2,650 during development. The research rules load only in projects with a research/ folder, and publishing guidance loads only when you publish.
The public eval suite has 48 cases, including point-in-time leakage, understanding, provenance and robustness. On the final changes (Sonnet 5.5 as agent and judge) it scores 0.93 to 0.95 at about $0.05 per case, and a private held-out set scores 0.95.
Try it
/plugin marketplace add withqwerty/plugins
/plugin install nutmeg@withqwerty
Already installed? Run /plugin update nutmeg. Then start a project with /nutmeg:research.
What would you want a research project to check that it doesn’t yet? I’m taking requests.