Edouard Tardif
FREN

Live · internal use · as of 10 October 2026

Recette — the tool that checks the agents' work

Check every factory delivery in three moves: Verified, Bug, Skip.

For whom
Me, the reviewer
Stack
  • Next.js
  • SQLite
Recette's station, drawn: deliveries roll by, a stamp approves, the gauge shows confidence, three buttons decide.

In figures (as of 10 October 2026)

  • 179

    agent PRs out of 183 merged

    A "PR" (pull request) is a proposed change that is tested, then merged into the software.

  • 194

    closed issues

Build timeline

From 19/07 to 10/10/2026, twelve weeks. The solid bar is Recette; the others show where it sits in the factory.

  1. Chartrium: from 18/08/2026 to 10/10/2026 (snapshot), 621 agent PRs
  2. Merkindium: from 18/08/2026 to 02/10/2026, 440 agent PRs
  3. Recette
  4. Bottrading: from 10/09/2026 to 02/10/2026, 201 agent PRs
  5. Raccourci: from 19/07/2026 to 16/08/2026, 16 agent PRs

01/09/2026 · first merge01/10/2026 · last merge

179 agent PRs out of 183 merged, as of 10 October 2026.

The problem

A factory that ships dozens of changes a day also ships mistakes. Automated tests don't catch everything: a human has to look, but cannot reread everything.

The answer

A verification queue:

  • each delivery arrives with its acceptance steps spelled out: where to go, what to do, what to expect;
  • three buttons: Verified, Bug, Skip;
  • Bug opens a new ticket on the original application, and the factory picks it up;
  • the queue is sorted by ascending confidence: I check first what the agent doubts most.

Three screens

① The queue, with confidence badges. ② An expanded check. ③ The Coverage page: what has been verified, repository by repository.

Recette, verification queue: ticket #940 scored 6/10, "to watch", with the Verified, Skip and Bug buttons
①The queue, with confidence badges: here ticket #940, scored 6/10.

Illustration drawing, not a screenshot.

②An expanded check.

Illustration drawing, not a screenshot.

③The Coverage page: what has been verified, repository by repository.

The first screen is a real screenshot. The other two are illustration drawings, not screenshots.

What was hard

Getting agents to state their doubts.

Every PR must end with "Confidence: N/10 — what was not checked". Across 1,013 declared scores, the average is 7.5, and only 3 are 10/10. A low score is not a failure: it is information.

Closing the loop.

A bug reported in Recette becomes a ticket, then a PR, then a new check.

What is left

A factory control console, being redesigned inside the same application.

Under the hood

Next.js, SQLite, the GitHub API. Some tickets were built by a third model provider, to test switching providers.