Edouard Tardif
FREN

How it works

An eight-step chain. Agents at every station, a human at the three places that matter.

The production chain

Scroll down: a ticket moves from station to station, and each station comes alive when you reach its text.

The production chain, in eight stationsEight stations arranged in a zigzag along a conveyor, from the specification (1) to the human verification station (8). A "Ticket" card moves from station to station. At the workbench (4) it becomes a "PR"; at stations 5 and 6 a "ci-ok" tag is hung on the card. Without animation, the card rests at station 8.1Specification2Pre-flight3Splitting4Workbench5Test bench6Supervision7Staging showcaseVerified8Human verification stationTicketPRci-ok
  1. Step 1 of 8

    The specification

    Requirements written so they can be tested, acceptance criteria, architecture decisions, mockups.

    Station 1 · Specification

  2. Step 2 of 8

    Adversarial pre-flight

    Three agents hunt for contradictions in the specification before anything starts. 59 findings on Merkindium, all resolved before the first ticket.

    Station 2 · Pre-flight

  3. Step 3 of 8

    Splitting the work

    The project-manager agent splits it into batches and tickets, with their dependencies.

    Station 3 · Splitting

  4. Step 4 of 8

    Building

    An agent takes a ticket and proposes a change (PR).

    Station 4 · Workbench

  5. Step 5 of 8

    Testing

    Every PR runs the automated tests, plus a real start of the application in a container.

    Station 5 · Test bench

  6. Step 6 of 8

    Supervision

    If everything is green, the PR is merged. Otherwise a fixer agent takes it back.

    Station 6 · Supervision

  7. Step 7 of 8

    Staging

    The new version deploys itself to a private copy.

    Station 7 · Staging showcase

  8. Step 8 of 8

    Human verification

    In Recette: Verified, Bug or Skip. A bug becomes a new ticket.

    Station 8 · Human verification station

What the human does

Three things that stay with the human.

Decides

the specification, the trade-offs when it contradicts itself, what goes to production.

Holds the secrets

payment keys, passwords, server access. Agents see none of them.

Verifies

in staging, starting with what the agents doubt.

Safeguards

Five protections built into the chain.

  • A mandatory check (ci-ok)

    Nothing is merged without green tests.

  • A daily budget per application

    Each one has a daily consumption cap.

  • No secrets in the code

    Checked automatically.

  • A stub mode for every external service

    The factory builds without holding any key.

  • A limited number of tickets in progress

    In dependency order.

The means

GitHub for tickets and tests, two families of agents (Claude Code and Codex), twelve build machines on servers rented from OVH, Coolify for deployments.

What it costs (order of magnitude)

about €300 a month

servers and agent subscriptions included.

No usage-based cost: agents run on fixed-price subscriptions.

Limits

  • Quotas.

    Subscriptions have caps. When they are reached, the factory slows down or stops, and the project-manager agent alone uses a large share.

  • What agents miss.

    They rarely check in a real browser. They can ship code that passes the tests and is still wrong on screen. 121 of their confidence scores are 6/10 or lower, and they say so.

  • Flaky tests.

    A real share of the work goes into fixing tests that fail at random: 40 PRs out of 440 in Merkindium.

  • Getting stuck.

    A ticket can stall several times: the one on the home page took five attempts. When the main version breaks, the factory can stay stuck until I decide.

  • The human remains the bottleneck.

    Deciding, checking, unblocking: that time cannot be delegated.

The figures, and how they are counted

Collected from the GitHub API on 10 October 2026; regenerated by script.

  • 1,457

    agent PRs, out of 1,486 merged PRs across the 5 applications shown

    A "PR" (pull request) is a proposed change that is tested, then merged into the software.

  • 1,681

    closed issues (tickets, bugs, dropped ones included)

  • 7.5/10

    average of the "Confidence: N/10" scores the agents report about themselves

  • 1,013

    agent PRs that state a confidence score

  • 121

    scores of 6/10 or lower (12%), and only 3 scores of 10/10

  • 12 weeks

    from 19 July 2026 to 10 October 2026; first repository on 19 July 2026

Self-reported confidence scores

The 1,013 agent PRs that state a "Confidence: N/10" score, grouped by score. Average: 7.5/10.
  1. 40/10: 4 scores
  2. 01/10: 0 scores
  3. 02/10: 0 scores
  4. 23/10: 2 scores
  5. 24/10: 2 scores
  6. 135/10: 13 scores
  7. 1006/10: 100 scores
  8. 2787/10: 278 scores
  9. 5238/10: 523 scores
  10. 889/10: 88 scores
  11. 310/10: 3 scores

121 scores of 6/10 or lower (12%)3 scores of 10/10

Pull requests (PR), closed issues and average confidence, by application, as of 10 October 2026
ApplicationAgent PRs / merged PRsClosed issuesAverage confidence
Chartrium621 / 6407207.5
Merkindium440 / 4455257.4
Bottrading201 / 2022167.7
Recette179 / 1831947.6
Raccourci16 / 1626no score
Total1,457 / 1,4861,6817.5

A PR counts as "built by an agent" when it comes from a ticket branch (ticket/…). All of them are opened under my GitHub account, which the factory uses. The confidence score is self-reported by the agent. The rule came in along the way: Raccourci did not have it, nor did some older Chartrium and Merkindium PRs. Collected from the GitHub API on 10 October 2026; regenerated by script.

5 applications shown.