Chartrium
In progress



Find a document fast, never lose a version, and be able to show what happened to it.
- For whom
- Small law firms (4 to 10 lawyers)
- Stack
- Go
- React + TypeScript
- PostgreSQL
- Temporal
- Apache Tika
Since 19 July 2026, AI agents have been turning written specifications into real applications: they split the work, write the code, test it and deploy it. I make the decisions, I hold the keys, and I check the work before it counts.
code changes built by agents and accepted
A "PR" (pull request) is a proposed change that is tested, then merged into the software.
applications built in 12 weeks
the agents' own average self-reported confidence, and only 3 scores of 10/10
human to decide and verify
121 scores of 6/10 or lower (12%)3 scores of 10/10
The journey of a real ticket
Here is a real Merkindium ticket from 19 September 2026, step by step. Nothing is staged: it got stuck, it was picked up again, and the agent said what it had not checked.
This ticket did not go through pre-flight: it was born along the way.
Step 1 ·
While working on another ticket, an agent notices that the forum's "New topic" screen does not match its mockup: two "Write / Preview" tabs where the mockup shows a single "Preview" button. It does not decide on its own: it opens ticket #940.
#940opened at the workbench, during ticket #907
a single "Preview" button
two "Write / Preview" tabs
Station 4 · Workbench, during ticket #907
Step 2
I decide: the mockup is the reference, the code gets aligned. The decision is written at the top of the ticket, with what to do and how to check it. The decision also goes into the specification, as question Q-137.
Station 1 · Specification
Step 3 ·
An agent schedules the work, in dependency order, with a limited number of tickets in progress at once.
· "ready" label set by the project-manager agent, in dependency order
Station 3 · Splitting
Step 4 ·
On the first try the agent gets stuck after four minutes. The factory hands the ticket to another model and starts again. Four retries later, the work is ready.
five attempts
Station 4 · Workbench
Step 5 ·
The agent opens PR #946: one button instead of two tabs, typed text never lost, labels translated into French, English and Spanish, and tests to prove it.
#940PR #946
From station 4 to station 5 · the ticket becomes "PR #946"
Step 6
Seven checks run. One of them, required by the ticket, could not run: another test was already failing on the main version. The agent says so in plain words instead of ignoring it.
6 checks ran1 could not run, and the PR says so
Station 5 · Test bench
Step 7 ·
Supervision accepts the PR, the ticket closes, and the new version deploys itself to the staging server.
"Staging" is a private copy of the application where things are checked before the public sees them.
Station 6, then station 7 · Supervision, then staging showcase
Step 8
The PR lists five steps to check it, and ends with an honest score:
"Confidence: 6/10 — […] no real browser check."
This is exactly why a human goes over the work.
Station 8 · Human verification station
What the factory has built
Each one keeps its own visual identity. The badge says where it stands as of 10 October 2026, without rounding up.
In progress



Find a document fast, never lose a version, and be able to show what happened to it.
In staging






The companion site of a home-DIY YouTube channel: sorted tutorials, courses, forum and gallery.
Live (internal use)

Check every factory delivery in three moves: Verified, Bug, Skip.
In progress · private access



Design, test and monitor trading agents on your own account, without the platform ever holding funds.
No performance, no returns, no investment advice: this shows the design, not gains.
Live

The very first application: a link shortener used to tune the factory.
What gets built, and what to do when the specification contradicts itself.
No password or payment key ever goes into the code: the factory builds without knowing them.
Every delivery goes through a human check before it counts.
Subscriptions have caps. When they are reached, the factory slows down or stops.
They rarely check in a real browser. 121 of their confidence scores are 6/10 or lower, and they say so.
A real share of the work goes into fixing tests that fail at random: 40 PRs out of 440 in Merkindium.
A ticket can stall several times: the one on the home page took five attempts.
Deciding, checking, unblocking: that time cannot be delegated.