Three automated reviews, nineteen findings, four real bugs
The studio holds itself to one rule: say only what it measures. This page is that rule turned around on the writing about the studio: what an adversarial pass found, including the things it found in the writing’s own claims, and the defects it found in the compiler underneath them.
Which version was checked
This is a historical record of the earlier tracing-and-hatching compiler and its publication drafts. The three reviewers were AI agents asked to inspect the repository from different disciplinary perspectives; this was an internal automated review, not external academic peer review. The test counts and measurements below belong to that review. They do not validate the newer 44-second drawing version, which adds vector pigment regions and timed contacts. Its approach and limits are described in What happens when you sit.
The method
The book was drafted thirty chapters at a time, each by an agent working from a written brief that named the sources it had to read before it wrote a word. Every finished chapter then went to a second agent whose instructions were the opposite of the first’s: not to improve the draft, not to praise it, but to find what was wrong with it, reading the repository directly rather than reading the chapter’s own account of the repository.
Findings came back in four classes, and the class mattered more than the wording:
- Wrong: the repository says something else.
- Unsupported: the repository says nothing either way.
- Overclaim: true, but stated more strongly than the evidence carries.
- Stale: true once, and quietly not true now.
Each was fixed in place rather than argued with. The full record (every drafting report and every verification report, verbatim, roughly forty-two thousand words of it) is kept as a working document rather than a summary, because a summary of a fact-check is exactly the kind of artifact that can quietly stop matching what it summarises.
The pass worked against its own editor
The interesting result was not the chapters. Agents drafting from briefs make ordinary mistakes and a checker catches them; that is the arrangement working as designed. What was worth reporting is that eight corrections came back up the chain and changed material written for the publication rather than by a chapter agent, the editorial writing that was supposed to be the reliable part.
- The design system’s own token count, given as 165, is 148, and two rows of its per-group table were wrong with it.
- The stated contrast ratios: none of the four pigment
deepstops annotated at about 7:1 actually reaches 7:1. --r-none, presented as the token that keeps the artwork square, is declared once and referenced nowhere. It is still declared once and referenced nowhere; it is simply no longer described as doing something.- The legacy alias block, dated to v5 when it is v3-born.
- The front matter’s “43,000 lines in eleven days,” when the first commit lands 41,691 insertions from a history kept somewhere else entirely.
- An appendix and a figure that both presented
headPassas a working mechanism, when the compiler never read it. - A figure caption naming a four-term scoring formula that is really five terms, and does not include stillness, which is the term the caption leaned on.
- A diagram of the import graph, redrawn after a checker found that the module at its centre is imported by two of the four renderers, not four.
None of these were load-bearing for the product. All of them were load-bearing for the claim that the documentation could be trusted, which is the claim the whole exercise existed to make.
Three reviewers, nineteen findings
The paper went through a deeper pass than the book did. Three reviewers, briefed with deliberately different lenses (empirical software engineering; computer graphics and non-photorealistic rendering; human-computer interaction and research ethics), each reading independently against the repository rather than against each other, so that agreement between them would mean something.
They returned nineteen major findings: four from the methods reviewer, six from the graphics reviewer, nine from the ethics reviewer. The revision record marked eighteen fixed by inspection of the paper at that time. Twelve minor and bonus findings were recorded fixed alongside them. Those are the review’s recorded outcomes, not a fresh verification of every claim in the current paper.
One is not. The ethics reviewer asked, on a second reading, for a paragraph positioning honesty contracts against the adjacent art the paper had skipped: banned-phrase CI checks, restricted-syntax lint rules, architecture fitness functions. The editor could not confirm that paragraph had been written before the session ran out, and it is recorded as open rather than assumed closed. It corrects no error; it is a missing piece of intellectual honesty about the technique’s neighbours. Reporting it as open costs nothing except the appearance of a clean sweep, which is not a thing worth buying.
Where the review stopped being about prose
The graphics reviewer’s questions about the compiler’s arithmetic did not stop at the paper. Answering them precisely enough to publish sent the revision back into the compiler itself, where it found three real defects, none of which the existing 173-test suite had ever caught, because none of them throws. Each one surfaces only as a shift in the compiled document’s own statistics, which is exactly what a suite that checks whether code runs rather than what it produces will never see.
A remainder that was supposed to be a modulo
The tone hatcher thinned its sweep with u % 3 < 1, intending to keep about one point in three. JavaScript’s % returns a remainder, not a modulo: for negative u the result lands in (-3, 0], which is always less than 1. Every point on the negative side of a sweep survived instead of one in three. Point counts up to tripled, and since darkness accumulates at every step but divides on the assumption of one-in-three sampling, the hatching was also reporting itself lighter than it was by close to the same factor.
A width measured to the wrong edge
Stroke width is recovered with a distance transform, which was seeded so that every non-traced pixel counted as background, including every pixel of tone mass. A traced line running up against a mass therefore measured its width to the mass’s edge rather than to the paper beyond it, and clamped near the compiler’s floor no matter how much real paper was actually there.
Five hands with no rim at all
The density split sent a mass’s outermost ink column to line work only once its local density fell under the artist’s gate, but under the compiler’s own 11×11 window that column already reads about 0.545. Five of the eight hands (Odile, Sol, Mira, Halia and Kavi) carry a gate below that number, so all five got no traced rim at any straight mass boundary, however thin the edge. The edge a viewer’s eye reads as the line was simply not being drawn.
All three are fixed. Because the fixes change what the compiler produces rather than whether it runs, the honest way to report them is with the numbers, measured on the same bundled fixture before and after:
| Measured | Before | After |
|---|---|---|
| Traced contour and construction strokes | 270 | 568 |
| Tone-stroke opacity, median | 0.61 | 0.64 |
| Strokes in the document | 2,509 | 3,874 |
The fourth bug, and the tests that did not catch it
Two of the three fixes ship with a regression test, each confirmed to fail when its own fix alone is reverted. The third does not. A synthetic scene built to isolate it would not discriminate reliably: the boundary pixels the rim fix adds are, by construction, already next to real paper, so it is covered by the suite passing and by the written analysis instead, and the test file says so in as many words rather than carrying a test that overclaims what it proves.
Then the same remainder-versus-modulo defect turned up a second time, one function over, in the wash extractor: u % 6 < 2 where the tone sweep had u % 3 < 1. The wash sweep’s angle guarantees negative coordinates, so every negative sample had been surviving there too. This one never corrupted a stored value (wash darkness is a fixed constant rather than an accumulated sum), but it had been laying wash up to three times denser on one side of a stroke than the other.
| Measured | Before | After |
|---|---|---|
| Wash path points | 4,191 | 2,382 |
| Points in the heaviest single stroke | 86 | 34 |
The sharper finding came next, and it is a finding about the tests rather than the code. Both decimation tests stayed green when either call site was reverted to its raw expression, because neither test exercised a call site: they tested the shared function, which was correct. A test can pin a formula perfectly and still not notice that nothing calls it. The fix was a third test at the integration level: compile the real fixture, and hold the actual tone and wash output to a per-stroke point budget.
What happened afterwards
One correction deserves its ending told, because the ending is better than the report. Several hands carry a pass named for the face itself (“The exact face,” “Eyes, set once,” “A single hair’s line”) and the name implies a mechanism: that the compiler gathers strokes from the measured region of the face into that pass, so the name is true of what the sitter is watching while it plays. The fact-check found the mechanism was not there. The passes were filled by the same length-and-darkness ranking as every other pass; what appeared during them was a scatter across the whole sheet.
The correction at the time was to fix the claim rather than the code. That was the right order: a pass name is a promise made to somebody watching in real time, and a promise repaired afterwards by quietly building the thing it promised is not the same promise. So the documentation stopped saying it, and the paper reported it as a claim withdrawn.
A mechanism was built four days later. That compiler estimated a head region from the width of the ink down the sheet, looking for the widening from head to shoulders without a server-side face detector. It gave the pass a spatial rule, rather than a name alone. It remained a heuristic: an unusual crop, pose, garment or subject can defeat that assumption. The current contact compiler uses its own proportional head estimate, so neither version establishes reliable recognition of facial features or a general drawing plan for non-face subjects.
The complete record is kept in the repository rather than summarised: every drafting and verification report, all three reviews in full, and the revision log with a verification command against each finding so it can be re-checked without trusting the summary. That is the point of publishing any of this. A claim you cannot check is just a nicer-sounding claim.