The Performance Is the Artifact: Drawing AI Portraits Stroke by Stroke, and Proving It
Abstract
Generative image models are very good at producing a finished picture and very bad at showing their work. When a product tells you “the AI is drawing your portrait,” what you almost always get is a still image with an animation layered on top of it afterward — a wipe, a fade, a fake pencil scribbling across pixels that were already decided. The animation is theater. Nothing about the process it shows is true.
This paper describes a system built to close that gap: a portrait studio where a sitter is observed for a few seconds by a webcam, and a portrait is then drawn in front of them, one stroke at a time, by one of eight distinct artistic hands. The central design decision is that the drawing the sitter watches and the picture they keep are not two different things connected by a promise — they are mechanically the same object. A generative model is used, but only to produce a private reference study that a deterministic compiler then traces into an ordered sequence of strokes. That sequence, and nothing else, is what gets performed, stored, replayed, and exported. I describe the compiler’s algorithm in enough mechanical and mathematical depth to be reproduced — the local-neighborhood threshold that separates a traceable line from a shaded mass, the classical thinning and simplification steps that turn a mask into paths, the seeded ordering that makes replay exact rather than approximate — and the small, named parameter table that gives each of the eight hands its character, published in full because making style auditable rather than opaque turns out to be a useful property in its own right. I also describe a technique I ended up calling “honesty by design”: encoding the specific claims a product is allowed to make, including claims buried in its own code, as tests that fail automatically the moment the system stops living up to them — and, holding this paper to the same standard, I report two real defects that writing it down carefully enough to publish found in the compiler itself. I report in full a fairness defect in the compiler’s own normalization step — an adaptive local background estimate that discarded a sitter’s face in proportion to how dark it was, in every one of the eight hands — together with the two changes that fixed it and the standing per-hand regression test that now guards it. Finally, I report plainly what I have not yet measured: no perceptual study, no cross-device performance benchmark, and no fairness evaluation on real faces rather than generated ones.
Implementation note — 7 September 2026
The manuscript below documents the earlier traced-line and synthesized-hatching compiler. Its parameter tables, timings, and evaluation results describe that version. They are retained as a record, not presented as measurements of the new drawing version.
The new drawing version uses a versioned source-colour contact extension. It retains the generated study’s ink and pigment as quantized vector regions, and deposits those materials through explicit drawing paths with a radius, contact duration, and lift interval. The complete replay artifact therefore contains both the vector materials and the timed contacts, rather than centerlines alone. The study raster is not embedded in the export. The live canvas, saved sheet, SVG, and film consume that recorded geometry; rasterization can still differ slightly between renderers.
This is an authored reconstruction from a generated study. It is not a recording of a human artist, a recovery of the image model’s original drawing process, or evidence that eight autonomous artists make decisions while the animation plays. Each residency supplies a medium, study prompt, instrument choices, and pacing. The new drawing window is 35 seconds; observation, image generation, and compilation happen before that window. Short texture contacts are necessarily compressed into a timelapse at this speed.
The reference sheets use a different generated study for each residency. The homepage uses its own Mira study through the same compiler, with no homepage-only deletion or reordering of marks. A version and source hash accompany these examples. This improves traceability; it does not establish quality on unseen portraits. The earlier fairness results and renderer checks below must be repeated for the new version before they can support a public-release claim about it.
1Introduction
Ask a generative image model for “a pencil portrait” and it will hand you a finished pencil portrait. It will not hand you a sequence of pencil marks, because there was never a sequence — a diffusion model does not draw the way a person draws, and nothing about its internal process resembles the motion of a hand across paper. That is fine for most uses of these models. It becomes a problem the moment a product tells its user that the machine is drawing them, live, right now, because that is a claim about how the picture is being made, not just about what it looks like when it’s done.
The common way products paper over this gap is decorative: generate the finished image first, then animate something over it — a reveal, a soft-focus fade-in, a scribble effect timed to look plausible. I think this is worse than it looks. It’s not a simplification of the truth, it’s a fabrication of a process that never happened, wrapped around a result that already existed before the animation started. A user watching it has no way to tell the difference between a system that is actually constructing their portrait and one that is running a canned effect over a photo it already generated.
I wanted to build the honest version of that promise: a system where the thing the sitter watches happen is the thing being made, such that the final stored portrait is not a separate artifact the system also happens to have — it is the same object, captured at its last frame. That single decision turned out to have real teeth. It forces determinism (the same drawing must be reproducible, not regenerated). It forces every place the drawing is shown — live to the sitter, saved to storage, exported as a video, exported as vector art — to agree on the exact same geometry, because if they didn’t, “what you watched” and “what you got” would quietly diverge and the whole premise would collapse back into decoration. And it makes the artifact itself replayable: because the drawing is a recorded sequence of moves rather than a rendered picture, watching it again later is not a new render, it’s a rerun of the exact same tape.
This paper is about the system that came out of that constraint, and about two things that came out of it almost by accident. The first is that once the drawing itself had to be described precisely enough to compile, style stopped being something the system merely produced and became something it could state — a short table of named numbers rather than a learned embedding nobody can read. The second is that once I was already committed to being honest about how the portrait was made, it became natural to be equally strict about every other claim the product makes, including claims made only to myself, inside the code, about what a given number or a given pass name actually does. I ended up encoding those claims as automated tests, so that a change to the product can’t silently make a true sentence false without the test suite noticing — and, as this paper’s own writing found, so that a false sentence gets caught even when it was never meant for a user at all.
Concretely, this paper reports four things. First, an architecture in which a generative model is deliberately confined to producing a private reference study, and a separate, deterministic algorithm is responsible for everything the sitter actually sees drawn — meaning the “art” the audience experiences is fully specified by ordinary, inspectable code, not by a model’s black box. Second, the tracing algorithm itself, in enough mechanical and mathematical detail to reproduce: how a raster image becomes an ordered set of strokes with believable weight, sequence, and timing, using a pipeline of fairly classical computer-vision techniques rather than a learned stroke-generation model. Third, an interpretable style vector: the entire artistic difference between the studio’s eight hands, at the level the tracer can see, is a short table of named numbers, which means a reader can predict what changing one of them will do — and I report a case where publishing that table exposed two hands that are nearly identical at that level even though their names and their media suggest otherwise. Fourth, honesty by design as a general method for keeping a product’s claims about itself true over time, extended in this paper past prose and into the arithmetic underneath the drawing itself.
2Related work
The idea of turning a picture into strokes is old and well studied under the name of non-photorealistic rendering. Haeberli’s “Paint by Numbers” and Hertzmann’s painterly-rendering work from the 1990s both convert a source image into a set of placed brush strokes to imitate a painted look; later work extended this to coherent line extraction and directional hatching guided by local image structure. What almost all of that literature optimizes for is the final image looking painterly. It generally does not treat stroke order as a first-class output, because nothing downstream needed to replay the strokes in sequence — the goal was a static picture, and the strokes were an intermediate representation on the way to one. My system inherits the low-level techniques from this tradition (a distance transform for stroke width, tone reproduced as directional hatching) but treats the order strokes are laid down in as the actual deliverable, because a person is watching it happen in real time and later re-watching it. Winkenbach and Salesin’s pen-and-ink illustration work is the closer ancestor of Section 3.3’s parameter table specifically: their system also exposes stroke behavior as a small set of named, tunable values rather than the emergent behavior of a learned model — though theirs shapes a single rendering style, where mine uses the same mechanism to differentiate eight.
A more recent and more direct point of comparison is work that generates stroke sequences from a learned model directly — Sketch-RNN, and later CLIPasso and similar systems that optimize a small set of vector strokes against a semantic objective. These are elegant and, in a real sense, closer to “the model is actually drawing.” I chose a different tradeoff deliberately: by keeping the raster-to-stroke step as a classical, deterministic algorithm rather than a learned one, I get exact reproducibility, predictable performance, and a result I can fully explain and audit, at the cost of the stroke placement being derived from a generated image rather than directly produced by a drawing-native model. I think that’s the right tradeoff for a product that has to make a truth claim about its own process to a paying customer; I’d make a different tradeoff for a research system whose goal was more expressive or stylistically novel stroke generation.
On the honesty side, the closest formal comparison is content provenance standards like C2PA, which cryptographically sign metadata about how an image was produced. That is a valuable, complementary idea, and this system does not implement it — nothing here is a signed attestation. What I built is much narrower and much easier: an internal discipline for making sure a product’s own marketing copy, UI text, and internal parameters cannot say something the code doesn’t actually do, checked automatically on every change. It’s closer in spirit to contract testing than to cryptographic provenance, applied to a wider set of claims — prose shown to a person, a name displayed mid-performance, and, as Section 5 reports, arithmetic that never surfaces to a user at all but is still a claim the code makes about itself.
3System design
The system is a three-stage pipeline: design, compile, perform.
3.1Design
A sitter is observed by webcam for a short window — long enough to get a decent frame, short enough to still feel like a moment rather than a shoot. The best candidate frame, by simple, disclosed image-quality measures — a fixed weighted sum of brightness, contrast, sharpness, subject placement and framing, with nothing in it about the person’s expression or identity beyond “is there a clear face here” — is sent once to a hosted image-generation model along with a fixed, versioned prompt specific to the chosen artistic hand. Inter-frame stillness is measured as well, but it is not a term in that score: it decides when an automatic capture is allowed to fire, not which frame ranks highest. What comes back is not shown to the sitter. I call it a study, borrowing the painter’s term for a private working sketch: its only job is to be something a downstream tracer can recover clean, distinct marks from — no gradients, no photographic shading, no dark background, dark line on light paper. The prompt is engineered entirely around that constraint. It is, in an important sense, not a request for a picture to look at; it’s a request for a picture that can be traced.
3.2Compile
This is the part of the system that does the real work, and the part I’ll spend the most space on, because it’s the part most people assume must be a learned model and isn’t. The compiler takes the study image and produces an ordered document of individual strokes — each with a path, a width, an opacity, an instrument, and a position in a fixed sequence. It works like this.
Normalize. Rather than compare every pixel in the study to one fixed brightness cutoff, the compiler compares each pixel to a locally-estimated background. A generative model’s raster, like a photographed sheet of paper, is not evenly lit; a single global cutoff would misread a light corner as blank paper and a shadowed one as solid ink. The estimate is taken per tile: the sheet is divided into 48-pixel squares, a high percentile of each tile’s luminance is taken as that tile’s paper level, and the surrounding tile estimates are bilinearly interpolated so the background varies smoothly across the sheet rather than in blocks.
That much is standard, and on its own it is unsafe. A tile is only a fair estimate of the paper when most of the tile is paper. Over a passage that is uniformly dark — a deep complexion, a black garment — the tile’s own percentile sits inside that passage, so the passage barely departs from what the compiler believes the paper to be, and it is discarded as though it were blank sheet. The local estimate is therefore clamped: it may not depart further than a fixed allowance from the sheet’s own global paper level, measured once over the whole image. A corner vignette drifts a little way from the global paper and is still absorbed, which is what the tiles were added for; a face drifts a long way and is now measured rather than explained away. Section 6 reports what this cost before the clamp was there.
The departure from that background is then divided by the absolute range from the sheet’s paper level to true black, which gives every pixel a value between zero and one: how much ink it is owed. The denominator is deliberately absolute rather than the darkest pixel that happens to be present. Normalized against the observed maximum, a dark face and a fair one both land near the top of their own sheet’s range and arrive at the same weight — the same erasure the unclamped tile estimate performs, one stage further down.
Split. Two questions decide what kind of mark a pixel belongs to, and the compiler asks them in order. The first is whether the pixel is part of a line at all. A drawn line is thin by definition — it is dark where the few pixels either side of it are not — so the compiler measures how far each pixel stands proud of the mean of its own neighborhood, and takes as line work the strongest such ridges, in whatever share of the sheet that hand traces (the line share of Table 1). Selecting line work by darkness rank alone cannot do this: on a study that models the whole figure in tone, the darkest ninth of the sheet is mostly broad tone, so real contour lines are displaced into the mass and the sheet loses its drawing.
The second question is density, and it decides what happens to everything else. Thresholding a soft passage does not give a region with a boundary; it gives a fractal band of pinholes and spurs a pixel or two wide, and thinning a band is not a contour — it is a scatter of short marks. The mask is therefore closed first (a dilation followed by an erosion), which turns the band back into a curve, and the local density of the closed mask — the mean over an 11×11 window, computed as a pair of one-dimensional running sums rather than a full two-dimensional convolution, since the compiler runs once per portrait and has no room for a slow measurement — is compared against a per-style constant, the density gate (Table 1). Above the gate the pixel belongs to fused tone: a mass too dense to be a single traceable line, more like a shaded area.
The split carries one deliberate exception, and it is load-bearing. The deep interior of a mass is surrendered to the hatcher, because a skeleton says nothing there that a hatch does not say better — but the mass’s own boundary, taken as exactly one pixel wide (the mass minus its own erosion), is traced even where that edge sits under the tracing bar. Without that rule a mass sitting right at the gate loses its entire traceable edge: the interior correctly reads as solid tone, and the boundary a viewer’s eye actually reads as a line disappears along with it. I found that gap by writing this section down precisely enough to state as a rule rather than a paragraph, which is as good an argument for describing a pipeline this carefully as I have.
Trace the line work. Line regions are reduced to a one-pixel-wide skeleton with Zhang and Suen’s 1984 parallel thinning algorithm: in alternating passes, a foreground pixel is deleted only if it has between two and six foreground neighbors, exactly one background-to-foreground transition walking around its neighbor ring, and one of two further conditions that swap between the two passes — a construction that erodes a thick stroke from opposite sides on alternating iterations until only its centerline survives, without ever severing a stroke that should stay connected. The skeleton alone has thrown away the stroke’s thickness, so a distance transform is computed alongside it: at every traced point, the distance to the nearest true background pixel becomes that point’s recovered radius, which is what gives the final stroke a believable, varying weight — thick where the original mark was bold, thin where it thinned to a whisper — rather than a uniform wire. The resulting polylines are then simplified with the Ramer–Douglas–Peucker algorithm: recursively find the point on a segment farthest from the straight line between its endpoints; keep it and split the segment there if that distance exceeds a tolerance, or discard it and collapse the segment to a line if it doesn’t. This keeps the geometry honest — every kept point was worth keeping — while cutting the point count to something a renderer can draw smoothly and a replay file doesn’t need to be enormous to store.
Synthesize the tone. Fused regions, which by definition have no single line to recover, are not traced at all — they are re-drawn as hatching. At each point in a mass, a structure tensor built from the local image gradient gives the dominant local orientation of the underlying form, and the hatch strokes there follow that orientation rather than running in one fixed direction regardless of what’s under them — a jaw reads differently from a cheek because the hatching actually follows the jaw and the cheek. Darker regions get denser hatching, and past per-style thresholds — a first crossing point, then a second — a second and, past the higher one, a third crossing layer of hatching is laid over the first: literal cross-hatching. Both thresholds are per-style constants (Table 1), and what they mean depends on a third: each hand carries a build temperament, and the compiler runs one of two different hatching programs accordingly. A hand that closes lays a single direction of ranks and tightens it as the plane darkens — the stepped ink of a nise-e sheet, yukrimun, pardakht — and crosses only once, in the deepest passage, so that paper still shows between the lines at every step. A hand that crosses lays one even rank and adds further directions over it, which is what Ashcan crayon reportage and kolam-rhythm hatching do. For a closing hand the two thresholds are the points at which its ranks tighten; for a crossing hand they are the points at which the second and third directions arrive. This is why the same pair of numbers reads so differently down the table, and why the two temperaments cannot be compared as though they were one scale.
Both thresholds are compared against the local value at each pixel, not against the average of the mass that pixel belongs to. Gated on the mass average, every pixel of a region receives the same number of crossings, and a face whose study carried a soft gradient comes back as one uniform crosshatch mat, edge to edge, at whatever weight the average happened to land on. Value is what hatching is for; it has to vary within the mass or the drawing has no light in it.
Stroke opacity follows a fixed rule with one deliberate asymmetry: the opacity of a mark is the smaller of 1 and a per-style opacity floor plus 1.1 times the mark’s measured darkness — but the floor applies at full strength only to traced line work, and at four tenths of its value to synthesized tone. The floor is a promise about the line: a contour is committed to, so it is never laid tentatively. Tone makes the opposite promise, because a rank in a lit plane has to read as light. Applying the same floor to both put every rank on the sheet at four tenths strength or more, which erased the study’s whole light end — a face carrying a soft gradient came back at one even weight, with hard white holes where the highlights fell under the floor.
Order. Every stroke — traced or synthesized — is assigned to one of seven named passes and a position within a single seeded sequence: nothing about the order is left to the renderer’s discretion at playback time. The seed comes from the study image itself, computed with the well-known FNV-1a hash run over a sample of the study’s own pixel bytes: start an accumulator at a fixed constant, and for every sampled byte, exclusive-or it into the accumulator and multiply the result by a fixed prime, wrapping to 32 bits, until the sampled bytes run out; the final accumulator value is the seed. Nothing about the sitter’s identity enters that computation — only the bytes of the image the model already returned — so the same study always produces the same seed, and the same seed always produces the same sequence.
3.3Eight hands, one tracer
The studio ships eight distinct artistic hands, each attached to a city the way a real atelier might name a resident: Mira in San Francisco works in a sparse, close-valued graphite economy with a fog-grey wash; Sol in New York throws down fast sanguine crayon reportage on warm ivory, no wash at all; Odile in Paris keeps a silverpoint-fine grey line with rose chalk used only where warmth belongs; Aya in Tokyo draws a fine sumi contour and builds every value from stepped ink ranks; Yeon in Seoul works a hair-fine sepia line modeled entirely by many short, separate parallel strokes with paper breathing between them; Halia in Singapore pairs a Nanyang ink-brush line with shophouse-pastel watercolour; Kavi in Chennai draws in kaavi red ochre with kolam-rhythm hatching; Ganga in Jaipur works a wiry contour modeled by ranks laid like a ploughed field. Every hand shares the same tracer described above. What makes them eight different hands rather than one hand in eight colors is a request written in language, sent once to the image model, and a short table of numbers read by the tracer afterward. The table is the part I can publish in full — unlike the language, which is the studio’s own creative work, the table is just the compiler’s input, and publishing it is itself a small honesty exercise, for reasons the next two paragraphs get to.
| Hand (city) | ink (RGB) | stroke width scale | opacity floor | density gate | hatch spacing | build | first cross | second cross | line share | wash |
|---|---|---|---|---|---|---|---|---|---|---|
| Mira (San Francisco) | 91, 97, 103 | 1.05 | 0.40 | 0.50 | 3.2 | cross | 0.55 | 0.80 | — | yes |
| Sol (New York) | 182, 92, 67 | 1.20 | 0.55 | 0.42 | 3.2 | cross | 0.42 | 0.62 | — | no |
| Odile (Paris) | 124, 130, 136 | 0.85 | 0.55 | 0.40 | 2.8 | cross | 0.50 | 0.74 | — | yes |
| Aya (Tokyo) | 61, 56, 49 | 0.85 | 0.46 | 0.72 | 3.0 | close | 0.30 | 0.52 | 0.03 | yes |
| Yeon (Seoul) | 75, 64, 52 | 0.95 | 0.55 | 0.82 | 4.2 | close | 0.32 | 0.50 | 0.03 | yes |
| Halia (Singapore) | 58, 56, 53 | 1.10 | 0.50 | 0.50 | 3.6 | cross | 0.60 | 0.84 | — | yes |
| Kavi (Chennai) | 156, 74, 51 | 1.10 | 0.50 | 0.50 | 3.0 | cross | 0.50 | 0.76 | — | yes |
| Ganga (Jaipur) | 91, 70, 54 | 1.05 | 0.78 | 0.85 | 4.4 | close | 0.30 | 0.48 | 0.03 | yes |
Line share is the fraction of the sheet a hand traces as line work rather than carrying as tone; left blank in the table, it defaults to 0.09. The three hands that lower it to 0.03 — Aya, Yeon, Ganga — are exactly the three whose media lay fine ranks close together: at the default share, the soft antialiased halo around each rank fuses into the next one and the whole passage traces as a single gravel-textured mass instead of separate lines; a smaller share keeps only each rank’s solid core, so the ranks survive as what they’re supposed to be.
Reading the rest of the table the same way: the five mass-driven hands — Mira, Sol, Odile, Halia, Kavi — sit at a low density gate (0.40–0.50), which sends more of the mask into synthesized tone and makes that tone heavy fast. The three line-driven hands — Aya, Yeon, Ganga — sit at a high gate (0.72–0.85), which keeps recovered line work out of the mass branch, and they are also exactly the three closing hands, which is not a coincidence: a medium that carries value by tightening one direction of ranks needs those ranks to survive as line work in the first place. Spacing does not split along that same line: Yeon and Ganga hatch wide (4.2, 4.4), but Aya — despite carrying the same high gate and the same closing build — hatches at 3.0, as tight as anything in the mass-driven group. The gate and the spacing are two different knobs, tuned separately per hand, and this is where the table says so rather than implying a pattern that isn’t there. Ganga carries the highest opacity floor of any hand, 0.78, so that even the faintest fragment of a traced line lands on the page as a committed stroke rather than fading toward invisible — wiry and even-weight, never hesitant.
The table is also where the studio’s honesty about itself gets uncomfortable, which is the point of publishing it rather than only the flattering medium descriptions above it — though not in the way an earlier draft of this paper claimed.
That draft reported that two of the eight hands, Halia and Kavi, were close to identical at the level the tracer can see: same width scale, same opacity floor, same density gate, same spacing, differing only by two hundredths in each crossing threshold and by their ink colour. One hand in two colours, I wrote, and I would rather say so than round them up to “distinct” because the studio gives them different names.
That paragraph was wrong, and the way it was wrong is worth more than the paragraph was. It had not been written from the code; it had been written from a transcription of the code into this table, and the transcription had drifted. Regenerating Table 1 from the shipped configuration rather than retyping it moved values in seven of the eight rows. Measured properly — across the twelve parameters the compiler actually consumes, ink colour excluded, since identity-but-for-colour is the claim under test — Halia and Kavi do remain the closest pair on the roster, and they agree on seven of those twelve. They differ in hatch spacing (3.6 against 3.0), in both crossing thresholds (0.10 and 0.08 apart, not two hundredths), in face-ink economy, and in whether they concentrate the head into a named pass at all. No pair on the roster today is near-identical: the next closest agree on six of twelve.
So the self-criticism does not survive, and something less comfortable takes its place. The original claim was an assertion about the system that the system did not support, in a paper whose subject is not making assertions a system does not support. It survived because Table 1 was a copy rather than a derivation — precisely the failure mode Section 5 describes for prose, which I had not thought to apply to my own numbers. Table 1 is now generated from the source it describes, and the web page carrying this paper is generated from this manuscript, with a check that fails when the two disagree. The class of error that produced that paragraph can no longer happen quietly.
| Hand | tool, early / middle / late | pass 1 | pass 2 | pass 3 | pass 4 | pass 5 | pass 6 | signature | total |
|---|---|---|---|---|---|---|---|---|---|
| Mira | graphite, stump, sable | 6.2 | 7.6 | 7.4 | 6.8 | 6.6 | 6.2 | 4.2 | 45.0 |
| Sol | graphite, chalk, stump | 5.4 | 5.0 | 5.6 | 5.4 | 5.0 | 5.2 | 3.8 | 35.4 |
| Odile | fine, chalk, sable | 5.6 | 7.2 | 7.6 | 5.2 | 5.6 | 6.6 | 4.2 | 42.0 |
| Aya | fine, sumi, sable | 6.2 | 7.6 | 6.6 | 6.4 | 6.8 | 4.8 | 4.2 | 42.6 |
| Yeon | charcoal, fine, sable | 6.2 | 6.4 | 6.8 | 8.2 | 6.6 | 6.4 | 3.8 | 44.4 |
| Halia | graphite, sumi, sable | 4.8 | 6.2 | 6.6 | 5.6 | 5.2 | 5.6 | 3.8 | 37.8 |
| Kavi | graphite, sumi, sable | 4.2 | 7.2 | 5.8 | 6.2 | 4.8 | 5.2 | 3.8 | 37.2 |
| Ganga | graphite, fine, sable | 5.8 | 6.6 | 7.6 | 7.4 | 6.2 | 6.4 | 4.8 | 44.8 |
Every hand performs in roughly the same window, thirty-five to forty-five seconds, regardless of how many individual strokes its version of a given face actually produced — a face with more traceable detail doesn’t take proportionally longer to watch, it takes the same time with more strokes packed into the same seven passes. Each pass also carries a name a sitter can read while it plays — “The long contour,” “Eyes, set once,” “A single hair’s line” — and that naming is where the table’s honesty properties get tested hardest, which Section 5 reports on directly.
3.4Perform
The compiled document — an ordered list, nothing more exotic than that — is what every downstream surface actually draws from. A canvas player executes it live, stroke by stroke, timed against the per-pass duration budget in Table 2, so the performance takes roughly the same amount of time no matter how many individual marks a particular face happened to produce. A server-side renderer produces the still image that gets saved, from the identical geometry, not a separate render pass with its own logic. A vector export and a short video export both walk the same document too. Because there is exactly one source of truth for the geometry, and every consumer reads from it rather than re-deriving it, “what the sitter watched” and “what got saved” are the same data by construction, not by careful synchronization.
4Determinism and replay
The property that makes all of this hold together is that compilation is deterministic: the same study image and the same seed produce the same stroke document, every time, with no randomness anywhere in the tracing or ordering logic. That single property is what lets me make a much stronger claim than “the drawing looks similar to what you saw” — I can claim it’s the same drawing, because it’s literally the same recorded sequence of moves, replayed.
This has a practical consequence I didn’t fully appreciate until it fell out for free: replaying a finished portrait later — for the sitter to watch again, or for a visitor to a public page — costs nothing beyond serving a small file. There is no second call to the image model, no re-generation, no risk that a re-run produces a subtly different result. The stored artifact for a completed sitting is the compiled stroke document itself (in a compact packed form — coordinates, widths, and colors encoded efficiently rather than as verbose text), and “replaying” it is exactly the same code path as the first performance, just fed a document that already exists instead of one just produced.
I want to be precise about the boundary of this claim, because it would be easy to overstate it. Determinism holds from the study image onward. It does not hold from the sitter onward: the upstream generative model is not seeded or otherwise made reproducible, so if a sitter’s photo were resubmitted, there is no guarantee the model would produce the same study twice. What’s guaranteed is narrower and, I think, still the important part: once a study exists, everything after it — the drawing, the order, the timing, every render of it forever after — is fixed and repeatable.
5Honesty by design
Somewhere in building this, I noticed that the discipline I was applying to the drawing claim — don’t say something is happening unless it mechanically is — was a discipline I wanted applied to every other claim the product made, and that I had no way to guarantee that discipline would survive future changes made under time pressure, by me or by anyone else working on it. The fix I landed on is simple to describe and, I think, underused: write the specific sentence you’re allowed to say as an automated assertion against the actual product surface, so that changing the product without updating the sentence — or writing a sentence the product can’t back up — breaks a test.
A concrete example: at one point I wanted the interface to reassure people that the system does not claim to know anything about them from their face beyond whether there’s a usable photo of one — no inferred mood, no age guess, nothing psychological. Rather than trust that this stays true as a matter of code review discipline, there’s a test that reads the actual user-facing copy and fails if certain classes of claim ever appear in it. Another checks that a number displayed anywhere in the product as a count is actually computed as a count, rather than derived from something that merely correlates with one — an easy mistake to make by accident, and a genuinely misleading one if it ships.
The same discipline turned out to apply to the compiler’s own internal claims, not only to the copy a sitter reads, and two examples from writing this paper are worth reporting plainly rather than tidying away. The first: several of the eight hands include a pass named for the face itself — “The exact face,” “Eyes, set once,” “A single hair’s line” — and the name implies a mechanism: that the compiler concentrates strokes from the measured region of the face into that particular pass, so the name is literally true of what a sitter watches while it plays. Writing Table 2 forced me to check that claim against the code doing the routing, and the mechanism wasn’t there — the pass received its strokes from the same length-and-darkness ranking as every other pass; nothing about it was actually face-aware. I could have quietly built the missing mechanism and made the claim true after the fact. I corrected the claim instead, because a pass name is a promise made to someone watching in real time as it plays, and a promise repaired retroactively by changing the code around it is not the same as a promise that was true when it was made.
The sequel is the part I did not expect to be reporting. Some weeks later the mechanism was built deliberately, and the names went back up on that basis. The compiler now finds where the head is on the sheet — from the shape of the ink itself, not from a face detector: ink begins at the crown, holds roughly the head’s width down the face, then widens sharply at the shoulder line, and the row half again as wide as the head is where the head ends. Strokes from that region are routed into the named pass, for the six hands that carry one. The two that work by mass rather than by feature carry no such pass and were left alone.
I want to be precise about why this is not the thing I refused to do two paragraphs earlier, because from outside the two look identical and from inside they are opposites. What I refused was making the claim true after the fact while leaving the claim standing, so that no one would ever learn it had been false. What happened instead is that the claim came down, the gap stayed visible for as long as the mechanism was missing, and the name went back up only once the code could carry it. The difference is not in where the system ends up — both routes end with working code and a true sentence — but in whether there was ever a window in which the product said something it could not back up, and in whether anyone outside could have known. Honesty by design is a claim about that window, not about the endpoint. A method that could not tell those two cases apart would not be worth much.
The second is smaller and sharper. Describing the density-gate split in Section 3.2 precisely enough to state as a rule, rather than a paragraph about it, is what surfaced that a mass sitting close to the gate could lose its traceable edge entirely — and, separately, that a decimation guard controlling how much of a hatched pass survives had been written expecting mathematical modulo, when JavaScript’s % operator returns a remainder, and the two disagree on negative inputs. A sweep whose coordinate could run negative had been keeping every point on the negative side of the sweep instead of roughly one in three, which had quietly been making some hatching up to three times denser than it was supposed to be. Neither defect threw an error; both only ever showed up as a subtle shift in a compiled document’s own statistics, which is exactly the kind of thing a test suite checking that the code runs, rather than checking what it produces, will never catch. Both are fixed now, with regression tests that fail if the fix is reverted. I’m reporting them here rather than only in a private changelog because a system whose honesty tests catch every claim in its marketing copy but missed the arithmetic underneath its own tone rendering is a system I should be more careful about, not less — and because an honesty method that only ever gets pointed at prose isn’t really a general method yet.
I don’t think any of this is exotic. It’s closer to ordinary contract testing than to anything novel in the honesty-technology sense — the trick is applying that discipline everywhere a claim is made: in prose shown to a user, in a name displayed mid-performance, or in the arithmetic underneath the performance itself, not only in the parts that are easiest to point a test at. I’m including it here because I suspect most products that make a process claim (“the AI is doing X live”) have no equivalent enforcement at all, and would benefit from one regardless of whether the underlying process claim is as literal as mine.
6Evaluation
I want to separate, cleanly, what I have actually measured from what I believe but haven’t verified, because conflating the two is exactly the kind of overclaiming this whole project is supposed to avoid making.
What is measured. Determinism of the compile step is a property of the code, not a statistic — given the same input and seed, the algorithm is not stochastic, so this is closer to a proof than an experiment, and it’s the property everything else in this paper leans on. Cross-consumer geometric agreement (the live player, the saved image, the vector export, and the video export all producing the same visible result from the same document) is checked as part of the ordinary build process. Reasonable performance ceilings exist and are enforced — a document cannot exceed a bounded number of strokes or path points — but I have not run a systematic latency or frame-rate benchmark across a real spread of devices, and I want to say that plainly rather than imply a number I don’t have.
What is proposed, not done. A perceptual study asking real viewers to judge whether a given output reads as “drawn by hand” versus “printed” would be the most valuable single experiment I don’t yet have, ideally as an ablation across the compiler’s main tunable parameters — the ones tabulated in Section 3.3, principally hatch spacing, the crossing thresholds, and the density gate that decides line versus mass — rather than a single yes/no verdict on the shipped defaults. A second, more targeted study could ask viewers to match a rendered portrait to its named hand, or to tell two hands apart — most usefully starting with the closest pair the parameter table admits, Halia and Kavi, which agree on seven of the twelve parameters the tracer reads. A controlled comparison of replay-from-recording against a hypothetical re-generate-from-scratch approach would put a real number on the cost and fidelity argument I make in Section 4 rather than leaving it as an architectural claim. And the fairness evaluation described below, which is a standing repeatable test today, still runs against a fixed pair of generated sitters rather than a real demographic sample set.
That last point deserves its own section, and it is the place where earlier drafts of this paper were most wrong about my own system.
An early bias evaluation found a real problem: a line-based style, applied without adjustment, systematically under-represented deeper skin tones — bare paper was standing in for skin where a lighter complexion would have gotten real ink. I changed the prompt language for the affected hands, and then closed the report on an assumption I never checked: that because the compile step samples darkness from the study, a fix made upstream would propagate to the finished sheet on its own. It did not, and nobody knew for two weeks.
A second audit found the cause in the compiler, and specifically in the normalization step of Section 3.2 — in the version of it that had no clamp. The background estimate is a percentile taken inside a small tile, and a tile is only a fair estimate of the paper when most of the tile is paper. Over a uniformly dark passage the percentile sits inside that passage, the passage barely departs from the estimated paper, and it is discarded as blank sheet. The drawing was being erased in proportion to how dark the sitter was. Measured across all eight hands with no exception, a deep-complexioned sitter’s face survived the trace at roughly half the rate a fair-complexioned sitter’s did, and five of the eight finished sheets rendered the deep sitter lighter than the fair one.
Two changes were needed, and neither was sufficient alone: clamping the local estimate so it cannot depart more than a fixed allowance from the sheet’s own global paper level, and normalizing a pixel’s owed ink against the absolute range from that paper level to true black rather than against the darkest pixel present. A third change, later, moved the compiler’s ink-economy budget onto rank spacing and away from opacity, on the same principle: coverage is the artist’s business and value is the sitter’s, and a loop that enforces one must not be allowed to spend the other.
What used to be a one-time pass is now a standing test. Eight hands are compiled against a fixed pair of eval sitters and each is held to its own separation floor, set at three quarters of what that hand measures today, so a hand that starts sliding fails the suite rather than degrading quietly. The measurement is taken from the rendered sheet’s own pixels — the way a sitter looking at their portrait would — rather than from the compiler’s internals, and it is run as the median of several fresh studies at production resolution rather than against frozen fixtures. That distinction earned itself: the frozen-fixture version of this test was passing while three hands were failing on fresh work. All eight clear their floors today.
What I still do not have is that evaluation on real faces. The eval sitters are generated, which means they come from the same family of models as the design stage, and the honest limit of the claim is narrow: the compiler no longer erases a complexion it is given. It is not a claim that the studio draws a real distribution of people fairly. Closing that gap needs consented photographs of real sitters scored against a published skin-tone scale, and it is the single most valuable thing missing from this paper.
7Limitations
Everything in Section 6’s second half is a limitation, and I won’t repeat it. Beyond that: this is a single system built by one person, without an external audit beyond ordinary self-review, so I have less confidence in claims about the system’s edge cases than I would if a second team had tried to break it. The upstream image model is a hosted third-party service; if it changes behavior, degrades, or is swapped for a different model — which I expect to happen, since the entire architecture is built so that the specific model behind the design stage is a replaceable implementation detail, not a load-bearing assumption — the character of the resulting studies could shift in ways the compiler’s current tuning doesn’t anticipate. While the compile step is deterministic, that only buys back reproducibility for a portrait that has already been made; it says nothing about whether the first draft a given sitter gets is a good one, which is still, at the margin, a matter of luck in what the upstream model happened to produce. And none of the eight parameter vectors in Table 1 were derived from a search, an objective function, or a perceptual study — they were set by hand and adjusted by eye, and the reasoning I give for reading the table in Section 3.3 is falsifiable but has not actually been tested against a viewer who doesn’t already know which hand is which. What I can report is that the values are distinct, that each is consumed exactly where Section 3 says it is, and that the pipeline is deterministic given a study and a seed. What I can’t report is that any one of the eight is right, or that a stranger could reliably tell two of them apart beyond the pair the table already admits are close.
8Conclusion
The premise of this system is a small, checkable idea: if you’re going to tell someone a machine is drawing their portrait, the drawing they watch should be the same object as the portrait they keep, not a performance layered over a result that already existed. Getting that right forced determinism, forced every surface that shows the drawing to agree on the same geometry, and turned a marketing promise into an architectural property that can be verified rather than merely asserted. It also produced a compiler whose entire behavior is described by ordinary code and a short table of named numbers rather than a learned style embedding — style you can read, not just observe — and applying the same standard to the rest of the product’s claims, including the arithmetic inside the compiler itself, turned into a second, more general technique: encode what you’re allowed to say as a test, so it can’t quietly stop being true. I think all of this is more useful outside this specific portrait studio than inside it, and I’ve tried to describe it here precisely enough that someone building something else entirely could take any part of it.
References
- Haeberli, P. “Paint By Numbers: Abstract Image Representations.” Proc. SIGGRAPH, 1990.
- Hertzmann, A. “Painterly Rendering with Curved Brush Strokes of Multiple Sizes.” Proc. SIGGRAPH, 1998.
- Hertzmann, A., Jacobs, C. E., Oliver, N., Curless, B., and Salesin, D. H. “Image Analogies.” Proc. SIGGRAPH, 2001.
- Winkenbach, G., and Salesin, D. H. “Computer-Generated Pen-and-Ink Illustration.” Proc. SIGGRAPH, 1994.
- Zhang, T. Y., and Suen, C. Y. “A Fast Parallel Algorithm for Thinning Digital Patterns.” Communications of the ACM 27(3), 1984.
- Ramer, U. “An Iterative Procedure for the Polygonal Approximation of Plane Curves.” Computer Graphics and Image Processing 1(3), 1972.
- Douglas, D. H., and Peucker, T. K. “Algorithms for the Reduction of the Number of Points Required to Represent a Digitized Line or its Caricature.” The Canadian Cartographer 10(2), 1973.
- Ha, D., and Eck, D. “A Neural Representation of Sketch Drawings.” ICLR, 2018.
- Vinker, Y., Pajouheshgar, E., Bo, J. Y., Bachmann, R. C., Bermano, A. H., Cohen-Or, D., Zamir, A., and Shamir, A. “CLIPasso: Semantically-Aware Object Sketching.” ACM Transactions on Graphics 41(4), 2022.
- Coalition for Content Provenance and Authenticity. C2PA Technical Specification, version 2.2, 2025.