The guarantee that made it correct is what made it slow
Image Horse is a photo editor that runs entirely in your browser. No server touches your pixels. The editing engine is Rust compiled to WebAssembly, and until this week every call into it happened on the main thread.
That last part is the problem. A heavy filter blocks everything — nothing paints, no click registers, the spinner doesn't even spin. The fix is to move the engine into a Web Worker.
It took eighteen days. It shipped on Thursday morning. By that afternoon the brush was seventeen seconds behind the cursor, edits were vanishing on refresh, and every single cause traced back to a property I had deliberately built, tested, and proved.
Here's the honest version.
The number that started it was wrong
The freeze I was fixing was 470 milliseconds — a 12-megapixel sharpen, measured on a standalone spike page. Half a second of a dead browser. That number is in the architecture decision record. It's the whole justification.
Except the app can't hold a 12-megapixel document. Every import runs through a function that downscales to 2048 pixels on the long edge. The largest document the editor will ever have is about 4.3 megapixels.
So the benchmark was taken on a document no user can produce. When I finally measured it properly at the real ceiling, the freeze was about 180 milliseconds.
Still worth removing. Not half a second.
I found this on day three, while checking something unrelated. Nobody had lied; the spike page predated the downscale, and the number just kept getting quoted. It had been sitting at the top of the document that justified the entire project.
Check the number that motivated the work. Especially if it's old, and especially if everyone agrees on it.
Counting the work was harder than doing it
Every synchronous call into the engine has to become asynchronous. So: how many are there?
I wrote a script to count them. Over four days it reported:
| Count | Why it moved |
|---|---|
| 121 | the original figure |
| 162 | the script only recognised three literal receiver names, so const t = toolRef.current; t.width() was invisible — 93 of 290 sites |
| 166 | it was counting engine calls written inside comments |
| 171 | its "is this value used?" test was single-line, so a call formatted as an argument on its own line read as a bare statement |
| 168 | actual work: six call sites became three |
Four of those five moves were the measurement getting more honest. Only one was code getting better.
There were more after that. Optional chaining, toolRef.current?.undo(), didn't match the pattern. A structural test I wrote to guard against that matched on useCanvasIdentity( while the real call site read useCanvasIdentity<HTMLCanvasElement>( — a type argument between the name and the paren, and my guard silently passed. And one file contained a literal NUL byte, which makes grep treat it as binary and skip it silently: a confident zero for a string that was demonstrably on line 44.
Nine detection failures. Every one found by cross-checking against a second, differently-broken method.
A scan of your own codebase is a measurement, and measurements need controls.
The plan was wrong, and the code said so
With the counting settled, the plan was mechanical: convert file by file, biggest first. The biggest was the save path — eighteen call sites in one function.
I opened it and found a comment I'd written months earlier:
Everything above reads the engine, and there is not a single
awaitin it. That is load-bearing, not incidental. If anyone ever adds anawaitabove this line, detaching stops being safe and this comment is the reason why.
The save path captures the document — canvas, undo stack, annotations, layers — and hands the bytes off. Because the capture is synchronous, the upload can be detached and run in the background.
Convert those eighteen reads to eighteen awaits and the capture stops being atomic. A photo switch completes halfway through. The second half of the archive describes the incoming photo, uploaded under the outgoing photo's key.
Silent cloud corruption, from following the plan exactly.
The fix wasn't to guard the reads. It was to delete the problem: one engine call that returns the whole capture at once. Several of those now exist, and they account for most of the conversions — one call taking out a dozen sites at a time.
When a plan and a code comment disagree, the comment is usually the one that had to survive production.
The best bugs had nothing to do with workers
Asking "could these two reads disagree?" of every call site found things that were already broken, for real users, with no worker anywhere:
Three composites to answer one question. The "export without the canvas background" path called three separate getters for pixels, width and height. Each internally composited the entire document and threw away two-thirds of the result. 69ms for the three-call form, 20ms for one call that keeps all three values. That had been shipping.
Dimensions in the wrong file's metadata. On the export path, width and height outlive the encode and get stamped into the exported file's EXIF. Read separately, one state's dimensions could end up recorded inside another state's pixels.
Share metadata I'd dismissed as a caption. I'd parked one item as "two integers on a button label, low priority." Tracing it showed those integers go into a database table backing public share links. A caption that's briefly wrong self-corrects. A torn pair stored against a share is wrong for the life of the link.
None of these needed a worker. They needed someone to ask a specific question about every place the code talks to the engine.
Then I actually ran it
Eighteen releases shipped with the worker code live and switched off. Every gate green: type checking, linting, 512 unit tests, a contract test suite built specifically for this migration.
The first time anyone drove it end-to-end in a real browser, it found two defects in one sitting.
Opening a second photo left the canvas blank. You can hand a canvas to a background thread exactly once — after that, that canvas can never be handed over again, ever. The app was building a fresh worker per photo, which destroyed the thread holding the canvas and left the replacement with nothing to draw on. Photo loaded, thumbnail correct, dimensions correct, picture absent, console empty. Counted during a reload: the discarded worker had the canvas and zero work; the live one had 339 pieces of work and no canvas.
There was a unit test asserting that the disposal was correct. It had sound-sounding reasoning about memory. It was green, and it was wrong.
The readouts stopped counting. Those atomic captures I was so pleased with return a Rust struct across the thread boundary. Structured clone keeps a value's own data properties and drops its prototype — and the struct's fields are prototype accessors. What arrived was { __wbg_ptr: 1114112 }: every field undefined.
That one is worth sitting with. The technique that made the migration tractable is the technique that broke it. Every capture I built moved call sites off the list my audit was counting, and onto a list nobody was keeping. The counter went down while the real problem grew, and reached its floor in a state that couldn't work.
Then a third: I flipped the kill switch to check it worked, and the whole app went white. The switch swapped the canvas element back but not the engine, so the render path called a method that now returned a Promise, threw inside render, and React unmounted the entire tree. The escape hatch was the crash.
The flip
Default on, Thursday morning. Total blocking time during a heavy operation went from 129–137ms to zero. Long tasks: one, to none. The freeze didn't move somewhere else. It went away.
Hours later, the brush was lagging.
Only when signed in — at first. Signed-in accounts autosave to the cloud 2.5 seconds after an edit, and that autosave asks the engine for the whole document: 29.5 MB on a photo with strokes. On the worker, that request goes through a queue that answers strictly in order. Brush strokes issued behind it wait for all of it. A call that normally takes a third of a millisecond measured at 407ms when stuck behind an autosave. And 2.5 seconds after your last stroke is precisely when your next stroke begins — nearly every stroke collided.
That queue answers in order because I made it answer in order, on purpose, and I have a proof that it does. The engine's undo history records operations in arrival order — no sequence numbers, no timestamps, just the order they show up. So if requests could overtake each other, the undo stack would silently stop reproducing. The whole architecture rests on one channel, strictly FIFO. I tested that claim before shipping: sixteen mutations fired with no waiting between them, and the resulting log was byte-identical to the single-threaded version.
That proof is correct. It's also the bug. One queue means ordering is guaranteed. One queue also means a 29.5 MB read sits in front of your brush stroke.
The fix is scheduling, not machinery: autosave waits until you lift the pointer. The data is no fresher mid-stroke, so waiting costs nothing.
I shipped it, verified it, and called it fixed.
The fix helped the wrong bottleneck
The brush was still seconds behind. Then the report got sharper: slow logged out too. That one sentence killed the signed-in theory, and everything I'd measured with it.
Here's what was actually happening, and why nothing I built could see it.
The old main-thread brush had flow control nobody designed. Painting blocked the page, so the browser could only deliver pointer positions as fast as the engine absorbed them. The blocking — the very thing this migration existed to remove — was the rate limiter.
The worker brush yields instead of blocking. So every position a real mouse produces, 120 to 420 per second, became a queued engine call plus a flush request. Arrival outran service, and the debt compounded: a 1.4-second stroke banked 17 seconds of queue. The ink eventually caught up, drawing a perfect stroke, seventeen seconds after your hand finished it.
And every instrument said everything was fine. Frame timing: a locked 60fps, because the main thread genuinely never blocked. The failure mode was latency, not jank — the two things a frame-gap monitor cannot tell apart, because it only ever sees one of them.
The test automation was blind for a different reason: Playwright moves the mouse a polite 25 times a second. Every verification run all week had been driving the pipeline at a twentieth of hardware rate. The gates weren't wrong; they were asking at the wrong frequency.
The fix made the accidental backpressure explicit: at most one paint call in flight, the newest cursor position replaces any unsent one, one screen flush per displayed frame. No stroke detail is lost — the engine draws the connecting segment between landed points, which is exactly the input the blocking version always fed it, for exactly the same reason. Same stroke after: 0.26 seconds behind at release, level with the cursor during.
Removing a block can remove a guarantee you didn't know the block was providing. The synchronous version was slow and self-pacing. I proved the replacement kept the ordering. I never asked what else the blocking had been doing.
Then edits started vanishing
Next report: clone stamp, emoji and Magic Eraser edits didn't survive a refresh. Brush strokes did. Signed in or out, either engine.
The pattern in that list is the diagnosis. The survivors are recorded in the operation log and replay on restore. The casualties aren't — they exist only as pixels, and pixels are saved by an autosave that waits 2.5 seconds after your last edit. Refresh inside that window and the work is gone.
This wasn't new. What was new is that the safety net under it had died — silently, on flip day. The old engine had a last-chance save at page close: synchronous, so it always completed before teardown. Behind the worker, that save is an await across a message port during page teardown. The reply arrives after the page no longer exists. The continuation never runs. No error, no log line — the handler simply stops mattering, and the entire 2.5-second window became a data-loss window.
The immediate fix watches for the dangerous state: when a document holds edits the log doesn't cover, autosave fires at 300 milliseconds instead of 2500. And the first version of that still lost the stroke — it checked whether the log was recording, and a log can be recording happily while a stamp walks straight past it. What works is arithmetic: every committed edit grows the undo stack, but only recorded edits advance the log. Undo count greater than op count is proof an unrecorded edit exists, no enumeration of tools required. Coverage, not activity.
The real fix — recording those tools in the log so they replay like paint — is filed, with a number, for its own session.
A synchronous escape path doesn't survive going async, because "runs during teardown" was a property of being synchronous. Nothing warns you. The code still looks like it runs.
Where the bugs actually lived
Five ship-day defects, and not one was in the code the tests covered. Line them up and they're the same finding five times:
| Bug | Where the instruments weren't looking |
|---|---|
| autosave × brush collision | the only path with no automated coverage — needs a signed-in session |
| 17-second ink lag | real hardware event rates; automation drives at 25 Hz |
| 17-second ink lag, again | latency, when every monitor measures jank |
| vanishing edits | the gap between two save mechanisms, open only for 2.5 seconds |
| dead last-chance save | page teardown, where nothing can log that it failed |
The night before the flip I ran a coverage sweep — every export format, save and reload, undo twenty-two deep, layers, text, pen, batch mode. Zero failures. The last line of that report reads: cloud sign-in → save → restore: not reached — needs credentials. I recorded the gap accurately and flipped the default anyway, because everything else was green.
The bug was in the gap. It's always in the gap — that's what makes it the gap. And the other four weren't even in a gap I knew to write down: they were in dimensions my instruments didn't have. Event rate. Latency versus jank. Teardown timing.
Writing a gap down is not the same as closing it. And your instrument list is itself a claim about where bugs can be — one that nothing in the test suite checks.
Was it worth it?
The freeze is gone, measured rather than asserted: 130 milliseconds of blocking per heavy operation to zero. The brush runs at 60fps for the first time in the app's life — the old blocking path never got above ~27, and nobody knew, because it was slow in a way everyone was used to. Along the way: exports 3.45× faster, a batch-export data-loss bug fixed, eight read-modify-write races removed, and a data-loss window that predated the whole project finally closed.
But the pattern I keep coming back to isn't any single bug. It's that every property has two faces, and testing only ever looks at one.
FIFO means the undo log is byte-perfect. FIFO means a 29.5 MB read blocks your brush. Blocking means frozen pages. Blocking means the mouse can't outrun the engine. Synchronous means slow. Synchronous means it finishes before teardown.
I proved the first face of each of those, with real tests, some of them mutation-tested. The second face shipped to production, three times, in one day. Not because the proofs were wrong — because a proof answers the question you asked it, and the user was asking a different one.
If you have a boundary in your codebase — a database, an engine, a service, anything you call synchronously and trust — ask what the synchrony is quietly doing for you besides being slow. Pacing? Ordering? Completing before something dies? That question is cheaper than the migration, and it's the one I'd start with next time.
Comments (0)
No comments yet. Be the first to share your thoughts!