LightPhotos

Six queues and a reserved worker

How LightPhotos loads images in the background, told as the problems that shaped it. Six queues and a worker that refuses the most common job sounds like overkill, I know. I added each piece because something visibly broke without it.

What I learned

One queue, several workers

LightPhotos loads (decodes) images on a few background threads, which I call workers, so the app stays responsive while they work. My first version was the simplest thing I could think of: one queue, a handful of workers, first come first served.

It worked until I scrolled. Scrolling the grid queues up a hundred thumbnails. Then you double-click a photo, and it goes to the back of the line behind a hundred thumbnails you're no longer looking at. The app looks frozen while it finishes work for part of the screen you already scrolled past, which is frustrating when all you want is to see your photo.

Some jobs matter more

So I split the one queue into six, ordered by how much you care about the result right now. The photo you just opened jumps ahead of the thumbnails.

SIX QUEUES, DRAINED IN THIS ORDER 1 First pass of the photo you opened 2 Sharp second pass of that photo 3 Full-resolution decode, on demand 4 Camera and lens metadata 5 Thumbnails for the grid 6 Capture times for burst grouping 8 DECODE WORKERS reserved · never takes thumbnails
The order that might surprise you is camera details (metadata) coming before thumbnails. The details panel sits next to a photo you're already looking at, so you'd notice a delay there, while a thumbnail two rows down isn't on screen yet.

I ordered the queues by what you're looking at right now. How quick a job is didn't factor in.

Being first in line isn't enough

With six queues, the freeze got shorter, but it didn't go away. Priority only decides which waiting job starts next. It does nothing about jobs already running. If all eight workers picked up thumbnails a millisecond before you double-clicked, your photo is at the front of the line and nobody is free to take it.

The usual answer is to interrupt a running job and give the worker to the urgent one. I couldn't do that, because the loading happens inside someone else's library, and it has no way to know I want it to stop.

What I could do was keep one worker aside. It never takes thumbnails at all. It sits idle while the others churn through thumbnails, and it's ready the moment you open a photo. On a machine with a single CPU core I turn this off, because setting aside the only worker would leave the grid with nothing.

A worker that can say no broke the wake-up

Workers sleep when there's nothing queued, and adding a job wakes one of them up. That's the standard way to do it, and it works as long as any sleeping worker can take any job. My reserved worker can't. It can say no.

WAKING ONE SLEEPER, WHEN ONE SLEEPER CAN SAY NO Thumbnail queued Wake one worker The reserved worker looks, declines, goes back to sleep the other seven were never woken No error. No crash. Just work that sits still, until something unrelated wakes somebody.
The fix is to wake all of them instead of just one.

Waking just one worker only works if every worker can handle every job. As soon as I let one of them refuse, that stopped being true. The app didn't crash or show an error. Work just sat there until something unrelated happened to wake a worker.

Marking work as started when nobody started it

The loader keeps a list of "already asked for this" markers so it doesn't queue the same photo five times, and it clears each marker when the job finishes. In the browser version, the worker pool doesn't start any threads at all, so nothing was ever going to finish those jobs or clear the markers.

There was no error message. Instead, the app woke up sixty times a second for the rest of the session, draining your laptop battery, to check on work that didn't exist.

I fixed it by treating each marker like a loan that someone has to pay back. Queueing a job now returns one of three answers: accepted, no workers exist, or the worker pool is broken. The code only marks the job as in progress on the first answer. I now think of any "loading" flag or "already requested" list this way. Only set it once whoever is responsible for clearing it has agreed to take the job.

In a browser, workers die

Moving the worker pool into a browser tab meant I had to babysit the workers much more closely. A newly started browser worker (a Web Worker) throws away any message sent to it before its script has finished starting up. So each worker now says when it's ready, and if one stays quiet for ten seconds I treat it as failed.

I replace a failed worker up to three times, then mark that slot as permanently gone. I kept that separate from "not ready yet" on purpose, because mixing them up caused a real hang. Exporting waits while the number of running jobs is at or above capacity, and capacity only counts ready workers. Without a "permanently gone" state, a slot stuck halfway through being replaced looks not ready, capacity drops to zero, and the export progress bar waits forever.

What made all of this testable

I can test all of the above without real threads, because the rules for who runs what live in a plain data structure. One test builds it, adds one job of each kind in reverse priority order, takes them all back out and checks the order. Another takes jobs out as the reserved worker and checks that thumbnails get skipped.

Code that runs things in parallel usually has a predictable core (the rules about who runs what, and in what order) wrapped in a layer whose timing changes from run to run. When I pulled the rules out on their own, I could test the decisions I'd made in microseconds, and the tests give the same result every time.