🧶Les pelotes de ficelle

Catch a loose end, pull the thread, unravel the whole ball.

← Back to the ball of yarn

The bug I set out to fix was already fixed — and the fix was wrong

Anatomy of a first contribution: three findings I wasn't looking for, one failed prediction, and what hardware says when code stays quiet

In two minutes

  • Starting point: I wanted to contribute and didn’t know where. I took the standard advice — a well-scoped task, in an active project, on hardware I own.
  • First surprise: the task I picked had been done six weeks earlier. The warning I was seeing on my phone wasn’t a gap, it was a version lag.
  • Second surprise: the recent fix contained an error. Two values swapped, in code reviewed by four people including the maintainer, and marked as tested.
  • The real work: a sensor missing from the tables. Values taken from the device, except one I couldn’t prove — and flagged as such in the commit message.
  • The failed prediction: I announced brighter photos after the fix. Measured: 39.2% before, 38.1% after. Nothing. The explanation is arithmetic, and it’s instructive.
  • The actual find: while chasing why the app kept freezing, a performance defect the package maintainer had instrumented himself, writing that it “indicates actionable issues in the V4L2 or GPU drivers”. Nobody had reported it.

This post is about method rather than cameras. It should read even if libcamera means nothing to you. The technical side is in the previous post.

The first-contribution problem

Everyone gives the same advice: “find an issue tagged good first issue”. It’s good advice and it doesn’t work very well, for a simple reason — those tasks are either trivial enough to teach you nothing, or already taken by someone faster.

What works better, I think, is to start from what you have that others don’t. In my case: an old phone I’d decided to run Linux on, and therefore physical hardware on which to run code other people write blind.

That’s a rarer advantage than it sounds. Plenty of developers work on drivers for hardware they don’t own, relying on datasheets and user reports. Whoever has the device plugged in on their desk can answer questions nobody else can settle.

The thread: a well-scoped task

The postmarketOS project keeps a per-device task list. For mine, an umbrella issue titled “Camera TODOs” listed a dozen items, all ticked except one:

imx355 driver (front camera) is missing features for libcamera, makes the later complain (e.g. when running cam -l)

Clear scope, reproducible symptom, one command to see it. Exactly what you look for.

I ran the command on the phone. It complained, as advertised. I opened the libcamera repository to write the fix.

And the entry was already there.

First lesson: an open issue isn’t necessarily open

Support for that sensor had been added on 17 July 2026, by a Raspberry Pi engineer. The version installed on my phone dated from 10 July.

Seven days apart. The warning I was seeing wasn’t a gap in the project, it was a lag between the released version and the repository.

That’s an easy confusion to make, and it deserves a reflex: before writing any code, check the problem still exists upstream, not just on your own machine. It takes two requests.

Going deeper: comparing a release to the repository

Most forges let you fetch a file at a given reference. Just compare the version you’re running with the main branch:

F="src/libcamera/sensor/camera_sensor_properties.cpp"
for ref in v0.7.2 master; do
  echo -n "$ref: "
  curl -s ".../repository/files/$(urlencode $F)/raw?ref=$ref" | grep -c '"imx355"'
done
# v0.7.2: 0
# master: 1

Zero occurrences in the release, one in the repository: the work is done but not shipped. There’s nothing to write.

The same reflex applies to the kernel. The second item on the list concerned a driver that didn’t answer a particular query; there too the fix already existed upstream — absent from releases 6.18, 7.0 and 7.1, present in the main branch. It simply hadn’t shipped yet.

Two tasks out of three had evaporated. I could have stopped there feeling I’d wasted my evening. Except that reading the freshly added entry, something was off.

Second lesson: review doesn’t see tables

The entry mapped two test patterns to numeric values. But the sensor’s kernel driver defines those values in the opposite order. “Solid colour” and “colour bars” were swapped.

I checked three times, by independent routes:

Source What it says
The driver on my phone, queried directly 1 = solid colour, 2 = colour bars
The driver’s source in the official kernel same order
Two other sensors with an identical menu, described right beside it correct mapping

And I actively looked for the context in which the author would have been right: his company maintains its own kernel, sometimes with different drivers. I checked — same order in both trees. The error was real, on his own hardware too.

What makes this interesting isn’t the error. It’s that the commit had been reviewed by four people, including the project maintainer, and carried a “tested” tag from a fifth.

How do five competent people miss two swapped values?

Because there’s no logic to review. It’s a correspondence between two documents that don’t live in the same repository: a table on one side, an array of strings in the kernel on the other. To check it you must open both files side by side, or have the device to hand. And “tested” here meant the camera works, I point it at a scene and get an image — which is true, and never exercises test patterns, which are a diagnostic tool.

The rest of the entry was right. Pixel size, sensor delays: everything exercised day to day was correct. Only the part nobody runs was wrong.

That’s where owning the device becomes a decisive advantage. Not because I’m better, but because I was the only one who could put the question to the hardware.

What a two-line patch has to contain

The patch is two lines. Its commit message is thirty, deliberately.

A reviewer should have nothing to go and look up. So I put in the kernel driver’s table quoted verbatim, the output of the command querying my phone, the concrete consequence in one sentence — asking for a solid colour produces colour bars, and vice versa — and the consistency argument with the two correctly described sensors.

And I pre-empted the question that was coming: “are the other entries wrong too?” No, and I checked driver by driver: two other sensors do use the reverse order, because their drivers define it that way. One sentence in the message saves a three-day round trip.

Going deeper: the commit-id trap

These projects use a convention to reference the commit you’re fixing:

Fixes: a1db25dabaee ("libcamera: camera_sensor: Add Sony IMX355 sensor properties")

The identifier is twelve characters. I had seven in front of me and filled in the missing five from memory. They were wrong. A made-up identifier that looks real is worse than a missing one: nobody checks it, and it points at nothing.

The correct form asks the tool rather than your memory:

git rev-parse --short=12 <reference>

It’s a small thing. It’s also exactly the kind of detail that gets a first patch sent back.

The real work, and the value you can’t prove

Something genuinely missing remained: the phone’s other sensor appeared nowhere. Not in the library, not in the official kernel — its driver was written at a chip maker eight years ago and never sent upstream.

I took the values from the device and from the driver’s source. Three follow rigorously. The fourth doesn’t: it’s a parameter you read in a datasheet, and datasheets for these sensors aren’t public.

I tried to back it up. I found a configuration file, on my own phone, declaring exactly the value I’d assumed. Excellent news for about three minutes — until I checked whether it was independent. It wasn’t: all ten equivalent files in the project declare the same value, including for sensors from three different manufacturers. That wasn’t ten concurring measurements, it was one default copied ten times.

So I wrote it as it stands in the commit message: this value follows what’s documented for sensors in the same family. Not “per the datasheet”, no phrasing that would suggest a measurement. A reviewer with the documentation will correct it in one message, and that’s fine.

Dressing up an uncertainty is the most efficient way to lose a project’s trust on your very first patch.

The prediction I got wrong

Reading the code, I’d worked out what the sensor’s absence caused: without it, the library no longer converts the camera’s gain. It writes a setting number where it should write an amplification factor, and reads the number back as if it were the factor. The auto-exposure loop is wrong in both directions.

From that I made a prediction: after the fix, low-light photos should be better. I set up a clean protocol — phone propped and never moved, objective luminance and noise measurements, a before series and an after series.

Result: 39.2% luminance before, 38.1% after. Nothing.

The explanation is arithmetic, and I had flagged it as a risk before running the test — without drawing the consequence, which was the mistake.

Gain requested Correct formula Broken formula
0 1.0× ≈ 1.0×
50 1.1× 50×
300 2.4× 300×

The error is only huge at high gain. My test scene had a saturated white area: auto-exposure had light to spare, stayed near zero gain, and at that point both formulas give the same result.

The patch is still correct, and it works — I verified that differently. After installation, the library emits no warning at all for that sensor, while the phone’s other sensor, unpatched, still emits every one of them. A clean negative control.

But I didn’t demonstrate a visible improvement, so I won’t write that I did. The “tested” tag I’ll attach will say the sensor is now recognised. Not that the images are better.

It’s a distinction that sounds pedantic and isn’t. A correct piece of reasoning about a mechanism says nothing about its magnitude under test conditions. I had the mechanism; I didn’t have the magnitude.

The find that was on no list

Throughout all this, the camera app kept freezing. A nuisance, which I worked around by relaunching.

Then I instrumented it, for lack of anything better. And it turned out that a single line, emitted once at initialisation, explained everything:

Importing input DMABuf failed, falling back to upload

Image processing runs on the graphics processor. Normally the raw frame reaches it without a copy, through shared memory. Here the import fails — and the library switches to copy mode: twelve megapixels transferred to the GPU on every frame. The memory bus saturates, the display can no longer get bandwidth for its own operations, and the pipeline ends up strangled.

It isn’t a crash. It’s a suffocation, which explains why it left no usable trace.

And here’s what gives that line its value. It comes from a patch added by the package maintainer, who had deliberately restored it after the upstream project removed it. His rationale:

failing imports majorly impact performance and indicate actionable issues in the V4L2 or GPU drivers.

So the maintainer had planted the detector, stating in writing that its firing warrants investigation. It fired on my phone. Nobody had reported it.

I haven’t proved causation yet — the correlation is clear, the mechanism coherent, and a simple test would settle it. But it is by far the most useful thing I found that morning, and it was on no task list.

What I take away, for next time

What you bring isn’t talent, it’s a position. I wasn’t better than anyone. I had the device plugged in and five competent reviewers didn’t. That’s all, and it’s enough.

Check the problem still exists before writing. Two tasks out of three had evaporated, fixed upstream but not yet released. Two requests would have said so immediately.

A confirming source is only useful if it’s independent. Ten documents copying the same default are worth no more than one.

Separate what’s demonstrated from what’s inferred — in a commit message as in conversation. I had a correct mechanism and a wrong prediction; stating one without the other would have been a polite lie.

And instrument what annoys you. The app freezing was a nuisance I’d been working around for hours. It was when I stopped working around it that I found the one thing nobody else was looking for.

It really is a ball of string: you pull on a two-line thread, and three things you never asked for come out with it.