Review of the redrafted articles and corrected studies
Findings
The reviewer read the first review note, LH004 and LH003 at version 0.2 with their briefs, reading-list.md and sources.md, and both articles and their SVG files. Each article was read once straight through as a newcomer would, then again against the piece rules in AGENTS.md and the Piece form in LH F05. Three central claims in each piece were checked against the record beneath it and, where reachable, against the public source. Six public pages were fetched as raw HTML with curl and read with tags stripped, not through a summarising tool. The first review's findings 1 to 7 were all acted on in the text of the records and the pieces, not only in the Corrections sections. The fetches found new problems, mostly in passages the first review did not examine, and two of them concern the competing explanation the xz piece leaves out: that the March warnings were answered by the attacker rather than passed over.
The first review, finding by finding:
| First review finding | Where answered | Made in the text? | What remains |
|---|---|---|---|
| 1. LH004 lets the framework explain history | LH004 line 27 and its bold lead; Corrections item 1 | Yes | The Answer (line 15) still calls the early signal "dismissed at the time" and "not yet believed"; see finding 1 |
| 2. LH004's opening overreaches | LH004 Answer, line 15; Corrections item 2 | Yes; "nearly every" survives only as a quotation in Corrections | Nothing |
| 3. LH003's headline does not follow | LH003 Answer, first Finding, Limits, Next; Corrections item 1 | Yes; "almost disjoint" survives only where it is withdrawn | "Nobody has measured" is broader than the record; see finding 9 |
| 4. The Economic Index entry is outdated | sources.md row and note; two new columns; brief.md amendment | Yes, by describing the introductory report explicitly | A quotation not on the page, and a sentence resting on the first review alone; see finding 10 |
| 5. A definition presented as a discovery | LH003 Answer and second Finding; Corrections item 4 | Yes | Nothing |
| 6. Both articles carry the studies' machinery | Both articles | Yes: no ids, no "the study", no labels, no hashes; both illustrations are drawn | The xz piece defines "propagation" and never uses it; see finding 6 |
| 7. The xz payoff is sound | The xz article, "Seen, but not recognised" | Yes: availability, noticing and recognition are kept apart | The account of the early signal omits what the record says happened to it; see finding 1 |
Central claims checked, three in each piece and one more in the survey because the first review corrected it:
| Piece | Claim | Checked against | Stated as the record supports? |
|---|---|---|---|
| xz | The backdoored releases sat in public for nearly five weeks, the trigger only in tarballs made by Jia Tan, never in Git | LH004 second Finding; Tukaani page; Cox timeline; S14 | Yes. 24 February to 29 March is 34 days. Tukaani: "These tarballs were created and signed by Jia Tan"; the trigger code "was never included in the Git repository" (observation) |
| xz | A slow login gave it away, about 0.8 seconds where 0.3 was expected | LH004 fourth Finding; S14 | The figures, yes: S14 shows real 0m0.299s before and 0m0.807s after. The attribution only in part: S14's first sentence names valgrind errors beside the slow logins; see finding 4 |
| xz | Valgrind errors in early March were seen but not recognised, and the record does not say why they were not followed up | LH004 third and fifth Findings; Cox timeline; S14 | No; see finding 1 |
| survey | Nobody yet knows how the sources fit; nobody has measured their overlap | LH003 Answer, Limits, Next | Stated more strongly than the record; see finding 9 |
| survey | The counts will not add: different units, one job in several windows, some sources joined by design | LH003 Answer and first Finding; deps.dev page | Yes. deps.dev "indexes the Cargo, Go, Maven, npm, NuGet, PyPI, and RubyGems package ecosystems, the GitHub, GitLab, and Bitbucket project hosts, and the security advisories from OSV" (observation). The one-job example is framed as a picture, as a scenario should be |
| survey | None of the seventeen can say which downloads, crawls or conversations came from an AI system; the crawler window sees nothing of why a page was fetched | LH003 Answer; sources.md Cloudflare row; Cloudflare page | No; see finding 7 |
| survey | The Economic Index's first report, 10 February 2025, covered only Free and Pro conversations and said it did not represent AI use in general | sources.md row and note; Anthropic page | Yes for date and scope: "We also only analyze data from Claude.ai Free and Pro plans, rather than API, Team, or Enterprise users" (observation). See finding 10 for the record's quotation and the later reports |
Public pages read for this review, all on 26 September 2026:
- S14, Andres Freund, post to oss-security, https://www.openwall.com/lists/oss-security/2024/03/29/4. Posted 29 March 2024; describes symptoms seen "over the last weeks" before that date.
- Russ Cox, Timeline of the xz open source attack, https://research.swtch.com/xz-timeline. Posted 1 April 2024, updated 3 April 2024; covers 29 October 2021 to 30 March 2024.
- Tukaani project, XZ Utils backdoor, https://tukaani.org/xz-backdoor/. Last updated 17 January 2025; no observation window beyond the incident.
- Anthropic, The Anthropic Economic Index, https://www.anthropic.com/research/the-anthropic-economic-index. Dated 10 February 2025; states no date range for the conversations analysed.
- Open Source Insights documentation, https://docs.deps.dev/. No page date shown; describes current coverage.
- Cloudflare, Analyze AI traffic, https://developers.cloudflare.com/ai-crawl-control/features/analyze-ai-traffic/. Last updated 23 April 2026; describes the product as of that date.
The xz article and LH004
-
Must fix before the keeper reads it. articles/the-xz-backdoor.md lines 33 and 37; the-xz-backdoor-timeline.svg, third lane; studies/LH004/LH004.md lines 15, 27 and 47; studies/LH004/reading-list.md line 31. The March warning was answered by the attacker, and the piece says the record does not say why it was not followed up. The record beneath the piece already holds the fact: LH004 line 23 reports a change on 8 March that the timeline calls a misdirection and a further change to the test files on 9 March. The primary source confirms it: S14 says the injected code "caused valgrind errors and crashes in some configurations", that "These issues were attempted to be worked around in 5.6.1", that "the exploit code was then adjusted", and that the committer "communicated on various lists about the 'fixes'" (observation). The Cox timeline dates it: Red Hat's distributions began seeing the errors on 4 March; Jia Tan committed a "purported Valgrind fix" on 8 March, called there "a misdirection, but an effective one", and on 9 March the updated backdoor files, "the actual Valgrind fix", released as 5.6.1 (observation). So the reports were followed up, by the person who planted the backdoor, and the 9 March release the piece mentions at line 19 was his answer to them. What must change: at line 37, replace "The record does not say how far the March reports travelled, or why they were not followed up at the time" with an account of that response, and say at line 19 or 33 that 5.6.1 is the release that made the errors go away. "We think a memory checker's complaint during testing is easy to read as an ordinary bug" can stay if it adds that the attacker supplied the ordinary-bug explanation himself. At line 33, "records valgrind errors reported from Fedora's build and test machines between about 4 and 9 March" should say what the timeline says: errors seen from 4 March until the 9 March release silenced them. In the timeline figure, join the valgrind band to the 5.6.1 mark, for example "Jia Tan's 'fix', 8 to 9 Mar". In LH004: the Answer's "dismissed at the time" and "seen earlier and not yet believed"; line 27's "no source this study read says why the reports went uninvestigated"; and Limits line 47, which says the early signal rests on secondary reconstruction alone, when the primary disclosure the study read corroborates the errors and the 5.6.1 workaround, though not the 4 March date or the attribution to Fedora. In reading-list.md line 31, the gist's aside shows that one Gentoo developer did not reproduce the errors, not that they "went uninvestigated". Correct the record first and draw the piece from it, as the drive procedure asks.
-
Must fix. articles/the-xz-backdoor.md line 25; studies/LH004/LH004.md line 29. "Rather than anything coordinated" is contradicted by a source the piece cites. The Cox timeline records that on 27 February Jia Tan "starts emailing Richard W.M. Jones to update Fedora 40"; that on 25 March "Hans Jansen is back (!), filing a Debian bug to get xz-utils updated to 5.6.1", while more "addresses that don't otherwise exist on the internet show up to advocate for it"; and that on 28 March Jia Tan filed an Ubuntu bug to the same end (observation). The distributions took new releases in their ordinary rhythm, and the attacker also pushed. The hedge "We think" does not rescue a reading the piece's own source contradicts. What must change: drop "rather than anything coordinated" and say once that the attacker lobbied at least Fedora and Debian to take the new versions. In LH004 line 29, the "because" clause gives the ordinary pipeline as the only reason; add the lobbying the timeline records.
-
Must fix. articles/the-xz-backdoor.md line 25; the-xz-backdoor-timeline.svg lines 3 and 55; studies/LH004/LH004.md lines 29 and 49. Debian carried the backdoor for about a month, not "for some days", and a cited source gives the dates. The Cox timeline: Debian added 5.6.0-0.1 to unstable on 26 February and 5.6.0-0.2 to testing on 5 March, updated to 5.6.1 on 27 March, and rolled back on 28 March, the day before the public post (observation). S14 says Freund saw the symptoms "on Debian sid installations over the last weeks". The figure's note, "Also carried it, on dates the sources do not give: Debian unstable and testing, openSUSE Tumbleweed", is wrong for Debian. What must change: in the piece, date Debian's window from the timeline; in the figure, give Debian dated marks in the middle lane (unstable 26 Feb, testing 5 Mar, rollback 28 Mar) and keep the undated note for openSUSE Tumbleweed alone; in LH004 line 29, replace "for some days" for Debian; in Limits line 49, "At the time of disclosure, Debian's own unstable and testing suites did carry the compromised 5.6.1 build" should give the timeline's dates instead. The correction sharpens the piece's own point at line 41: Freund's machine ran a distribution that had carried the backdoor for a month.
-
Should fix. articles/the-xz-backdoor.md lines 5, 35 and 43. Freund's own symptoms included the valgrind errors. S14 opens: "After observing a few odd symptoms around liblzma (part of the xz package) on Debian sid installations over the last weeks (logins with ssh taking a lot of CPU, valgrind errors) I figured out the answer" (observation). The piece quotes those words at line 11, then gives the test errors to Fedora and the felt slowdown to Freund (line 35) and closes on "one person noticed half a second and followed it" (line 43). What must change: at line 35, say that Freund had also seen valgrind errors on his own machine. The half second can keep its place in the title and the last line, since it is what he measured and followed, as long as the text does not make it the only signal. The change strengthens the piece's case for watching test output.
-
Should fix. studies/LH004/LH004.md line 25; articles/the-xz-backdoor.md line 15. The identifier was assigned on 28 March and published on 29 March. The Cox timeline records "RedHat assigns CVE-2024-3094" on 28 March; S15, as LH004 reads it through the NVD API, was published on 29 March. LH004's "assigned CVE-2024-3094 the same day" should read published that day, assigned the day before. The piece's "given the public identifier" is defensible; "published under the identifier" is exact.
-
Should fix. articles/the-xz-backdoor.md lines 25, 33, 39 and 47; studies/LH004/reading-list.md. Four passages that trip or fall short of the piece rules. (a) Line 25 defines propagation and never uses the word again. The fact carries the paragraph without the label, and the rule is that the three layers appear only where they earn their place; residue and activity earn theirs, in the contrast between evidence at rest and code running. (b) Line 33 moves from "he" to "I" inside one reported sentence and trails ", another Linux distribution" after the quotation. Quote Sam James directly and put the gloss beside Gentoo's first mention. (c) Line 39, "would not have shown which distributions built the releases in either", has to be read twice; for example, "would not have shown, either, which distributions had built the releases into their packages". (d) The colophon at line 47 dates four of its eight sources. Add the dates the record can give: the Tukaani page, last updated 17 January 2025; the Cox timeline, posted 1 April 2024 and updated 3 April 2024; the Red Hat CVE page and the Fedora Magazine article once the record has them. The same gap is in reading-list.md, which gives no publication date for most entries, against the records rule in AGENTS.md. LH003 added those columns for its sources, and LH004 should do the same.
The survey article and LH003
-
Must fix before the keeper reads it. articles/what-we-can-see.md lines 19 and 39; studies/LH003/LH003.md line 15; studies/LH003/sources.md line 14. The crawler window does say which crawls came from AI systems, and it sorts them by purpose. Line 19 rightly says Cloudflare's feature counts requests "from crawlers it has matched to a named AI company". Line 39 then says none of the seventeen "can say which downloads, crawls or conversations came from an AI system rather than a person", which the piece's own description of the crawler window contradicts. Cloudflare's page lets a site owner group AI-crawler traffic by crawler, by operator and by "Category: Analyze crawlers by their purpose or type" (observation). Line 19's "nothing of why a page was fetched: to train a model, or to answer someone's question" is therefore wrong about the product as a whole. It holds only for a single fetch, because the category belongs to the crawler, not to any one request (interpretation). Line 39's "None sees work in private projects" is also too broad: the broker and the vendor see requests that come from private software, even though they do not see the projects. What must change: line 39 should say what the record supports, for example that no source can tell a person's download or code change from an agent's, that the crawler counts cover only crawlers already identified, and that the conversation figures do not say whether a person or a program was typing. Line 19 should say that the feature sorts identified crawlers by their stated purpose but cannot say why any single page was fetched. LH003's Answer ("What stays invisible across all seventeen") and the "What it cannot see" cell of the Cloudflare row in sources.md change to match, citing the page.
-
Must fix. articles/what-we-can-see-map.svg; articles/what-we-can-see.md lines 21 and 23. The illustration draws the two things the record says nobody has measured. Four small, separate patches in a mostly bare field tell the eye that the windows do not overlap and that together they see little of the whole. LH003 withdrew the first as "almost disjoint by construction", and its Limits (line 49) say it cannot tell how much of the wider population any source covers. The piece's own example at line 31, one job entered in up to four windows, contradicts the picture, and the caption has to argue against its own drawing. The alt text at line 21 and the SVG's desc (line 3) say "most of the field is empty", which contradicts the caption's "unobserved, which is not the same as empty" and the point of line 41. What must change: redraw so that neither overlap nor coverage is asserted. Two options: draw the four windows overlapping with the shared regions marked unknown, or draw the one-job picture, a single task passing four windows with nothing joining the four entries. Then correct the alt text and the desc to say "unobserved", not "empty".
-
Should fix. articles/what-we-can-see.md lines 3, 23, 35 and 43; what-we-can-see-map.svg line 27; studies/LH003/LH003.md line 66. "Nobody knows" and "Nobody has measured it" claim more than the record. LH003 shows that no source read states its overlap with another and that no joint sample was taken (line 15). It did not look for anyone else's measurement. At line 35 the piece's next two sentences give the true scope. What must change: at each place, say that none of the sources says and that we have not measured it, or words to that effect. In LH003's Next (line 66), "today nobody has measured" should become "no source read states, and this study did not measure". The title can stand, because its "we" is accurate.
-
Should fix. studies/LH003/LH003.md lines 19 and 23; studies/LH003/sources.md lines 22 and 37; articles/what-we-can-see.md line 17. The Economic Index row quotes words the page does not contain, and one sentence in the piece rests on the first review alone. The page reads: "we don't argue that the uses in our dataset are a representative sample of AI use in general". The phrase LH003 puts in quotation marks, "not representative of all AI use generally", is not on the page. Nor is the privacy floor of 15 conversations or 5 accounts that line 23 attributes to the introductory report (observation, from the raw page; sources.md line 37 already warns that its quotations came from a model-written extract). The row's "image-generation conversations (excluded)" misreads the page's remark that Claude cannot generate images: those uses are absent, not excluded. What must change: quote the page exactly or drop the quotation marks, and attribute the privacy floor to the paper if that is where it appears, or mark it unconfirmed. In the piece, "Later reports widened the scope" (line 17) has only the first review note behind it, and that note gives no URL. Either read and cite a later report in the record, or attribute the sentence in the piece. "It said plainly that it did not represent AI use in general" is a fair paraphrase; "it did not claim to represent" would be exact.
-
Should fix. articles/what-we-can-see.md lines 3 and 7. The opening does not begin with the thing itself. It opens on a question, and its third sentence, "We tried to put the pieces together", is the method. The piece's best surprise waits until the second section (line 31): one agent's single job is counted up to four times, and nothing joins the entries. Consider opening with it, because it gives the reader a reason to care in one image.
-
Should fix. articles/what-we-can-see.md lines 33, 41 and 45. Three passages trip when read aloud. (a) Line 33, "The third is that the windows are not simply separate, either", comes straight after a paragraph that has just shown they are not separate. Its evidence comes from two sources the reader has not met, when the reader is holding four windows. Say that these joins are among the other sources in the seventeen, and say whether any of the four windows shown shares an identifier; LH003 line 19 records that the broker, the vendor reports and Cloudflare state none. (b) Line 41 brings in harm and danger, which nothing earlier prepares. Keep "empty" and drop "harmful" and "dangerous", or first say why someone would read a gap as danger. (c) The closing question at line 45, "when the same piece of work passes in front of two of these windows, how often do both of them see it?", answers itself, because passing in front of a window means being seen by it. For example: "when one of these windows sees a piece of work, how often does another see it too?"
Both pieces
-
Note. Layout and colophons. Both pieces meet the Piece form and the piece rules. Each has a title that says the finding, an opening with a reason to care, a discreet date, byline and Draft label, three headings, a drawn illustration with a one-line caption, and a colophon with sources, method in a sentence, authorship, review, version, corrections and one line on Lighthouse. No document number, id, "the study", "the keeper", record label, charter citation or commit hash appears in either. Small differences would show to someone reading both: the date line runs "Draft · date · Lighthouse" in one and "date · Lighthouse · Draft" in the other; the survey uses LH F05's one-paragraph colophon while the xz piece splits it into labelled paragraphs; one names Claude's maker and the other does not. Neither colophon says where a reader can find the record beneath it, although the handbook says each rung "points down to it in its colophon, so a reader can descend". A link at release would do it.
-
Note. Smaller points of fact and wording. In the survey, line 33's "links each package to its project" is a little stronger than the deps.dev page, which says the service indexes the ecosystems, the project hosts and OSV's advisories; "links packages to their projects" would be exact. Line 13 puts download counters under "Public code forges": the SVG's sublabel handles this, but the heading does not. In the xz piece, line 19 attributes Jia Tan's first appearance in October 2021 to "the project's own account"; the date comes from the Cox timeline, not the Tukaani page. Line 3 sets "a lot of processor time" beside an elapsed-time figure. S14's own measurement shows the client's CPU unchanged at 0.202 seconds, with the extra time spent in the server. In the timeline figure, "Fedora takes 5.6.0" is placed at 26 February, before the 27 February date on which the Cox timeline says Jia Tan began asking Fedora to take it. Either date the mark from the Red Hat and Fedora sources or move the estimate later. The Cox timeline says "RedHat distributions start seeing Valgrind errors"; neither it nor S14 says the errors came from build and test machines. Finally, "Fedora notified" is a notification, not something noticed, and sits oddly in a lane called "What people noticed".
-
Note. The SVG files, read as text. Both parse as XML. Each sets a viewBox that matches its width and height (800 by 300, and 800 by 400), paints its own background, loads nothing external (the only URL is the SVG namespace) and uses system fonts. Measured with FreeSans metrics, a Helvetica-like face, no text runs outside the viewBox or collides with a mark or another label (derived measurement). The xz date marks sit where the stated scale puts them: x = 165 + 17.5 per day from 24 February, with 2024 a leap year, gives 1 March at 270, 9 March at 410 and 29 March at 760. The smallest text is 11 pixels in an 800-pixel drawing, about 5 pixels on a phone-width screen, so consider larger type for phones. Apart from findings 1, 3 and 8, the marks match the captions.
-
Note. Record housekeeping. studies/LH004/LH004.md lines 27 and 31 say that I-0002 and LH003 remain
todo. Both are now done in harbour/tickets.json. The statements were true at grounded_at, and version 0.2 records one later change (the Correctness surface), so it could record these too. LH004 version 0.2 gives no grounded_at, author or model for its corrections, which LH003's header does, and its Corrections section is not in LH F05's correction-record form, which LH003's is. LH003's "earlier edition retained at: git history" should name a path and commit as LH004's does: studies/LH003/LH003.md@af3cf58, the commit that created version 0.1, which is unchanged at e0b7ca4 where the first review read it. The first review's findings 8 and 9 were outside this task and were not re-checked.
Verdict
- articles/the-xz-backdoor.md: not yet; findings 1, 2 and 3 block it, and LH004 should be corrected first so the piece is drawn from the corrected record. It is otherwise close: the opening, the three-events passage and the survivorship caution read well.
- articles/what-we-can-see.md: not yet; findings 7 and 8 block it, with LH003 and sources.md corrected first for finding 7. The middle section, why the numbers will not add, is the strongest writing in either piece.
- The records: LH004 and LH003 answer the first review's findings in their text. Both need a version 0.3 before the redraft: LH004 for findings 1, 2, 3 and 5, and LH003 for findings 7, 9 and 10.
What the review could not check
- Not fetched, because of the six-fetch limit: Sam James's gist (the quotation at the xz piece's line 33, the ignore-list mechanism, "last updated 9 September 2026"); Red Hat's blog and CVE page and Fedora Magazine (Fedora's dates, RHEL never affected, the 28 March notification and 29 March reverts); the NVD API (the score of 10.0 and the 29 March publication); GH Archive's "every public event ... since February 2011"; OSV.dev's upstream sources; OpenRouter's licence; PyPI Stats's 180 days; and any Economic Index report later than the first.
- The 4 March start of the valgrind reports rests on the Cox timeline alone. The xz-utils git history was not read.
- The SVGs were not rendered, because this container has no SVG renderer. Text extents were computed from FreeSans metrics, and a browser's system-ui face may be somewhat wider or narrower.
- Reading aloud was done as a silent read for rhythm, once straight through and once against the rules. No audio was produced.
- This review's own collection was mostly untagged. The six pages were fetched with curl from this container, five with curl's default user agent and one with a user agent naming Lighthouse. The cost of the review is not measured here.
From notes/R-0001.md in the repository, last changed 26 September 2026.