Lighthouse

Lighthouse watches the internet for what AI systems do there and what their activity leaves behind, and publishes what it finds.

Studies LH003

LH003 survey table

The two tables below are the deliverable LH F04's LH003 design section and the dispatch prompt both ask for: one row per candidate observatory, in the column order the dispatch prompt gave (source, layer seen, population, boundary, unit, cadence, licence, what it cannot see, S number or URL), and a second table for candidates considered and excluded, with the reason. Version 0.2 added two columns to the included table, publication date and observation window, beside a last column that now gives each row's date read (see LH003.md's Corrections).

Layer names follow Lighthouse_Research_Design.md (LH F02): activity is computation happening, residue is the altered structure and state it leaves behind, propagation is residue enabling later activity elsewhere and is not itself a surface. F02's "Outcomes" row sits outside the three layers and scores activity and residue against a declared purpose; a few candidates below fit there better than in any of the three, and the layer cell says so rather than forcing a fit. Where a source spans more than one surface, the cell names the primary one and notes the other.

Every cell that could not be confirmed from a source read on the date below is left blank or marked "not stated", per AGENTS.md; nothing here is invented. A fuller note follows the table for any row whose cells need a qualification too long to fit them honestly. Every row was read on 26 September 2026, and the last column says so row by row; where a row also cites an LH F02 source (S number), F02's own check of that link was made on 25 September 2026. Publication date is the date a source's page gives for the version read; observation window is the period of activity its data covers, as the source states it. Both are filled only where the page, as this survey recorded it, states them; elsewhere the cell is blank, with a note in brackets saying why. Only the Anthropic Economic Index page was read again for version 0.2. Version 0.3 takes its changes to the Cloudflare AI Crawl Control, Open Source Insights and Anthropic Economic Index rows from notes/R-0001.md, whose reviewer read those pages as raw HTML on 26 September 2026; no page was fetched again for it.

Included: observatories with a row

Source Layer seen Population Boundary Unit Cadence Licence What it cannot see Publication date Observation window S number or URL, and date read
GH Archive Residue (Structure); some event types (comments, stars, watches) read closer to Activity's Communication surface All public GitHub events since Feb 2011 (Timeline API to 2014, Events API from 2015) One forge, public events only; payload shape varies by event type and can change without notice One JSON event record per GitHub event, 15+ event types, aggregated into hourly files Hourly, continuous since 2011 Not stated on gharchive.org itself; redistributed via a public Google BigQuery dataset (1 TB/month free tier) and Snowflake Marketplace Private repositories and activity; no code diffs beyond a payload's own fields; no deployment or runtime data; no reliable AI-vs-human attribution (our inference from what the records hold; not stated on the page) (blank; a continuously updated archive, no page date recorded at the read) February 2011 to the read date, in hourly files (Timeline API to 2014, Events API from 2015) S5; https://www.gharchive.org/ read 26 Sep 2026
Cloudflare AI Crawl Control Activity (Communication): identified AI-crawler HTTP requests Traffic reaching a Cloudflare zone that has the feature enabled, from crawlers Cloudflare has fingerprinted to a named AI operator Per zone (domain), not global; only crawlers Cloudflare has identified; path grouping to 2 levels, not aggregated across hostnames Requests (count), bytes transferred, HTTP status buckets, referral visits; traffic can be grouped by crawler, by operator and by category ("purpose or type") Queryable by date range; refresh interval not stated Included with a Cloudflare account for that zone; "total referrals" and "top referrers" need a paid plan Unfingerprinted crawlers; traffic off Cloudflare entirely; the reason for any single fetch (the page groups identified crawlers by "purpose or type", a label on the crawler, not on a request; see note); zones with the feature off 23 Apr 2026, last updated, per notes/R-0001.md's raw read of the page, 26 Sep 2026 (blank; a query covers a date range the zone owner chooses; how far back data is kept was not recorded) S7; https://developers.cloudflare.com/ai-crawl-control/features/analyze-ai-traffic/ read 26 Sep 2026; read again as raw HTML by notes/R-0001.md, 26 Sep 2026
Cloudflare Radar Activity (Communication: requests, DNS queries, attacks); a technology-adoption section reads closer to Residue Traffic through Cloudflare's own network plus queries to its 1.1.1.1 public resolver Cloudflare's network and resolver only; a site or query never touching either is invisible to it Aggregated, anonymised counts by category, geography, time; specific native units not itemised on the page read Not stated on the page read CC BY-NC 4.0 for published data; free API for academic, technical and general use Traffic that never crosses Cloudflare; individual users (aggregated by design); the main radar.cloudflare.com site refused this study's automated check, per F02's existing note (blank; no page date recorded at the read) (blank; not recorded at the read) S18; https://developers.cloudflare.com/radar/ read 26 Sep 2026
Software Heritage Residue (Structure): archived source code and development history Public repositories on the forges its crawlers cover, plus anything explicitly deposited or save-requested Source-code artefacts and history only (commits, releases, directories, files), addressed by SWHID; not runtime behaviour or private repositories One SWHID-addressed object (revision, release, directory or content) per record Continuous automated crawling, plus on-demand save requests A stated API terms-of-use page exists; exact text not read Anything not yet crawled or deposited; whether a commit was authored by a human or an agent; softwareheritage.org's own main site errored at F02's check (blank; no page date recorded at the read) (blank; not recorded at the read; the archive grows by continuous crawling) S10; https://docs.softwareheritage.org/ read 26 Sep 2026
Open Source Insights (deps.dev) Residue (Structure): resolved dependency graphs Packages on Cargo, Go, Maven, npm, NuGet, PyPI, RubyGems, linked to GitHub/GitLab/Bitbucket project metadata and OSV advisories What a manifest resolves to, not what is installed or deployed anywhere; only the seven named ecosystems One resolved package-version graph node/edge per record Described as updated "regularly"; no interval stated Website, HTTP/gRPC API and a public BigQuery dataset offered; no explicit open-data licence stated on the page read Deployed or running code; private packages; whether a dependency bump was made or reviewed by a human or an agent (blank; no page date shown, per notes/R-0001.md's raw read, 26 Sep 2026) (blank; current resolved graphs, updated "regularly"; how far back history goes was not recorded) S11; https://docs.deps.dev/ read 26 Sep 2026
OpenRouter rankings Activity (Inference): tokens processed by model, through one broker Developers/apps that route calls through the OpenRouter API; not inference generally One commercial broker's own traffic; models/providers it does not proxy are invisible to it Tokens processed per model; benchmark and pricing figures shown alongside Usage shown "through" a stated date; refresh interval not stated Rankings data CC BY 4.0; a separate Data API documented for programmatic and cited access Direct-to-provider inference, self-hosted models, or traffic through another broker; identity or purpose beyond "app" attribution (blank; the page shows usage "through" a date it states, which the read did not record) (blank; ends at that stated date; start not recorded) S19; https://openrouter.ai/rankings read 26 Sep 2026
Vulnerability databases: CVE.org + NVD Residue (State/Structure): disclosed vulnerabilities, enriched with severity scoring Software assigned a CVE ID by a CVE Numbering Authority (400+ CNAs under MITRE and CISA, per a 26 Sep 2026 search read); NVD then enriches CNA records with CVSS Only disclosed, formally CVE-assigned flaws; undisclosed or unassigned flaws are invisible by construction One CVE record per vulnerability (e.g. CVE-2024-3094, CVE-2021-44228) Continuous; CNAs issue from assigned ID blocks, "tens of thousands" a year per the search read Not confirmed; both cve.org and nvd.nist.gov required scripting and returned only a bare header to this study's fetch tool Exploitation in the wild; patch adoption; propagation beyond disclosure; a known-but-unassigned flaw; publish-date lag behind actual discovery (blank; both primary pages failed to fetch) (blank; both primary pages failed to fetch) S15, S16; most cells from a WebSearch read, 26 Sep 2026; https://www.cve.org/About/Overview (fetch failed, 26 Sep 2026); https://nvd.nist.gov/general/News/change-timeline (fetch failed, 26 Sep 2026)
OSV.dev Residue (State/Structure): aggregated open-source vulnerability records 50+ ecosystems at the read (npm 228,730 entries, PyPI 25,153, Maven 7,040 among them), several Linux distributions, sourced from partner databases (GHSA, PyPA, RustSec, others) An aggregator, not a discoverer; coverage is a function of its upstream sources' own gaps One OSV-schema entry per vulnerability, keyed to affected versions or commit ranges Not stated on the page read Described as an open-source project; specific data licence text not read Anything not already published by an upstream source; no independent discovery of its own (blank; no page date recorded; entry counts are as at the read) (blank; not recorded at the read) https://osv.dev/ (read 26 Sep 2026)
GitHub Security Advisories (GHSA) Residue (State/Structure) 12 ecosystems (Composer, Erlang, Go, GitHub Actions, Maven, npm, NuGet, Pip, Pub, RubyGems, Rust, Swift) "GitHub-reviewed" advisories are validated before publication; "unreviewed" ones come straight from the NVD feed with no validity check and get no Dependabot alert; "malware" advisories are sourced from npm security and OpenSSF's malicious-packages list One GHSA-ID record (GHSA-xxxx-xxxx-xxxx) per advisory, published in OSV format Continuous; EPSS scores sync daily No licensing or reuse terms stated on the page read Vulnerabilities in ecosystems outside the 12 listed; anything not submitted to or reviewed by GitHub; malware advisories carry no fix version by design (blank; no page date recorded at the read) (blank; not recorded at the read) https://docs.github.com/en/code-security/security-advisories/working-with-global-security-advisories-from-the-github-advisory-database/about-the-github-advisory-database (read 26 Sep 2026)
Anthropic Economic Index, introductory report Activity (Inference): conversations classified by occupational task In the introductory report: Free and Pro conversations on Claude.ai only, explicitly excluding API, Team and Enterprise traffic; the first review note says later reports widened the scope, not checked here (see note) One vendor's consumer tiers in the introductory report; the page says "we don't argue that the uses in our dataset are a representative sample of AI use in general"; coding is noted as overrepresented Percentage of conversations mapped to O*NET occupational tasks; an augmentation/automation split Periodic named reports (dated entries found through June 2026 at the read); "we'll repeat... over time" Underlying dataset stated as open-sourced on Hugging Face; exact licence text not read In the introductory report: API/Team/Enterprise traffic; other providers entirely; whether output was used at work; image generation (absent, because Claude does not generate images, not excluded); cells below a privacy floor (a floor of 15 conversations or 5 accounts was recorded in version 0.1 but is not on the page; unconfirmed) 10 Feb 2025, the introductory report, as its page states (page read again 26 Sep 2026 to date it) (blank; the page gives no date range for the conversations it analysed) https://www.anthropic.com/research/the-anthropic-economic-index (read 26 Sep 2026; read again the same day for the version 0.2 correction); read again as raw HTML by notes/R-0001.md, 26 Sep 2026
OpenAI, "How People Use ChatGPT" Activity (Inference-adjacent): message volumes and classified conversation content Growth data: all paying ChatGPT users, Nov 2022 to Sep 2025; a classified sample of ~1M de-identified messages, May 2024 to Jun 2025; an employment/education sample of ~130,000 users via a "secure data clean room" OpenAI's consumer ChatGPT product only; reported at 700M users and 18B messages/week by Jul 2025 Messages, classified topics, self-reported demographic/occupational categories A one-off published study (NBER working paper w34255), not a continuously queryable instrument Not established from a direct read; see Limits Enterprise/API ChatGPT use; other providers; demographics outside the drawn sample; no stated open dataset (blank; NBER working paper w34255, whose date was not read; see the row's note) Growth data Nov 2022 to Sep 2025; classified message sample May 2024 to Jun 2025; per a WebSearch summary, not the primary text https://www.nber.org/system/files/working_papers/w34255/w34255.pdf (found via WebSearch, 26 Sep 2026; direct PDF fetch and the companion openai.com summary page both failed the same day, see Limits)
Common Crawl Residue (State/Structure): archived web pages; not AI-specific Pages CCBot could reach across the open web; sites can opt out via robots.txt Over 300 billion pages since 2007, 3-5 billion added monthly, per the page read; language/demographic gaps are noted in research the page cites but not itemised in it One archived page/capture per record; CDXJ, URL, host and domain-level indices Monthly crawl releases, recent ones spanning three-month windows Described as "free and open"; separate Terms of Use and Privacy Policy referenced but not read Anything CCBot could not or was not allowed to crawl; whether content was AI-authored or altered; dynamic/authenticated content (blank; no page date recorded at the read) 2007 to the read date; monthly releases, recent ones spanning three-month windows https://commoncrawl.org/ (read 26 Sep 2026)
PyPI download stats Activity (Communication): package-download requests Downloads of packages on the Python Package Index Excludes known mirrors (e.g. bandersnatch) unless noted; does not distinguish a human install from a CI job (our inference from what is counted; not stated on the page); 180-day retention Aggregate download counts, JSON API for current and historical series Derived from PSF's public BigQuery download logs; refresh interval beyond the 180-day window not stated Run by the Python Software Foundation, described as open source; specific data licence not read Which downloads feed an AI system vs ordinary use; private-mirror installs; anything beyond 180 days without a separate BigQuery query (blank; no page date recorded at the read) A rolling 180 days https://pypistats.org/about (read 26 Sep 2026)
npm download-count API Activity (Communication): package-download counts Downloads of any package on the public npm registry, from log data held since 10 Jan 2015 Bulk queries: at most 128 packages, 365 days; single-package queries: at most 18 months; per-version counts: previous 7 days only; scoped packages unsupported in bulk queries Downloads per day/period, via api.npmjs.org's point and range endpoints Computed once daily, shortly after UTC midnight, from the previous day's logs Not stated in the documentation read What exactly counts as one download (the doc defers to an external post not read here); mirror or CI-inflated counts; private registries (blank; no page date recorded at the read) Logs from 10 Jan 2015 to the previous UTC day; one query reaches at most 18 months (single package), 365 days (bulk) or 7 days (per version) https://github.com/npm/registry/blob/main/docs/download-counts.md (read 26 Sep 2026)
Epoch AI Residue (Structure): a maintained catalogue built from public information, not a live feed; a separate polling dataset falls outside the three layers (see note) Over 3,600 ML models, 1950 to present, at the read; data centres, GPU clusters, chip sales, organisations; a polling dataset on employed Americans' workplace AI use A curated catalogue, not a census of all models trained; physical-infrastructure entries use methods such as satellite and permit data One model, data centre, chip or benchmark-evaluation record per entry Most datasets monthly or quarterly; some (e.g. GPU clusters) less often CC BY, per the page read Models never publicly disclosed; the workplace-adoption figures are self-reported, not usage telemetry (blank; no page date recorded at the read) Models from 1950 to the read date; other datasets' windows not recorded https://epoch.ai/data (read 26 Sep 2026)
AI Incident Database Outside F02's three layers; closest to its separate Outcomes category, and built from voluntary reporting rather than direct observation Incidents someone submitted and the Responsible AI Collaborative indexed as real or near harm from deployed AI, per the page's own wording Only incidents that reached public reporting and were then submitted; no denominator of total AI deployments exists One incident record per case, with a unique identifier and linked source reports Continuous; new incidents added as submitted and reviewed Open for browsing and submission; no data-reuse licence text read Unreported or unrecognised harms; any incidence rate (no denominator); coverage shaped by who chooses to submit (blank; no page date recorded at the read) (blank; not recorded at the read) https://incidentdatabase.ai/ (read 26 Sep 2026)
Hugging Face Hub Residue (Structure: models/datasets that exist) and Activity (Communication: download counts as a usage proxy) Public models and datasets anyone uploaded; ~3M public models (milestone announced 18 Aug 2026) and datasets past 1M during 2026, per a search read Self-hosted uploads only; existence says nothing about whether a model is ever run; 85.6% of models have under 200 lifetime downloads and the top 50 entities take over 80% of all downloads, per the same search read One model/dataset repository per record; a documented API splits recent activity from a cumulative downloadsAllTime figure Publisher-analytics figures described as daily; no stated refresh interval for the headline model/dataset counts Per-repository, set by the uploader (Apache 2.0, MIT, OpenRAIL variants, others); Hub Terms of Service effective 15 Sep 2022, per the search read Local or re-hosted mirrors that bypass the Hub's own counter; private/gated repositories; whether a downloaded model was ever actually run (blank; page content thin; no page date recorded) (blank; the API separates recent activity from an all-time download total; neither period's dates recorded) https://huggingface.co/docs/hub/models-the-hub (fetched directly 26 Sep 2026, page content thin); supplementary figures from a WebSearch read the same day, sources not directly fetched

Notes on rows above

GH Archive. LH002's own design section already names the qualification that matters most for this survey: event records alone do not give reliable AI attribution, deployment records, or every code diff.

Vulnerability databases. F02 already records that nvd.nist.gov "needs scripts to render" for its S15 entry; this study's own attempt at both nvd.nist.gov and cve.org on 26 September 2026 hit the same wall (a bare page header, no body text a fetch tool could read), so the licence and exact update cadence for both are left blank rather than guessed. The CNA/NVD split (who assigns an ID versus who scores it) came from a WebSearch summary, not a direct page read, and should be checked against a primary CVE Program page before a study relies on it.

Anthropic Economic Index. Version 0.1 gave this row's scope as Free and Pro only without saying which report the scope belonged to. The page read is the introductory report, which its page dates "Feb 10, 2025" and which states: "We also only analyze data from Claude.ai Free and Pro plans, rather than API, Team, or Enterprise users." notes/R-0001.md's raw read of the page on 26 September 2026 confirmed both. That read also found three errors carried from version 0.1: the page says "we don't argue that the uses in our dataset are a representative sample of AI use in general", not the words earlier versions put in quotation marks; the privacy floor of 15 conversations or 5 accounts is not on the page, so it is unconfirmed and may come from the accompanying paper, which this survey has not read; and image generation is absent because Claude cannot generate images, not excluded by the index. The earlier reads, like version 0.2's, used a fetch tool that returns a model-written extract of a page rather than the page itself. The page gives no date range for the conversations it analysed, so the observation-window cell is blank. That later reports widened the scope is the claim of the first review note, notes/review-2026-09-26-first-articles.md, which says later reports include first-party API traffic and that the June 2026 report separates chat, Cowork and API usage; it gives no URL, and this survey has read no later report, so the row describes the introductory report only.

Cloudflare AI Crawl Control. Versions 0.1 and 0.2 listed "the purpose behind a crawl (training vs a live agent fetch)" among what this feature cannot see. notes/R-0001.md read the page as raw HTML on 26 September 2026 (last updated 23 April 2026) and found that it lets a site owner group AI-crawler traffic by crawler, by operator and by "Category: Analyze crawlers by their purpose or type" (observation). The feature therefore attributes traffic to AI operators, for crawlers Cloudflare has identified, and sorts it by each crawler's stated purpose. What it still cannot say is why any single page was fetched, because the category belongs to the crawler, not to the request (interpretation, following the review). No fetch was made for this correction; the page wording above is the review's.

OpenAI, "How People Use ChatGPT." Every figure in this row traces to a WebSearch summary of press coverage and the paper's own abstract, not a direct read of the NBER PDF (which this study's fetch tool could not parse as text) or of the companion openai.com page (which returned HTTP 403). The "10% of the global adult population" figure in particular is press commentary on the paper, not a number this study confirmed in the primary text; treat it as a secondary claim, not a derived measurement, until someone reads the primary source.

Hugging Face Hub. Only the models-the-hub page was fetched directly, and its returned content was too thin to source the download-concentration figures against; those came from a WebSearch summary citing a Hugging Face blog post and a download-analytics doc page, neither fetched directly here. A future check should read those pages directly before citing the 85.6%/80% figures as more than a secondhand report.

GitHub Security Advisories and npm download-count API. Both rows above were confirmed by a direct page fetch after an earlier WebSearch-only pass; they are accordingly more reliable than the OpenAI and Hugging Face rows, which rest on search summaries alone.

Excluded: candidates considered and not given a row

Candidate Reason for exclusion
Kephart and White, computer-virus epidemic models (S1) An academic model/paper on infection thresholds, not a standing observatory with a population, boundary or cadence of its own to record.
Ray, Tierra digital-organism paper (S2) A historical account of one contained experiment, not a live data source.
Anderson and Moore, economics of information security (S3) An academic paper on incentives; nothing to catalogue as a source's population or cadence.
Clymer, Wijk and Barnes, rogue replication threat model (S4) A threat-model essay with stated author disagreement about likelihood, not an observational instrument.
OpenTelemetry semantic conventions for generative AI (S6) A vocabulary/specification for other systems' telemetry; it produces no data of its own to survey.
Kessler and Cour-Palais, orbital debris (S8) The physics paper behind Lighthouse's own Kessler analogy; not a computational observatory.
The Menlo Report (S9) Research-ethics guidance, not a data source.
Harbour's own dispatch, proxy and feedback records (S12) Lighthouse's close-observation subject (LH F01's "nearest star"), already covered by registers/instruments.md's I-0001 and by LH000/LH001, not an outward third-party observatory; folding it in here would blur the close/outward distinction LH F01 draws.
Harbour's papers standard (S13) A publication-format standard, not a data source.
The xz backdoor oss-security disclosure post (S14) A first-hand account of one incident, feeding LH004's case study; not a standing observatory with an ongoing population or cadence.
The npm left-pad removal blog post (S17) A single incident account, also an LH004 alternative; not a standing instrument.
Stanford HAI AI Index Report Read directly on 26 September 2026 (https://hai.stanford.edu/ai-index): its own text says it "aggregates existing sources rather than conducting primary data collection." Giving it a row would double-count the primary sources already listed above rather than add a new one.
Air Street Capital, State of AI Report Known by reputation as a similarly annual secondary synthesis of other sources; not fetched for this survey, so no row is given rather than describing its cells from memory.
Shodan and Censys (internet-wide scanning) A 26 September 2026 search read found both gate most data behind free tiers with limited queries, reserving bulk or richer datasets for paid plans; neither is AI-specific, and confirming an exact population and boundary would need a paid query this reading-only study did not make.
SimilarWeb- and Ahrefs-style traffic estimators Proprietary estimation models whose methodology is not published in enough detail to fill this table's cells honestly; not fetched for this survey.
W3Techs technology-usage surveys Surveys web-server and CMS technology adoption generally, not AI activity specifically; not fetched for this survey.
Stack Exchange / Stack Overflow data dump A 26 September 2026 search read (https://archive.org/details/stackexchange; https://devclass.com/2024/07/30/stack-exchange-restricts-access-to-dump-of-user-contributed-data-as-critics-complain-license-permits-reuse-for-any-purpose/) found it CC BY-SA 4.0 licensed, formerly mirrored quarterly on the Internet Archive, with public access moved behind a login in mid-2024. It records human Q&A activity; any decline in it is at most an unconfirmed, indirect signal of AI's effect on this layer, not a direct observation of AI activity, so it is excluded rather than misrepresented as one.

Every fetch attempted for this study, 26 September 2026

Recorded here so a later check can see what was tried and what failed, independent of which row it fed.

URL Result
https://www.gharchive.org/ Read
https://developers.cloudflare.com/radar/ Read
https://developers.cloudflare.com/ai-crawl-control/features/analyze-ai-traffic/ Read
https://docs.softwareheritage.org/ Read
https://docs.deps.dev/ Read
https://openrouter.ai/rankings Read
https://nvd.nist.gov/ Failed: page returned only a bare "NVD - Home" header, no body text, to the fetch tool
https://www.anthropic.com/economic-index Failed in substance: page returned only a "Loading data" placeholder
https://www.anthropic.com/research/the-anthropic-economic-index Read
https://osv.dev/ Read
https://commoncrawl.org/ Read
https://epoch.ai/data Read
https://incidentdatabase.ai/ Read
https://pypistats.org/about Read
https://huggingface.co/docs/hub/models-the-hub Read, but content thin
https://openai.com/index/how-the-world-is-putting-chatgpt-to-work/ Failed: HTTP 403
https://cdn.openai.com/pdf/a253471f-8260-40c6-a2cc-aa93fe9f142e/economic-research-chatgpt-usage-paper.pdf Failed: fetch tool could not extract text from the PDF binary
https://nvd.nist.gov/general/News/change-timeline Failed: same bare-header problem as the NVD home page
https://www.cve.org/About/Overview Failed: same bare-header problem
https://hai.stanford.edu/ai-index Read (redirected from https://aiindex.stanford.edu/report/)
https://docs.github.com/en/code-security/security-advisories/working-with-global-security-advisories-from-the-github-advisory-database/about-the-github-advisory-database Read
https://github.com/npm/registry/blob/main/docs/download-counts.md Read
https://www.anthropic.com/research/the-anthropic-economic-index (again, for version 0.2; not counted among the twenty-two above) Read: page dated "Feb 10, 2025"; no date range given for the conversations analysed

Several rows above also draw on WebSearch result summaries rather than a direct page fetch (the OpenAI paper's population figures, the GHSA licence claim before the direct-fetch confirmation, the Hugging Face download-concentration figures, the Shodan/Censys and Stack Exchange exclusion reasons, the CNA/NVD split). Each such case is flagged in its row or note above; a WebSearch summary is a secondary read of the underlying page, not a primary one, and is recorded as such rather than presented as equivalent to a direct fetch.

From studies/LH003/sources.md in the repository, last changed 26 September 2026.