Sources
One source, named, with its refresh cadence and its reuse terms. Everything here is derived from it. Nothing is bought from a data vendor and nothing is scraped from another site.
Independence
This site is not affiliated with, endorsed by, or connected to the U.S. Securities and Exchange Commission or any government body. The Commission's registered marks, including the name of its filing system, are used here only to say where the documents came from.
The source
Annual reports on Form 10-K, filed with the U.S. Securities and Exchange Commission and published through its EDGAR system at sec.gov/search-filings. Each filing's primary document is read, along with the filer's submission history, which is where the standard industrial classification code and the date of the preceding 10-K come from.
Nothing else is used. No analyst notes, no news, no press releases, no vendor feed, no company questionnaire. If a company did not write it in its own annual report, it is not on this site.
How often it updates
- Four editions a year, on 15 March, 30 April, 31 July and 31 October. Each covers every filer's latest 10-K in the trailing twelve months, compared with its prior 10-K.
- This edition, September 2026, covers 10-Ks filed from 18 September 2025 to 17 September 2026.
- Every page states the run it was built from. A page carrying an older run id has not been refreshed yet; it has not been quietly changed underneath you.
Reuse terms
The Commission's website policy states that information presented on sec.gov is public information and may be copied or further distributed by users of the site without the Commission's permission, with attribution requested and its seal, logos and registered marks excluded. That policy is at sec.gov/about/privacy-information, under "Website Dissemination".
So the filings themselves are free and yours to take. What is sold here is the reading of them: the count, the denominator, and the comparison with last year. The filings are the record and they cost nothing.
How the filings were fetched
Collection follows the Commission's automated access rules, published at sec.gov/search-filings/edgar-search-assistance/accessing-edgar-data: a declared user agent naming this site and a contact address, and a global rate of 4 requests a second against the Commission's stated ceiling of 10. Requests come from our own connection and are never routed through a rotating proxy, and a rate-limit response stops every worker for ten minutes rather than being retried around.
The organisation and contact address declared in that user agent are not published here.
How the labels are derived
Of each filing, only Item 1A, the risk-factor section, is read. It is split into paragraphs. Every risk family gets a free keyword count. In this edition one family, tariff exposure, was also put to a language model: each paragraph is asked whether it names that family as a risk to this company, and a threshold, calibrated by reading sample rows on both sides of the line, turns the answer into a yes or a no. The other 13 families carry keyword counts only.
Both years' paragraphs are put to the same questions. A family counts as new for a company when it is found in this year's filing and not in last year's, where last year is judged at a lower bar and, for tariffs, only on paragraphs that use tariff wording (the method page gives both). The model is never asked whether something is new: that comparison is arithmetic, done afterwards.
The labels are machine-generated. The validation read behind the error rate is described on the methodology page.
- Measured error rate
- 2 of 300 named companies wrong (95% interval 0.2% to 2.4%); about 1 in 10 tariff-risk paragraphs missed; checked by a language model