Methodology
On this page
What was counted, over which filings, by what rule, and how often it is wrong. Every figure below comes out of the run's own files. Last updated 29 September 2026.
The run
| Run id | tariff-universe-2026-09-r5 |
|---|---|
| Generated | 2026-09-29T13:59:16Z |
| Filing window | 2025-09-18 to 2026-09-17, 12 months |
| Filers in scope | 6,575 |
| Comparable filers | 4,551 |
| Source | SEC EDGAR full-text filings and the submissions API |
| Named on this site | 3,090 filers for tariff exposure (288 first time, 2,802 carried over) |
| Measured error rate | 2 of 300 named companies wrong (95% interval 0.2% to 2.4%); about 1 in 10 tariff-risk paragraphs missed; checked by a language model |
The three counts
Every family page publishes the same three counts on the same denominator: the comparable filers, meaning companies with both a 10-K in the window and a preceding one to compare it with. A company with only one 10-K in the record has nothing to compare and is not counted.
For tariff exposure in this run, 333 comparable filers name it for the first time this year, 3,731 name it as a risk, and 3,841 mention it. Of those that name it, 3,090 are named on this site (288 first time, 2,802 carried over); the others are withheld until both checks agree.
Mention
The family's keywords appear somewhere in the filer's Item 1A. A regular expression decides it and it costs nothing. It is the widest count, published beside the others so the difference shows: mentioning tariffs in a list of things that could go wrong is not naming tariff exposure as a risk to the business.
Names it as a risk
At least one paragraph of the filer's Item 1A crossed this family's threshold. This run put one family (tariff exposure) to a language model; the other thirteen carry keyword counts only. Each paragraph is asked a yes-or-no question of the form "does this paragraph name <family> as a risk to this company". The model returns a probability, and the threshold below turns it into a yes or a no. A company's result is the union over its paragraphs.
First time this year
Found in this year's Item 1A and not found in the prior filing's, where the prior year is decided at its own bar. For tariffs, a prior paragraph also has to use tariff wording (tariffs, duties, trade barriers, trade policy, protectionism and similar) to count: the model's broad question also scores sanctions and export-control paragraphs, and those alone do not make a tariff risk last year. The model is never asked whether something is new: newness is arithmetic on the two years' results, so a single model error changes one company's status rather than the finding.
Where a new label has a close textual match in the prior year's filing, the matching paragraph is carried on the row as evidence. That is the false positive the validation read looks for: a company that appended one sentence to old boilerplate has not newly disclosed anything. Labels are never inherited across such a match.
Names its own exposure
Not measured in this run; when it is, it will be the narrowest count. The question asks whether a paragraph names the company's own exposure rather than describing the risk in general.
The thresholds
A current tariff paragraph counts when the original question scores at least 0.30 and a stricter question scores at least 0.70. Last year counts only if a paragraph scores at least 0.30 and uses tariff wording (tariffs, duties, trade barriers, trade policy, protectionism and similar).
The rule as the run records it: tariff present = orig>=0.30 AND strict>=0.70.
| Family | Threshold |
|---|---|
| Tariff exposure | 0.3 |
The prior year is decided at 0.3. A family counts as new only when the prior filing fails to reach that bar, which makes "first time this year" harder to claim rather than easier.
The families in the pack:
- AI dependency Keyword count only
- Climate Keyword count only
- Customer concentration Keyword count only
- Cyber incident Keyword count only
- Financing Keyword count only
- Geopolitical Keyword count only
- Going concern Keyword count only
- Intellectual property Keyword count only
- Key person Keyword count only
- Labor Keyword count only
- Litigation Keyword count only
- Regulatory Keyword count only
- Supply chain Keyword count only
- Tariff exposure Read by the model
The question packs
The exact wording of every question is fixed in a pack file, and a run records the hash of each one it used. Two runs with the same hashes asked the same questions; a changed hash is a changed question and the counts either side of it are not comparable.
| Pack | SHA-256 |
|---|---|
| pack-specific.json | 064b3782d29f75aeed7bc73bafbdb2780ada02b4914f535b7c71e522f19ac0e4 |
| pack-tariff-only.json | 174d89bfd133d5a0a427611a8bf138a3e45e55f1133c6c1ae89e7d54a351f4e7 |
| pack-tariff-strict.json | 797272ea9ffab53b6fe31f2bac5d76caf76887f144e3b4f072034bfae5b06330 |
| pack.json | 858196128798b1a78494881cbc302862a0d0766f88cce859d72dcb1cb6812a14 |
The free keyword baseline is a separate list, hashed the same way: 41600159af084240371863d5079bfb236bb1ff3fc0c3aa8ef8d8983b21f38445.
Which filings
The most recent annual report on Form 10-K for each filer, filed in the window above, one per company. Amendments (10-K/A) and transition reports (10-KT) are excluded, because an amendment restates part of a filing and would double-count the company. Each of those filings is compared with the same company's preceding 10-K, and both accession numbers appear on the filer's page.
Which text, and how it is cut
Item 1A, "Risk Factors", only. No other part of the filing is read, and nothing outside the filing is used: no analyst note, no news, no press release, no vendor feed. In the dataset every paragraph carries its sequence number within that section, so it can be found again in the original.
How wrong it is
- Measured error rate
- 2 of 300 named companies wrong (95% interval 0.2% to 2.4%); about 1 in 10 tariff-risk paragraphs missed; checked by a language model
| Named companies read in full | 300 |
|---|---|
| Named companies wrong | 2 of 300 (95% interval 0.2% to 2.4%) |
| Tariff paragraphs missed | About 1 in 10 |
How it is measured, and by whom:
- Companies. A random sample of the companies the rule names, each read in full at its strongest paragraph and compared with the label.
- Paragraphs. A held-out sample never used to write the questions, drawn by stratum so that the paragraphs the rule rejects but most likely should have accepted are read too.
- Who reads. These reads are made by a language model, not by a person. Separately, before any company is named, every claim is checked by code against the filing on EDGAR and two independent language models read both years' risk factors without seeing our label; a company is named only when both agree and every check passes. Disputed claims are withheld.
Until these figures clear the quality gate, the finding is published as a sample and nothing is offered for sale.
Source, reuse and access
Reuse terms
The Commission's website policy states that information presented on sec.gov is public information and may be copied or further distributed by users of the site without the Commission's permission, with attribution requested and its seal, logos and registered marks excluded. That policy is at sec.gov/about/privacy-information, under "Website Dissemination".
So the filings themselves are free and yours to take. What is sold here is the reading of them: the count, the denominator, and the comparison with last year. The filings are the record and they cost nothing.
How the filings were fetched
Collection follows the Commission's automated access rules, published at sec.gov/search-filings/edgar-search-assistance/accessing-edgar-data: a declared user agent naming this site and a contact address, and a global rate of 4 requests a second against the Commission's stated ceiling of 10. Requests come from our own connection and are never routed through a rotating proxy, and a rate-limit response stops every worker for ten minutes rather than being retried around.
The organisation and contact address declared in that user agent are not published here.
What this cannot tell you
- It reads disclosure, not exposure. A company that is exposed and says nothing looks the same as a company that is not exposed.
- Risk-factor language is written by lawyers and moves for legal reasons. A new disclosure can mean a new risk, a new template, or new counsel.
- Timing follows fiscal years. Two companies filing eight months apart are describing different worlds, and a trailing-twelve-month window mixes them.
- The labels are machine-generated and wrong some of the time. Read the filing before you act.
What we do not claim
Stating the limits plainly is part of the method. Each line below is something a reader could reasonably assume and should not.
- The error rate covers this question on these filings, nothing else. It was measured by reading a sample of paragraphs from this run. It does not carry over to another question, another family, another year or another corpus.
- A claim was reviewed by machines, not audited. Every named claim was checked against the filing text by code and read by two language models before publication. No person reads the claims. That is review, not an audit, and no accountant has signed it.
- The probabilities are not calibrated. A score of 0.9 does not mean nine in ten such paragraphs are right. Scores order paragraphs; the published cut-offs were chosen from read samples, and only the measured precision and recall should be relied on.
- Families other than the one measured are data, not claims. A family is only published once it has been through the same reading and measurement.
- Nothing here is investment, legal or tax advice, and no forecast is implied. It is a count of what companies wrote in their own filings.
How to cite this dataset
The run id is the version. Two numbers from two runs are two measurements, so a citation that does not name the run cannot be checked. Suggested form:
risk.readevery.co. Risk factor disclosure, filings 2025-09-18 to 2026-09-17. Run tariff-universe-2026-09-r5. Generated 2026-09-29T13:59:16Z. https://risk.readevery.co/methodology/Quoting a filing instead? Cite the filing, not this site: the accession number and the sec.gov link are in the references at the foot of every filer page, and the Commission asks for attribution to itself rather than to whoever fetched the document.
Reproducing it
Every published result names its company, its central index key and the accession number of the filing it came from, and the dataset gives the paragraph within Item 1A, so any number here can be checked against the original on EDGAR, which is free. Questions about the method are welcome at [email protected]; corrections more so.