Methodology
How this works, and what the data is
Everything a reviewer needs to decide whether the method is sound and where it would break. The uncomfortable parts are at the top rather than in a footnote.
The peer corpus is synthetic
Peer figures come from a synthetic corpus of 11,000 modelled policies. They are not market data: no carrier, filing, bordereau or real policy is represented here. The corpus is generated by a documented model (seed 20260826) that varies adoption, limits, premiums and deductibles by NAICS sector, revenue band and state, weighted toward Florida and the South Atlantic. It is built to be internally coherent and byte-for-byte reproducible so the analysis behaves the way it would on real data - but every peer statistic on this page is modelled. Replace this corpus with bordereau or rating-bureau data before quoting any figure to a client.
Corpus peer-corpus-2026.08-r1 · 11,000 generated policies · deterministic seed 20260826 · minimum cohort size 40 before widening.
01
Classification
A free-text description is matched against the Census NAICS index — the same index file the Census Bureau publishes to help a business self-classify. Matching is lexical and deterministic: no model call and no third-party API, and every match is traceable back to the index term that produced it, which is why the classification card shows those terms rather than a score alone.
NAICS codes
1,012
2022 vintage
Index terms
20,373
Searchable activity descriptions
Sectors covered
20
Two-digit NAICS sectors
Review threshold
55%
Below this, nothing proceeds
What confidence means here
Retrieval produces a raw score that is only comparable within one query — a long description matches more terms and scores higher regardless of how right it is. Confidence is that score normalised so it is comparable across queries, which is the only form in which a threshold means anything. A result under 55% is marked needsReview and the UI stops rather than benchmarking against a cohort it does not trust.
The runner-up matters as much as the winner. “We install docks and boat lifts” matches 237990 on Dock construction and Marine construction, and 238990 on Boat lift installation. Both are defensible and they sit in different peer cohorts, so the interface shows the margin and offers the override instead of presenting one code as settled.
Two numbers, and the held-out one is the real one
Accuracy measured only on the set used during development tells you how well a method was fitted, not how well it works. So the classifier is measured twice: once on the 184-case development set it was tuned against, and once on a 66-case held-out set written before the tuning and never inspected case-by-case afterwards.
Development set 84.8% top-1 · held-out set 86.4% top-1 · a gap of 1.6%.
A small gap means the method generalised rather than memorised. Quote the held-out number. Small buckets inside either set are noise and should not be quoted at all.
Classifier evaluation
held-out set (never tuned against)
Measured over 66 business descriptions the classifier was never tuned against.
Top-1 accuracy
86.4%
57 of 66 descriptions
The correct code ranked first
Top-3 accuracy
92.4%
61 of 66 descriptions
The correct code appeared in the first 3
Top-5 accuracy
92.4%
61 of 66 descriptions
The correct code appeared in the first 5
Eval set composition
66 descriptions across 9 buckets. Reporting a single headline accuracy over an unbalanced set hides which sectors the method is weak on, so the slice is shown alongside it.
| Bucket | Cases | Correct | Top-1 | |
|---|---|---|---|---|
| marine | 34 | 32 | 94.1% | |
| construction | 8 | 5 | 62.5% | |
| food_beverage | 5 | 3 | 60.0% | |
| healthcare | 4 | 4 | 100.0% | |
| professional | 4 | 4 | 100.0% | |
| logistics | 3 | 1 | 33.3% | |
| manufacturing | 3 | 3 | 100.0% | |
| personal_services | 3 | 3 | 100.0% | |
| retail | 2 | 2 | 100.0% |
Buckets under the overall 86.4% are marked. Those are where the index terms are thinnest.
Where it gets things wrong
A sample of failures, shown rather than summarised. Most are cases where the description names an activity that the Census index files under a different sector than a broker would expect.
| Description | Expected | Returned | Confidence |
|---|---|---|---|
| we move sailboats on a lowboy from the yard across town to the ramp | 484220 | 336612 | 44% |
| we deliver diesel by truck to boatyards, marinas and job sites | 457210 | 713930 | 93% |
| stucco and plaster contractor working custom homes | 238140 | 236115 | 73% |
| stump grinding and lot clearing for homeowners | 561730 | 113310 | 32% |
| sheet metal shop fabricating custom duct for mechanical contractors | 332322 | 332312 | 84% |
| drive thru coffee kiosk, two locations, no seating | 722515 | 722513 | 88% |
| ice plant, we make block and cube ice and deliver to convenience stores | 312113 | 457110 | 75% |
| flatbed fleet hauling steel and machinery over the road | 484230 | 484121 | 39% |
| last mile parcel contractor, we run 30 vans for a national carrier | 492110 | 484110 | 94% |
Run recorded 2026-08-31 18:23 UTC
Classifier evaluation
development set (tuned)
Measured over 184 hand-labelled business descriptions used during development.
Top-1 accuracy
84.8%
156 of 184 descriptions
The correct code ranked first
Top-3 accuracy
95.1%
175 of 184 descriptions
The correct code appeared in the first 3
Top-5 accuracy
97.3%
179 of 184 descriptions
The correct code appeared in the first 5
Eval set composition
184 descriptions across 12 buckets. Reporting a single headline accuracy over an unbalanced set hides which sectors the method is weak on, so the slice is shown alongside it.
| Bucket | Cases | Correct | Top-1 | |
|---|---|---|---|---|
| marine | 43 | 42 | 97.7% | |
| construction | 27 | 21 | 77.8% | |
| professional | 18 | 16 | 88.9% | |
| retail | 15 | 12 | 80.0% | |
| food_beverage | 14 | 11 | 78.6% | |
| healthcare | 14 | 11 | 78.6% | |
| manufacturing | 12 | 9 | 75.0% | |
| personal_services | 12 | 10 | 83.3% | |
| logistics | 11 | 9 | 81.8% | |
| real_estate | 7 | 5 | 71.4% | |
| technology | 6 | 6 | 100.0% | |
| auto | 5 | 4 | 80.0% |
Buckets under the overall 84.8% are marked. Those are where the index terms are thinnest.
Where it gets things wrong
A sample of failures, shown rather than summarised. Most are cases where the description names an activity that the Census index files under a different sector than a broker would expect.
| Description | Expected | Returned | Confidence |
|---|---|---|---|
| we write marine insurance for yacht owners, hull and liability | 524210 | 524126 | 61% |
| we do residential roofing and gutters, 12 guys, mostly insurance work | 238160 | 238170 | 69% |
| commercial refrigeration service for restaurants and grocery stores | 238220 | 811310 | 71% |
| interior and exterior painting crew, repaints and new construction | 238320 | 238130 | 53% |
| framing carpenters, we frame tract homes for a couple of builders | 238130 | 113110 | 30% |
| we build outdoor swimming pools, gunite shells and decks | 238990 | 561790 | 51% |
| handyman service, small repairs and odd jobs around the house | 236118 | 811411 | 79% |
| neighborhood breakfast and lunch spot with servers, no dinner | 722511 | 721191 | 21% |
| ice cream and frozen yogurt shop in a strip center | 722515 | 311520 | 68% |
| we make and bottle our own hot sauce and sell it wholesale | 311941 | 311421 | 54% |
| furniture showroom, sofas and dining sets | 449110 | 337126 | 55% |
| independent pharmacy filling prescriptions with a small front end | 456110 | 811114 | 84% |
| bike shop, we sell bicycles and do tune ups | 459110 | 811114 | 56% |
| management consulting, we do operations and strategy work | 541611 | 541614 | 63% |
| title agency handling closings and title searches | 541191 | 519290 | 54% |
| we own and rent out about 60 apartment units | 531110 | 236116 | 61% |
| family medicine practice, three doctors and two nurse practitioners | 621111 | 621399 | 59% |
| optometry practice, exams and we sell frames | 621320 | 611691 | 39% |
| outpatient counseling practice, licensed therapists only | 621330 | 621399 | 67% |
| refrigerated trucking, we run produce up the east coast | 484230 | 484220 | 42% |
Run recorded 2026-08-31 18:23 UTC
02
Extraction confidence
A declarations page is parsed in the browser. Each field carries its own confidence, the literal substring it was read from, and the page it appeared on — so a reviewer can check a number against the document without opening the document.
Signals
What goes into a field confidence
Per-field confidence combines the signals below. The weights live in the extractor; this page renders whatever it returns.
Label match
How closely the text preceding the value matches a known label for that field. An exact 'EACH OCCURRENCE LIMIT' scores far above a bare number in a column.
Value plausibility
Whether the parsed value is in range for its field and line. A $250 general liability occurrence limit is a parse error, not a cheap policy.
Layout proximity
Distance between the label and the value in the page's coordinate space. A value three columns away from its label is a weaker read than one immediately to its right.
Text-layer quality
Whether the page had an embedded text layer or had to be OCR'd. Every field on an OCR'd page is discounted, which is why the sample document's page 5 fields are the ones in the review queue.
The threshold is the same 55% used for classification. A field below it is flagged needsReview and enters the review queue with the reason it was flagged. The document-level figure is the weighted mean of the field confidences, so one bad page visibly drags it down instead of being averaged away.
03
Peer cohorts and benchmarking
A cohort is selected on NAICS code, revenue band and state. If fewer than 40 policies match, the cohort is widened — state is dropped first, then the code is rolled up to its sector — and the report says so on its face, along with what was dropped and what the sample size became. A benchmark computed on eleven peers and presented as a market rate is worse than no benchmark.
What is reported per line
- Adoption rate
- The share of the cohort carrying the line at all. This is the number that finds missing coverage, and it is independent of limits.
- Limit and premium percentiles
- Where the subject sits in the cohort's distribution, shown as the p10-p90 band with the interquartile range solid. A percentile without the spread behind it is not interpretable.
- Quartiles
- p10, p25, p50, p75 and p90 rather than a mean. Insurance limits cluster on round numbers and premiums are long-tailed; a mean of either is misleading.
- Gap severity
- Critical, material or advisory. Severity is assigned from the size of the gap and the exposure it leaves open, and each gap states the peer statistic it rests on so it can be argued with.
04
The proposed program
The comparison on the report reprices the program at peer-median limits using peer-median premiums from the same cohort. It is not a quote. Nothing has been submitted to a market, and no carrier has seen this risk.
How a proposed line is priced
Rank matching, stated plainly
The one assumption the model rests on, named rather than buried.
The cohort gives two separate distributions - what peers carry, and what peers pay - with nothing linking them. So the model assumes the peer at the kth percentile of limit sits at the kth percentile of premium. That assumption is the model. Reads between the quartile knots are linear; outside the 10th to 90th percentile the value is clamped rather than extrapolated, because premium tails are heavy.
What changes is decided from the findings above, not from adoption alone. That distinction matters in three real cases: a restaurant on a packaged policy is not missing general liability; charter liability sits at low adoption and is still the critical exposure on a chartered vessel; and a hull deductible is judged as a share of the insured value rather than in dollars.
The result is signed, and it is frequently positive. Of the five worked accounts, four cost more after the proposal - a dock contractor carrying no workers compensation while 92% of its peers do goes up $82,390. A tool that only ever finds savings is a discount calculator. This one will tell a client to buy more cover, which is the answer more often than not.
The comparison reports two axes, not one. Premium is what the program costs; total limits and lines carried are what it buys. A proposal that raises the premium and raises the cover is a different statement from one that raises the premium alone, and collapsing both into a single number hides which one happened. Total limits sum each line’s own limit — coverage lines do not stack in a claim, so it is a measure of cover in place, not one tower, and the report says so beside the figure.
Two things it will not do: price a line no peer in the cohort carries, and close a retention gap measured as a share of value when the peer dollar median equals the client’s own deductible. Both are left visibly open rather than filled with a number.
05
Putting real data behind it
Nothing above depends on the peer data being synthetic. The corpus sits behind one interface — a cohort query returning adoption rates and quartiles per line — so replacing it is a data-loading change, not an architectural one. Three routes, in order of how quickly they could be stood up.
Route A
FastestA broker's own book
An AMS export of in-force policies. Needs line, limit, deductible, premium, NAICS or SIC, revenue band and state per policy. Cohorts are only as wide as the book, so small classes fall below minimum size and widen — which the report already handles and discloses.
What blocks it: Cohort depth in narrow classes; PII handling; NAICS is often absent or stale in an AMS and has to be derived.
Route B
Slower, deeperAn industry data partner
A licensed benchmarking dataset gives national depth and real class breadth. The loader maps their schema onto the same cohort interface; nothing above the data layer changes.
What blocks it: Commercial terms and licence restrictions on redistributing quartiles to clients.
Route C
Narrow, high trustCarrier-provided aggregates
Some carriers will share class-level aggregates for classes they want to write. Highest credibility per number, but the coverage is patchy and the aggregates arrive pre-binned rather than as policies.
What blocks it: Coverage gaps between classes; pre-binned data cannot be re-cut into different revenue bands.
Where the swap actually happens
lib/ui/data.ts— the single seam every page reads through. Currently bound to liveDataProvenance— carried on every report and rendered at the bottom of it, so a report built on real data says so and a report built on synthetic data cannot quietly stop saying so.PeerCohort.widened— already models the case where real data is too thin, which is the case synthetic data never forces you to confront.
06
Limitations
- Peer numbers are generated, not observed. Every quartile, adoption rate and median premium in this prototype comes from a seeded generator. They are internally consistent and externally meaningless.
- Classification is lexical. It matches words, so a description that avoids the vocabulary of its own industry will score low — correctly, and it will route to review rather than guess.
- Limits are not the whole policy. Two accounts at the same limit can have materially different cover once forms, endorsements and exclusions are read. Benchmarking limits finds missing lines and thin towers; it does not read wording.
- Extraction handles declarations pages, not policies. A dec page is structured enough to parse. The forms behind it are not, and nothing here attempts them.
- A retention change is shown but not priced. Lowering a deductible normally costs more. The proposal reads premium from where the limit sits in the peer distribution, and pricing the retention as well would need a joint distribution of limit, deductible and premium that this corpus does not carry. The change is recommended; its premium consequence is left at zero rather than estimated.
- A premium nobody stated is not treated as zero. A declarations page often carries one premium for all sections and none per line. Where a line the program already holds is being changed and its own premium is unknown, the movement on that line cannot be computed, so it is left out of the totals rather than counted as the full new premium. The proposed total and the annual difference are then shown as lower bounds, with the number of excluded lines beside them.
- Severity is a heuristic. It encodes a defensible reading of exposure, not an underwriter’s judgement, and it does not know anything about the account that is not on this page.