TBF Insure

Methodology

How this works, and what the data is

Everything a reviewer needs to decide whether the method is sound and where it would break. The uncomfortable parts are at the top rather than in a footnote.

The peer corpus is synthetic

Peer figures come from a synthetic corpus of 11,000 modelled policies. They are not market data: no carrier, filing, bordereau or real policy is represented here. The corpus is generated by a documented model (seed 20260826) that varies adoption, limits, premiums and deductibles by NAICS sector, revenue band and state, weighted toward Florida and the South Atlantic. It is built to be internally coherent and byte-for-byte reproducible so the analysis behaves the way it would on real data - but every peer statistic on this page is modelled. Replace this corpus with bordereau or rating-bureau data before quoting any figure to a client.

Corpus peer-corpus-2026.08-r1 · 11,000 generated policies · deterministic seed 20260826 · minimum cohort size 40 before widening.

01

Classification

A free-text description is matched against the Census NAICS index — the same index file the Census Bureau publishes to help a business self-classify. Matching is lexical and deterministic: no model call and no third-party API, and every match is traceable back to the index term that produced it, which is why the classification card shows those terms rather than a score alone.

NAICS codes

1,012

2022 vintage

Index terms

20,373

Searchable activity descriptions

Sectors covered

20

Two-digit NAICS sectors

Review threshold

55%

Below this, nothing proceeds

What confidence means here

Retrieval produces a raw score that is only comparable within one query — a long description matches more terms and scores higher regardless of how right it is. Confidence is that score normalised so it is comparable across queries, which is the only form in which a threshold means anything. A result under 55% is marked needsReview and the UI stops rather than benchmarking against a cohort it does not trust.

The runner-up matters as much as the winner. “We install docks and boat lifts” matches 237990 on Dock construction and Marine construction, and 238990 on Boat lift installation. Both are defensible and they sit in different peer cohorts, so the interface shows the margin and offers the override instead of presenting one code as settled.

Two numbers, and the held-out one is the real one

Accuracy measured only on the set used during development tells you how well a method was fitted, not how well it works. So the classifier is measured twice: once on the 184-case development set it was tuned against, and once on a 66-case held-out set written before the tuning and never inspected case-by-case afterwards.

Development set 84.8% top-1 · held-out set 86.4% top-1 · a gap of 1.6%.

A small gap means the method generalised rather than memorised. Quote the held-out number. Small buckets inside either set are noise and should not be quoted at all.

Classifier evaluation

held-out set (never tuned against)

Measured over 66 business descriptions the classifier was never tuned against.

Held out

Top-1 accuracy

86.4%

57 of 66 descriptions

The correct code ranked first

Top-3 accuracy

92.4%

61 of 66 descriptions

The correct code appeared in the first 3

Top-5 accuracy

92.4%

61 of 66 descriptions

The correct code appeared in the first 5

Eval set composition

66 descriptions across 9 buckets. Reporting a single headline accuracy over an unbalanced set hides which sectors the method is weak on, so the slice is shown alongside it.

BucketCasesCorrectTop-1 
marine343294.1%
construction8562.5%
food_beverage5360.0%
healthcare44100.0%
professional44100.0%
logistics3133.3%
manufacturing33100.0%
personal_services33100.0%
retail22100.0%

Buckets under the overall 86.4% are marked. Those are where the index terms are thinnest.

Where it gets things wrong

A sample of failures, shown rather than summarised. Most are cases where the description names an activity that the Census index files under a different sector than a broker would expect.

DescriptionExpectedReturnedConfidence
we move sailboats on a lowboy from the yard across town to the ramp48422033661244%
we deliver diesel by truck to boatyards, marinas and job sites45721071393093%
stucco and plaster contractor working custom homes23814023611573%
stump grinding and lot clearing for homeowners56173011331032%
sheet metal shop fabricating custom duct for mechanical contractors33232233231284%
drive thru coffee kiosk, two locations, no seating72251572251388%
ice plant, we make block and cube ice and deliver to convenience stores31211345711075%
flatbed fleet hauling steel and machinery over the road48423048412139%
last mile parcel contractor, we run 30 vans for a national carrier49211048411094%

Run recorded 2026-08-31 18:23 UTC

Classifier evaluation

development set (tuned)

Measured over 184 hand-labelled business descriptions used during development.

Tuned

Top-1 accuracy

84.8%

156 of 184 descriptions

The correct code ranked first

Top-3 accuracy

95.1%

175 of 184 descriptions

The correct code appeared in the first 3

Top-5 accuracy

97.3%

179 of 184 descriptions

The correct code appeared in the first 5

Eval set composition

184 descriptions across 12 buckets. Reporting a single headline accuracy over an unbalanced set hides which sectors the method is weak on, so the slice is shown alongside it.

BucketCasesCorrectTop-1 
marine434297.7%
construction272177.8%
professional181688.9%
retail151280.0%
food_beverage141178.6%
healthcare141178.6%
manufacturing12975.0%
personal_services121083.3%
logistics11981.8%
real_estate7571.4%
technology66100.0%
auto5480.0%

Buckets under the overall 84.8% are marked. Those are where the index terms are thinnest.

Where it gets things wrong

A sample of failures, shown rather than summarised. Most are cases where the description names an activity that the Census index files under a different sector than a broker would expect.

DescriptionExpectedReturnedConfidence
we write marine insurance for yacht owners, hull and liability52421052412661%
we do residential roofing and gutters, 12 guys, mostly insurance work23816023817069%
commercial refrigeration service for restaurants and grocery stores23822081131071%
interior and exterior painting crew, repaints and new construction23832023813053%
framing carpenters, we frame tract homes for a couple of builders23813011311030%
we build outdoor swimming pools, gunite shells and decks23899056179051%
handyman service, small repairs and odd jobs around the house23611881141179%
neighborhood breakfast and lunch spot with servers, no dinner72251172119121%
ice cream and frozen yogurt shop in a strip center72251531152068%
we make and bottle our own hot sauce and sell it wholesale31194131142154%
furniture showroom, sofas and dining sets44911033712655%
independent pharmacy filling prescriptions with a small front end45611081111484%
bike shop, we sell bicycles and do tune ups45911081111456%
management consulting, we do operations and strategy work54161154161463%
title agency handling closings and title searches54119151929054%
we own and rent out about 60 apartment units53111023611661%
family medicine practice, three doctors and two nurse practitioners62111162139959%
optometry practice, exams and we sell frames62132061169139%
outpatient counseling practice, licensed therapists only62133062139967%
refrigerated trucking, we run produce up the east coast48423048422042%

Run recorded 2026-08-31 18:23 UTC

02

Extraction confidence

A declarations page is parsed in the browser. Each field carries its own confidence, the literal substring it was read from, and the page it appeared on — so a reviewer can check a number against the document without opening the document.

Signals

What goes into a field confidence

Per-field confidence combines the signals below. The weights live in the extractor; this page renders whatever it returns.

  • Label match

    How closely the text preceding the value matches a known label for that field. An exact 'EACH OCCURRENCE LIMIT' scores far above a bare number in a column.

  • Value plausibility

    Whether the parsed value is in range for its field and line. A $250 general liability occurrence limit is a parse error, not a cheap policy.

  • Layout proximity

    Distance between the label and the value in the page's coordinate space. A value three columns away from its label is a weaker read than one immediately to its right.

  • Text-layer quality

    Whether the page had an embedded text layer or had to be OCR'd. Every field on an OCR'd page is discounted, which is why the sample document's page 5 fields are the ones in the review queue.

The threshold is the same 55% used for classification. A field below it is flagged needsReview and enters the review queue with the reason it was flagged. The document-level figure is the weighted mean of the field confidences, so one bad page visibly drags it down instead of being averaged away.

03

Peer cohorts and benchmarking

A cohort is selected on NAICS code, revenue band and state. If fewer than 40 policies match, the cohort is widened — state is dropped first, then the code is rolled up to its sector — and the report says so on its face, along with what was dropped and what the sample size became. A benchmark computed on eleven peers and presented as a market rate is worse than no benchmark.

What is reported per line

Adoption rate
The share of the cohort carrying the line at all. This is the number that finds missing coverage, and it is independent of limits.
Limit and premium percentiles
Where the subject sits in the cohort's distribution, shown as the p10-p90 band with the interquartile range solid. A percentile without the spread behind it is not interpretable.
Quartiles
p10, p25, p50, p75 and p90 rather than a mean. Insurance limits cluster on round numbers and premiums are long-tailed; a mean of either is misleading.
Gap severity
Critical, material or advisory. Severity is assigned from the size of the gap and the exposure it leaves open, and each gap states the peer statistic it rests on so it can be argued with.

04

The proposed program

The comparison on the report reprices the program at peer-median limits using peer-median premiums from the same cohort. It is not a quote. Nothing has been submitted to a market, and no carrier has seen this risk.

How a proposed line is priced

Rank matching, stated plainly

The one assumption the model rests on, named rather than buried.

The cohort gives two separate distributions - what peers carry, and what peers pay - with nothing linking them. So the model assumes the peer at the kth percentile of limit sits at the kth percentile of premium. That assumption is the model. Reads between the quartile knots are linear; outside the 10th to 90th percentile the value is clamped rather than extrapolated, because premium tails are heavy.

What changes is decided from the findings above, not from adoption alone. That distinction matters in three real cases: a restaurant on a packaged policy is not missing general liability; charter liability sits at low adoption and is still the critical exposure on a chartered vessel; and a hull deductible is judged as a share of the insured value rather than in dollars.

The result is signed, and it is frequently positive. Of the five worked accounts, four cost more after the proposal - a dock contractor carrying no workers compensation while 92% of its peers do goes up $82,390. A tool that only ever finds savings is a discount calculator. This one will tell a client to buy more cover, which is the answer more often than not.

The comparison reports two axes, not one. Premium is what the program costs; total limits and lines carried are what it buys. A proposal that raises the premium and raises the cover is a different statement from one that raises the premium alone, and collapsing both into a single number hides which one happened. Total limits sum each line’s own limit — coverage lines do not stack in a claim, so it is a measure of cover in place, not one tower, and the report says so beside the figure.

Two things it will not do: price a line no peer in the cohort carries, and close a retention gap measured as a share of value when the peer dollar median equals the client’s own deductible. Both are left visibly open rather than filled with a number.

05

Putting real data behind it

Nothing above depends on the peer data being synthetic. The corpus sits behind one interface — a cohort query returning adoption rates and quartiles per line — so replacing it is a data-loading change, not an architectural one. Three routes, in order of how quickly they could be stood up.

  1. Route A

    Fastest

    A broker's own book

    An AMS export of in-force policies. Needs line, limit, deductible, premium, NAICS or SIC, revenue band and state per policy. Cohorts are only as wide as the book, so small classes fall below minimum size and widen — which the report already handles and discloses.

    What blocks it: Cohort depth in narrow classes; PII handling; NAICS is often absent or stale in an AMS and has to be derived.

  2. Route B

    Slower, deeper

    An industry data partner

    A licensed benchmarking dataset gives national depth and real class breadth. The loader maps their schema onto the same cohort interface; nothing above the data layer changes.

    What blocks it: Commercial terms and licence restrictions on redistributing quartiles to clients.

  3. Route C

    Narrow, high trust

    Carrier-provided aggregates

    Some carriers will share class-level aggregates for classes they want to write. Highest credibility per number, but the coverage is patchy and the aggregates arrive pre-binned rather than as policies.

    What blocks it: Coverage gaps between classes; pre-binned data cannot be re-cut into different revenue bands.

Where the swap actually happens

  • lib/ui/data.ts — the single seam every page reads through. Currently bound to live
  • DataProvenance — carried on every report and rendered at the bottom of it, so a report built on real data says so and a report built on synthetic data cannot quietly stop saying so.
  • PeerCohort.widened — already models the case where real data is too thin, which is the case synthetic data never forces you to confront.

06

Limitations

  • Peer numbers are generated, not observed. Every quartile, adoption rate and median premium in this prototype comes from a seeded generator. They are internally consistent and externally meaningless.
  • Classification is lexical. It matches words, so a description that avoids the vocabulary of its own industry will score low — correctly, and it will route to review rather than guess.
  • Limits are not the whole policy. Two accounts at the same limit can have materially different cover once forms, endorsements and exclusions are read. Benchmarking limits finds missing lines and thin towers; it does not read wording.
  • Extraction handles declarations pages, not policies. A dec page is structured enough to parse. The forms behind it are not, and nothing here attempts them.
  • A retention change is shown but not priced. Lowering a deductible normally costs more. The proposal reads premium from where the limit sits in the peer distribution, and pricing the retention as well would need a joint distribution of limit, deductible and premium that this corpus does not carry. The change is recommended; its premium consequence is left at zero rather than estimated.
  • A premium nobody stated is not treated as zero. A declarations page often carries one premium for all sections and none per line. Where a line the program already holds is being changed and its own premium is unknown, the movement on that line cannot be computed, so it is left out of the totals rather than counted as the full new premium. The proposed total and the annual difference are then shown as lower bounds, with the number of excluded lines beside them.
  • Severity is a heuristic. It encodes a defensible reading of exposure, not an underwriter’s judgement, and it does not know anything about the account that is not on this page.

Back to the benchmark