Splittr

Where the splitter cuts in the wrong place — measured against the final package, which is what support left behind after fixing it

snapshot
Range

Split accuracy vs the final package

Share of uploads whose final page ranges exactly match what splittr proposed · dashed = trend, bars = uploads scored that day (right axis) · the final ranges reflect support remediation, so this is the splitter's real-world hit rate · the y-axis floors just below the worst day rather than at 0, so small moves are visible — read thin-volume days with care · starts

Package accuracy

Share of packages where every upload split cleanly — up is good, and the gap below 100% is the work support absorbs · one bad split anywhere in a package pulls a human in, so this sits below the per-upload rate because a package carries several uploads · bars = packages (right axis)

What changed after the split

Only the uploads that did not match · merged = final has fewer documents (splittr over-split), fragmented = more (under-split), reshaped = same count, different boundaries · the line is bad-split support tickets filed that day, an independent count of what the team actually hit — ticket dates lag the split, so read the two as trends, not row-for-row

Packages to look at

The most recent packages whose final splits differ from what splittr proposed · these are the ones a human had to fix, listed so they can be pulled up directly

Accuracy by upload size

Bigger uploads have more seams to get wrong — this is where to look for the next splitting win · counts are uploads we could compare against the final package

Accuracy by customer

Uploads split in the joinable window · a customer with low exact% is one whose documents the splitter finds hard — or whose support team re-cuts aggressively

Bad splits per package

How concentrated the damage is

Hard job failures

Only jobs that crashed and produced nothing — these are rare. The far more common failure is a job that succeeds but cuts in the wrong place: see bad splits above · error text is bucketed, never quoted

Strategy league table

Trailing window · ok% = job-level success · exact% = splits that survived unchanged (only uploads in the joinable window)

Split volume & strategy

Jobs per day stacked by the strategy that actually ran (failover included) · line = documents found per upload (right axis)

Throughput

Time per page (left axis, p50 and p95) and total split time per package (right axis) · per-page cost is the number that has to come down as uploads get bigger · per-package figures start

Gemini vs Infrrd

Model calls per day by provider (bars) with each provider's error rate (lines, right axis) · gemini is the primary splitter; infrrd is the fallback and the large-file path

Pages by engine

Which engine actually carved each page · infrrd runs a small share of the jobs but takes the big documents, so its page share runs far ahead of its job share

Provider reliability

Call-level, whole window

Failover paths

Jobs where the strategy that ran wasn't the one selected

Repair signals

Work the splitter did to fix itself mid-run · rising counts mean harder documents or a regression upstream

Model calls

Per provider & call type · p50 latency and error rate

Failure classes

Bucketed error classes — raw messages are never inlined