published 2026-09-29 — four measured fp32/int8 pairs from our own benchmark runs, not a survey of other people's tables.
Everyone ships INT8 for the speed and the file size, and the usual answer to “what does it cost?” is a shrug: “a point or two of mAP, on average”. We stopped guessing and measured it. Same held-out images, same letterboxing, same matching rule, FP32 and INT8 side by side, on two detection domains and two Apache-2.0 architectures. The spread in the answer is the story: one model lost under 1% F1, the other lost 7–8%.
Each row is the same model at the same resolution, FP32 versus its published INT8 ONNX export, scored on a 250-image held-out subset. The F1 columns show stock → calibrated: stock is the vendor default confidence threshold, calibrated is after sweeping per-class thresholds and NMS IoU on the held-out set. Cost is the relative F1 drop of INT8 after calibration.
| domain | model | F1 fp32 | F1 int8 | int8 cost | speedup |
|---|---|---|---|---|---|
| vehicles in traffic | RF-DETR-Nano | 0.588 → 0.653 | 0.586 → 0.652 | −0.2% | 1.45× (325 → 224 ms) |
| vehicles in traffic | DETR ResNet-50 | 0.597 → 0.663 | 0.551 → 0.614 | −7.8% | 1.21× (1606 → 1332 ms) |
| people | RF-DETR-Nano | 0.706 → 0.765 | 0.699 → 0.759 | −1.1% | 1.45× (325 → 224 ms) |
| people | DETR ResNet-50 | 0.666 → 0.768 | 0.616 → 0.719 | −7.5% | 1.21× (1606 → 1332 ms) |
Vehicles: COCO val2017 vehicle subset (250 images, 971 boxes; car/motorcycle/bus/truck). People: COCO val2017 person subset (250 images, 926 boxes). 800 px letterbox, 1-core x86, runs of 2026-09-28. Latency is mean ms/image for the INT8 export versus its FP32 counterpart on the same core. Full per-class numbers are on the pack pages below.
Same recipe, same data, two architectures, and a 40× spread in relative accuracy loss. One INT8 export is near-lossless on both domains; the other gives up roughly seven points of F1 on both. We measured two models, so we will not claim a law — but whatever the pattern is, it is not “INT8 costs a point or two”. The only number that matters is the pair on your domain, and you cannot know it without running both.
Recalibrating thresholds on the INT8 outputs recovered almost nothing: the calibrated INT8 rows still trail their FP32 counterparts by about the same gap the stock numbers show. Threshold sweeps and NMS tuning buy real F1 (they bought 6–10 points on these very domains), but they buy it for both precisions. They move your operating point; they do not put back what the rounding took.
On a single x86 core at 800 px, the INT8 export ran about 1.45× faster for the nano model and 1.21× faster for the ResNet-50 — worth having, not the 2–4× that gets repeated. Much of the folklore comes from GPU tensor-core benchmarks or from comparing against unbatched FP32 with different thread counts. On a CPU at the edge, measure your own box before promising your manager a number.
Every yolomax pack contains both exports — FP32 and INT8 ONNX — plus the calibration report with exactly this table for the winning architecture, so the decision you are making today is the one we already measured. On vehicles in traffic and people you can also run the winner in your own browser on your own photo before paying a cent, which is the only benchmark we would trust either.
One honest caveat, the same one we put on every page: these subsets are COCO val2017, so the numbers are evidence about our pipeline and about these architectures, not about your factory line or your drone footage. If your domain is not one we have measured yet, that is what the calibration report in a pack is for — and if it is not in the catalogue at all, the licence guide covers why we build on Apache-2.0 architectures in the first place.