published 2026-10-05 — eight measured head-to-head runs from our own benchmark harness, not a leaderboard of other people's COCO numbers.
The two permissively-licensed detector families everyone actually deploys are Roboflow's RF-DETR and Facebook's DETR ResNet-50, both Apache-2.0, and both publish COCO numbers that tell you almost nothing about a real domain. So we raced them head to head on the four detection domains we have measured so far — vehicles in traffic, people, sheep from the air, and birds — same held-out images, same matching rule, stock settings first and then after calibration. The short answer: stock, it is a coin flip; calibrated, DETR wins all four; and RF-DETR is five times faster, which sometimes buys it the decision anyway.
Each row is RF-DETR-Nano versus DETR ResNet-50 on the same held-out subset, scored at F1 (IoU 0.5). Both columns read stock → calibrated: stock is the default confidence threshold with no per-class tuning, calibrated is after sweeping per-class thresholds and NMS IoU on the held-out set. The winner column is decided on calibrated F1, because nobody ships stock. Latency is mean milliseconds per image for the fp32 models, RF-DETR first.
| domain | RF-DETR-N F1 | DETR-R50 F1 | calibrated winner | ms/img |
|---|---|---|---|---|
| vehicles in traffic | 0.588 → 0.653 | 0.597 → 0.663 | DETR | 325 vs 1606 |
| people | 0.706 → 0.765 | 0.666 → 0.768 | DETR (+0.003) | 325 vs 1606 |
| sheep | 0.644 → 0.697 | 0.708 → 0.782 | DETR | 422 vs 2126 |
| birds | 0.527 → 0.568 | 0.479 → 0.604 | DETR (+0.035) | 323 vs 1619 |
Vehicles: COCO val2017 vehicle subset (250 images, 971 boxes). People: COCO val2017 person subset (250 images, 926 boxes). Sheep: COCO val2017 sheep subset (65 images, 354 boxes). Birds: COCO val2017 birds subset (125 images, 427 boxes). All at 800 px letterbox, 1-core x86, runs of 2026-09-28, 2026-09-30 and 2026-10-01. One honest sizing note: RF-DETR-Nano is the smallest member of its family and DETR ResNet-50 is the smallest mainstream DETR, so this is the cheap-vs-cheap matchup, not a like-for-like capacity match. Full per-class tables are on the pack pages below.
At default settings each architecture wins two of the four domains: RF-DETR takes people (+0.040 F1) and birds (+0.048), DETR takes vehicles (+0.010) and sheep (+0.064). If you picked your model on stock numbers — which is what downloading a checkpoint and running it at threshold 0.5 amounts to — you would be right half the time and have no way of knowing which half. That is exactly the choice paralysis the COCO leaderboards create: two Apache-2.0 models, each credible, and no evidence for your domain.
After per-class threshold and NMS sweeps, DETR fp32 wins all four domains — and the flip is most extreme exactly where RF-DETR looked best. On birds, RF-DETR's 0.048 stock lead becomes a 0.035 calibrated deficit: its fixed 384 px input resolves fewer small birds in a flock, and no threshold buys back what the resolution cannot see. The people row is the instructive near-miss — 0.765 versus 0.768 is a three-thousandths gap. Calibration is also where DETR earns most of its lead: on birds it added +0.124 to its own stock F1, on people +0.103. The operating point you ship matters more than the checkpoint you downloaded.
RF-DETR-N runs 4.9–5.0× faster on the same single x86 core (325 ms versus 1606 ms per image on three of the four domains) and its INT8 export is near-lossless, giving up at most about a point of calibrated F1 across these domains, while DETR's INT8 costs another 3–5 points on top of its calibrated deficit. We sell the calibrated winner, so take this from us as the honest counter-case: if your edge box is a Jetson or a mini-PC with a frame budget, RF-DETR on people — a 0.003 F1 gap for five times the throughput — is a defensible purchase that DETR cannot answer. The pack page tells you which situation you are in, with the numbers under it.
The winners above are live packs — vehicles in traffic, people, sheep and birds — each with the full race table under it and the winner running client-side on your own photo, stock versus calibrated side by side, before you pay anything. Birds is the page to visit if you only check one: it is the domain where the stock leaderboard and the calibrated one disagree most.
The most-asked version of this comparison adds YOLOv8, and we leave it out of the table on purpose: we only benchmark architectures whose licence survives shipping in a commercial product, and Ultralytics' AGPL-3.0 does not — the licence guide covers what that means for a fine-tuned checkpoint. These subsets are COCO val2017, so the numbers are evidence about these two architectures and our pipeline, not about your factory line or your drone footage — the calibration report in every pack carries the same sweep run on the domain you actually ship.