Object detection finds what is in an image and where it is. This guide explains backbones, region proposals, NMS, and mAP in plain terms, then covers the part most guides skip: how bounding box annotation rules, class taxonomy, and label quality decide what the model can actually learn.

Object detection is a computer vision task that finds every instance of a target object in an image or video frame and returns its position. For each detection, a model outputs four bounding box coordinates, a class label, and a confidence score between 0 and 1.

A classifier can tell you a photo contains a forklift. A detector tells you there are three forklifts, where each one is, and how sure it is about each. That difference is why detection sits underneath most production computer vision: counting, tracking, inspection, safety monitoring, and driver assistance all need coordinates, not just categories.

This guide covers how detectors work, how their accuracy is measured, and what it takes to train one on your own data. That last part gets the most space, because it is where most projects actually succeed or fail and where most explanations stop short.

Object detection takes an image and returns a list. Each entry in that list has the same shape: coordinates, a class, and a confidence value.

A single prediction looks like this in practice:

class: forklift
box: x=412, y=196, width=178, height=240
confidence: 0.91

Everything downstream is built on that list. Counting is the length of the list. Tracking is matching entries across frames. Zone alerts are geometry against the coordinates.

It helps to be blunt about what the model is doing. It has no concept of a forklift. It has learned a pattern of shapes, edges, textures and spatial relationships that appeared inside boxes labeled “forklift” in its training data, and it is matching regions of a new image against that pattern. IBM’s explainer makes the same point: detectors classify aggregates of visual properties, not objects.

Two properties matter for planning. First, detection is instance-level: ten cars produce ten predictions, not one “cars” label. Second, it is closed-set by default. A detector trained on 12 classes will never return a thirteenth, no matter how obvious the object is to you.

Diagram comparing image classification, object detection and instance segmentation on the same photo

These three tasks get confused constantly, and picking the wrong one is an expensive mistake because each needs a different kind of annotation.

Task Question answered Output Annotation needed Relative labeling cost
Image classification What is in this image? One label per image Tag per image Lowest
Object detection What is here and where? Box, class and confidence per instance Bounding box per object Low to moderate
Instance segmentation What is here and what is its exact shape? Pixel mask per instance Polygon or mask per object High
Semantic segmentation Which class does each pixel belong to? One class per pixel Full-scene pixel labeling Highest

Pick detection when position and count are enough. Pick segmentation when shape, area or boundary precision is the point, such as measuring a defect or separating touching objects. The accuracy and cost trade-off between the two is covered in detail in this comparison of bounding box and polygon annotation.

A common mistake is specifying boxes and then discovering the product needs area measurement. Re-labeling a dataset from boxes to polygons is not an adjustment. It is a new annotation project.

Flow diagram of the object detection pipeline from input image to final filtered detections

Every detector, from Faster R-CNN in 2015 to the transformer models shipping now, runs five stages. The architectures differ in how they handle stages 2 and 4.

Stage 4 is worth understanding because it has a direct operational consequence. In a crowded scene with genuinely overlapping objects, such as a dense parking lot or a shelf of stacked boxes, aggressive NMS deletes real detections. Loose NMS returns duplicates. There is no setting that is right for both, which is one reason newer models have moved away from NMS entirely.

The old one-stage versus two-stage split no longer covers the field. Query-based transformer detectors are a third family, and since 2025 they have led the accuracy benchmarks.

Family Examples How It Localizes Typical Strength Typical Weakness
Two-stage R-CNN, Fast R-CNN, Faster R-CNN Region proposal network, then per-region classification Strong localization, still a solid research baseline Slower, more complex to deploy
One-stage YOLO family, SSD, RetinaNet Dense prediction over anchors or feature map points in a single pass Fast, small, deployable on CPU and edge devices Historically weaker on small and overlapping objects
Query-based transformer DETR, RT-DETR, RF-DETR Fixed set of learned queries attending over the whole image Highest accuracy, no NMS needed, strong on crowded scenes Larger parameter counts, GPU-oriented

Current published benchmarks give a sense of the spread. RF-DETR is the first real-time detector to exceed 60 mAP on COCO, with its largest variant reporting 60.1 COCO AP at 17.2 ms and its L variant 56.5 AP at 6.8 ms, measured on an NVIDIA T4 with TensorRT at FP16. On the efficiency side, YOLO26 reports 40.9 to 57.5 mAP on COCO at 1.7 to 11.8 ms T4 latency.

Treat those numbers as a starting point for shortlisting, not as a prediction about your data. COCO has 80 everyday classes and large, well-lit objects. A model’s COCO ranking says little about how it will handle gray scale ultrasound frames or overhead warehouse footage. For a deeper architecture-by-architecture comparison, Hitech BPO maintains a benchmark roundup of object detection models that tracks the same leader-board figures.

The practical guidance is short. If you have GPU headroom and accuracy leads, start with a transformer detector. If you are deploying to CPU or an edge device, start with a small YOLO variant. Then stop tuning the architecture and go work on the dataset, which is where the remaining accuracy is hiding.

Annotation interface showing a tight bounding box drawn around a partially occluded object

Here is the part most explanations skip. A detector has no innate sense of where objects are. It learns coordinates by being shown thousands of examples of humans drawing boxes and adjusting its weights until its predicted coordinates match theirs.

That means the annotation is not preparation for the model. It is the specification of the model.

A single training label contains less than people expect:

That is all. The model never sees your intent, your edge case discussions, or the reasoning behind a judgment call. It sees the rectangle.

The three annotation formats

Format File Structure Coordinate System Used By
COCO JSON One JSON file for the whole dataset Absolute pixels: x, y, width, height from the top-left corner Detectron2, MMDetection, most research code
YOLO txt One .txt file per image Normalized 0 to 1: centre x, centre y, width, height Ultralytics YOLO models
Pascal VOC XML One .xml file per image Absolute pixels: xmin, ymin, xmax, ymax Older pipelines, some annotation tools

Conversion between them is routine, and every serious image annotation services workflow handles it as a matter of course. The failure mode is not the geometry, which converts cleanly. It is the class index mapping. A silent off-by-one in the class list produces a dataset where every label is wrong by one category, and the model trains happily on it. Spot-check a handful of converted files against the original class list before you kick off a training run.

These five decisions belong in your annotation guidelines before labeling starts. Each one is a fork, and the wrong choice is expensive to reverse because it means re-labeling.

Two operational practices catch problems in these areas before they become dataset-wide. A gold set of pre-answered items mixed into the queue measures accuracy continuously. Inter-annotator agreement, where the same items are labeled independently by two people and compared, measures whether the guidelines are actually unambiguous. Low agreement is a guidelines problem, not an annotator problem, and the fix is upstream.

I would add one observation from running annotation programs: teams consistently underestimate the calibration phase. The first two to four rounds of guideline revision, where the ML team and the annotation leads argue about edge cases, produce more accuracy improvement per hour spent than anything that happens later. Skipping that phase to start production faster is the single most common cause of a re-labeling pass.

Detection metrics all begin with Intersection over Union. IoU is the overlap area between a predicted box and a ground truth box, divided by their combined area. A prediction counts as correct only if its IoU clears a threshold and its class is right.

Diagram showing overlapping predicted and ground truth boxes with the intersection and union areas shaded
Metric What It Measures When to Look at It
IoU Overlap between predicted and true box Setting the correct threshold
Precision Of everything detected, how much was real When false alarms are costly
Recall Of everything real, how much was found When misses are costly, such as safety systems
AP Area under the precision-recall curve for one class Diagnosing per-class weakness
mAP@0.5 Mean AP across classes at a loose IoU threshold of 0.5 Did the model find the objects?
mAP@0.5:0.95 Mean AP averaged over IoU thresholds from 0.5 to 0.95 Are the boxes tight? This is COCO’s headline metric
AP_S / AP_M / AP_L AP split by object area: under 32², 32² to 96², over 96² Diagnosing small-object failure specifically

The COCO protocol defines AP at IoU 0.50:0.05:0.95 as the primary challenge metric, with AP small, medium and large computed over those pixel area bands.

The most useful diagnostic is the gap between mAP@0.5 and mAP@0.5:0.95. A model at 0.89 mAP@0.5 and 0.52 mAP@0.5:0.95 is finding almost everything but placing boxes loosely. Sometimes that is the model. Often it is the training data, because the model learned box tightness from annotations that were not tight. Before retraining with a bigger backbone, pull 100 training images and measure box tightness by hand against the objects. It takes an afternoon and frequently explains the gap.

One caution on reported numbers. The benchmark itself is not perfect. A 2024 ECCV study inspected thousands of COCO masks and found imprecise boundaries, non-exhaustively annotated instances and mislabeled masks, then released COCO-ReM as a corrected annotation set. The authors evaluated fifty detectors against it. Models producing visually sharper predictions scored higher on the corrected labels, which means they had been penalized by errors in COCO-2017 rather than by their own mistakes.

A 2026 survey in Artificial Intelligence Review reaches a broader version of the same conclusion. Across MS COCO, Open Images, Pascal VOC, LVIS, Object365 and DOTA, it documents missing labels, misaligned boxes, inconsistent occlusion treatment and ambiguous category definitions as systematic rather than occasional problems.

If the field’s flagship benchmarks carry that much label noise, your in-house dataset almost certainly does too.

A model at 0.92 mAP in validation can be useless in deployment. The failure patterns are consistent.

Four difficult detection scenarios including small objects, occlusion, truncation and low light

Small objects. After down-sampling, a 20-pixel object may occupy a single cell of the feature map. Higher input resolution, feature pyramid networks and tiled inference all help. So does a labeling policy that does not silently skip small objects.

Occlusion and crowding. Overlapping instances collide with NMS, as described earlier. Query-based detectors handle this better than anchor-based ones because they do not rely on suppression.

Class imbalance. If 94% of your boxes are one class, the model optimizes for that class and learns to ignore the rest. Rare-class recall will be poor no matter what the overall mAP says. Read per-class AP, not just the mean.

Domain shift. A model trained on daytime footage degrades at night, in rain, or on a different camera. This is a dataset composition problem. Training data needs to cover the conditions the system will actually meet, which is usually a data collection question before it is a labeling question.

Label noise. Missing labels are the most damaging kind, because an unlabeled object is not neutral. It is an explicit instruction that this appearance is background, and the model is penalized for detecting it correctly.

Temporal inconsistency in video. Frame-by-frame detection produces flickering identities across frames. If your product needs stable tracks, the annotation needs consistent object IDs across frames, which is a different and more demanding specification than per-frame boxes. Video annotation services handle this with interpolation between keyframes, and the tooling used has a large effect on both cost and ID consistency.

There is no formula, but there are workable planning ranges for fine-tuning a pretrained detector.

Scenario Images per class Notes
Proof of concept, visually distinct class 300 – 500 Enough to prove feasibility, not to deploy
Production, controlled environment 1,000 – 3,000 Fixed camera, stable lighting, limited variation
Production, variable real-world conditions 2,000 – 10,000 Multiple lighting conditions, angles, occlusion levels
Safety-critical or regulated 10,000+ Plus deliberate over-sampling of rare and dangerous cases

Two adjustments matter more than the headline number. Variation beats volume: 2,000 images spanning every lighting condition, camera angle and occlusion level outperforms 10,000 near-duplicate frames from one shift. And rare cases need over-representation relative to their natural frequency, because a class appearing in 0.5% of frames will be learned poorly if you sample naturally.

The budget follows from box count, not image count. HabileData’s published data annotation cost benchmarks put bounding box work at an indicative $0.03 to $0.12 per box. A 5,000-image pilot averaging eight objects per image is 40,000 boxes, roughly $2,000 of labeling at $0.05 per box, and around $2,700 once guideline development, QA and project management are included. Those figures are indicative planning ranges rather than quotes, and density is the variable that moves them most. A crowded retail shelf image with 40 objects costs eight times what a three-object image costs, at identical per-box rates.

Run a paid pilot of 500 to 2,000 representative items, including your ugliest edge cases, before committing to a full dataset. Score it against a gold set your own team labels independently. A pilot is cheap. A 40,000-box dataset built on misunderstood guidelines is not.

Partly, and not in the way the marketing suggests.

Open-vocabulary detectors accept free-text prompts instead of a fixed class list. Grounding DINO reports 52.5 AP on COCO without seeing any COCO images during training, and 63.0 AP on COCO test-dev after fine-tuning on COCO. YOLO-World-L reports 45.1 AP on COCO without COCO data. Those are genuinely impressive zero-shot numbers for general vocabulary.

They are also below what a fine-tuned detector achieves on a narrow domain, and the gap widens as your classes get further from everyday objects. A model that knows “car” and “person” from web-scale pretraining does not know your specific weld defect, your SKU packaging variant, or the difference between two visually similar surgical instruments.

Where these models earn their place is pre-labeling. The model drafts boxes, annotators correct them, and human time shifts from drawing to verifying. Pre-labeling with a reasonable model typically cuts human time by 30% to 70% on suitable tasks.

It carries a specific risk that is easy to miss. Annotators tend to accept the model’s suggestion, including its mistakes, because verifying is a lower-attention task than drawing. Systematic model errors then propagate into the ground truth and get reinforced in the next training round. The countermeasure is straightforward: seed gold-set items where the pre-label is deliberately wrong, and track how often annotators catch them. If catch rates fall, the review is becoming rubber-stamping.

Industry Application Typical Detection Challenge
Automotive Pedestrian, vehicle and sign detection Occlusion, night conditions, safety-critical recall
Manufacturing Defect detection, assembly verification Small defects, extreme class imbalance
Retail Shelf auditing, planogram compliance Dense scenes, visually similar SKUs
Healthcare Lesion and anomaly localization Expert annotation required, low tolerance for misses
Agriculture Crop, pest and yield monitoring Cluttered backgrounds, variable lighting
Logistics Package, pallet and container tracking Motion blur, identity consistency across frames
Security Intrusion detection, PPE compliance False alarm cost, wide camera variation

The pattern across these is consistent. The architecture choice is rarely the hard part. The hard part is a dataset that represents the deployment environment, labeled to rules that hold up under the messy cases.

Object detection is a well-solved problem at the architecture level. Pretrained models are open source, documented, and good enough that most teams can fine-tune something workable in a week.

What is not solved is the ground truth. Every coordinate a detector predicts is a reflection of coordinates a human drew, under rules someone wrote, checked by a process someone designed. When a detection model under-performs, the cause is more often in that chain than in the network.

So the sequence worth following is: define the classes a downstream decision actually needs, write the guidelines for occlusion, truncation and minimum size before labeling begins, pilot on a few hundred representative images, measure inter-annotator agreement, and only then scale. The architecture decision can wait. It is the easy one.

What is the difference between object detection and image classification?

Image classification assigns one label to a whole image. Object detection finds every instance of the target classes and returns a bounding box, class label and confidence score for each. Detection answers what and where; classification answers only what.

How does an object detection model know where an object is?

Through regression against human-drawn ground truth. During training the model compares its predicted coordinates with annotated boxes and adjusts weights to reduce the error. The annotated boxes are its only source of positional truth.

What is a good mAP score?

It depends entirely on the dataset. On COCO, the strongest real-time detectors currently sit around 55 to 60 mAP@0.5:0.95. On a narrow single-class industrial dataset, 0.90 mAP@0.5 is often reachable, and below 0.70 usually indicates a data problem rather than a model problem. Never compare mAP across different datasets.

Which object detection model should I use?

For GPU deployment where accuracy leads, start with a query-based transformer detector such as RF-DETR. For CPU or edge deployment, start with a small YOLO variant. Then spend your remaining effort on the dataset, which will move your accuracy further than switching between modern architectures.

How many images do I need?

Around 300 to 500 annotated images per class for a proof of concept on a visually distinct class. Production systems in variable conditions typically need 2,000 to 10,000 per class. Coverage of conditions matters more than raw volume.

Why does my model miss small objects?

Down-sampling in the backbone leaves small objects with very little signal in the feature map. Annotation contributes as well, because annotators skip tiny or blurry objects more often than large ones, and every skipped object trains the model to treat that appearance as background. Check AP_S separately from overall mAP.

What annotation format should I use?

Use the one your training framework expects natively. COCO JSON for Detectron2 and MMDetection, YOLO txt for Ultralytics models, Pascal VOC XML for older pipelines. Conversion is routine, but verify the class index mapping on a sample after every conversion.

How much does bounding box annotation cost?

Indicatively $0.03 to $0.12 per box, so a 40,000-box dataset runs roughly $1,200 to $4,800 in labeling, plus 30% to 40% for guidelines, QA and project management. Object density per image drives the total far more than image count does.

Can I use segmentation labels to train a detector?

Yes. A tight bounding box can be derived from a polygon or mask automatically. The reverse is not possible, which is why teams unsure about future requirements sometimes label polygons from the start and derive boxes as needed. That decision costs more up front and is worth weighing against the likelihood that you will actually need masks. The trade-offs between mask types are covered in this breakdown of semantic, instance and panoptic segmentation.

Planning a detection dataset and want the labeling spec right the first time?

Talk to our annotation team   »

Leave a Reply

Your email address will not be published.

Author Biju Peter

About Author

is a Senior Project Manager with 22+ years in the BPM industry, specializing in large-scale data operations and annotation-driven projects. He brings deep expertise in data processing, web research, scraping, and multi-modal annotation across image, text, audio, and video domains. He has successfully led 200+ projects, managed large teams, and delivered scalable, high-quality solutions for global AI and machine learning initiatives for clients across the globe. 🔗Connect with Biju on LinkedIn