Video annotation cost runs $0.02 to $0.15 per frame for bounding boxes and $20 to $150 per minute at 30 fps. Price is set by annotated frame rate, object density, interpolation leverage and QA depth. This guide gives the arithmetic behind per frame annotation pricing so you can rebuild any quote.

Search for video annotation pricing and you will find per-minute quotes ranging from roughly $0.50 to $150. A 300x spread, for what reads on every website as the same service.

Buyers usually assume someone is padding. In my experience scoping these projects, that is almost never what is going on. The spread comes from one variable that sits underneath both pricing models and that hardly any quote states out loud.

Below is the arithmetic. Once you have it, you can rebuild any quote you receive and see what you are actually being sold.

Video annotation costs roughly $0.02 to $0.15 per annotated frame for bounding boxes, or $20 to $150 per minute at 30 fps with interpolation. Segmentation runs 5 to 10 times higher. Price is set by object density, seconds per object, how many frames a human actually touches, and QA depth.

An image quote has one unit. A video quote has three competing ones, and vendors pick whichever flatters them.

Video also carries a cost that images do not: temporal identity. Every object needs the same track ID in frame 900 as in frame 1. That constraint, not the drawing, is what makes video expensive. If the annotation types themselves are still unfamiliar, our complete guide to video annotation for machine learning covers the fundamentals, and it is worth reading before you price anything.

Any per-frame or per-minute price is the same four inputs wearing a trench coat:

Annotator hours = (object instances x seconds per instance) / 3600 / interpolation leverage

Delivered cost = Annotator hours x loaded hourly rate x (1 + QA overhead)

Input What it means Typical range
Object density Labeled objects per frame 1 to 30+
Seconds per instance Draw or adjust time, including track ID 2s (adjust) to 18s (fresh box, dense scene)
Interpolation leverage Frames delivered per human-touched frame 1x to 8x
Loaded hourly rate Annotator + QA + supervision, blended $6 to $12 offshore, $25 to $60 onshore

Seconds per instance is the input buyers get most wrong, and the published research is worth knowing before you assess a vendor’s throughput claim.

The reference protocol used to annotate ImageNet reports median times of 25.5 seconds to draw a box, 9 seconds to verify it, and 7.8 seconds to check whether other objects of the same class remain. The annotation literature treats 34.5 seconds as the standard reference for one high-quality box, rising to 55 seconds once rejected boxes are redrawn. Other reviews of the same literature put the median anywhere between 7 and 35 seconds, depending on image quality and how tight the box has to be.

So when a vendor implies 5 to 8 seconds per box, is that fantasy? No, and the reason matters. Those research figures are crowd workers meeting an unfamiliar image with a blank canvas, and most of that time is visual search rather than drawing. Give an annotator a proposed box to correct instead, and the University of Amsterdam’s street-scene study measured a median 8.9 seconds per box adjustment, falling to 4.4 seconds by the third pass as annotators internalized the ontology.

That gap between drawing and adjusting is the entire economic argument for interpolation, and it is why the same footage can be quoted at wildly different prices without anyone lying.

Ask a vendor for those four numbers. One who cannot produce them is guessing, and you will meet the difference later as a change order.

Where the model breaks down: temporal annotation. Action and event labeling is priced by decision, not by object. An annotator watching for the exact start frame of a fall or a near-miss spends most of their time scrubbing, not labeling, and object density tells you nothing about that. Price temporal work per hour or per event, and treat any per-frame quote for it with suspicion.

Per frame annotation pricing: how it works

You pay per annotated frame, regardless of source duration.

Common in autonomous driving and any dataset assembled from extracted frame sequences rather than continuous video. It suits work where each frame is labeled independently and interpolation would introduce error, such as low frame rate footage or high-action sequences.

What it gets right: the unit matches the labor when object density is stable.

Where it fails: a frame holding one parked car and a frame holding 30 pedestrians cost the same on paper and differ 30x in reality. Per-frame pricing without a stated density band hands the vendor an unpriced option.

Pin the quote to a density ceiling. Something like $0.06 per frame up to 8 objects, tiered above that.

Per minute video annotation cost: how it works

You pay per minute of source footage. The vendor absorbs frame count and interpolation strategy.

This works for long continuous sequences with predictable motion, where keyframe annotation plus interpolation does most of the work. Surveillance, sports and retail footage usually fit.

The variable that explains the 300x spread sits in the frame count. Hold the rate constant at $0.03 per frame:

Annotated frame rate Frames per minute Cost per minute
1 fps 60 $1.80
5 fps 300 $9.00
15 fps 900 $27.00
30 fps 1,800 $54.00
Per Minute Video Annotation Cost Frame Sequence

Same rate. Same quality. A 30x price difference produced entirely by annotation density in time.

So when one vendor quotes $8 a minute and another quotes $95, they are rarely competing. The first is likely annotating 2 to 5 frames per second. The second is annotating all 30. Buy the first for a tracking model and you will hand it footage with nothing to learn temporal continuity from.

A per-minute quote without a stated annotated frame rate is not really a quote. It is a number.

Ranges below assume offshore delivery, trained annotators, two-stage QA, and no rush premium. They reflect quotes we see in the market rather than a published survey, so treat them as a plausibility check, not a rate card.

Task type Pricing model Cost range Assumptions
Bounding box, sparse Per frame $0.02 to $0.04 1 to 3 objects, clear scene, interpolation used
Bounding box, dense + track IDs Per frame $0.05 to $0.15 5 to 15 objects, occlusion handling, MOT-format output
Object tracking Per minute $20 to $60 5 to 10 fps annotated, predictable motion, ID persistence
Object tracking, full rate Per minute $80 to $180 30 fps, dense urban or crowd scenes
Polygon / semantic segmentation Per frame $0.20 to $1.20 20 to 50 frames per annotator hour, temporal mask consistency
Keypoint / pose Per frame $0.08 to $0.35 15 to 25 joints per subject, sports or clinical footage
3D cuboid / LiDAR Per frame $0.50 to $5.00 Multi-sensor scenes, multiple annotation passes

The 3D range looks absurdly wide until you look at the labor behind it. The SUN RGB-D dataset needed 2,051 annotation hours for 64,595 3D instances, roughly 114 seconds each by hand. Model-assisted methods on KITTI have brought that to under 4 seconds per cuboid. A 30x labor difference produces a 30x price difference, so ask which method you are buying.

If your project mixes video with image, text or sensor data, price each modality on its own model rather than blending them. A blended per-unit rate across image, text and video annotation services hides which stream is actually consuming the budget.

One correction worth making: the widely quoted $0.02 to $0.15 per-frame band is a bounding box band. Segmentation does not fit inside it and never has. If a vendor quotes segmentation at $0.12 a frame, ask how many frames per hour they expect an annotator to complete.

Frames per second. The single largest lever, and the one buyers control most easily. Most tracking models train fine on 5 to 10 fps. Going to 30 fps triples or quadruples your bill for data your model may not use.

Object density. Cost scales with object instances, not frames. A 12-object intersection costs four times a 3-object highway shot at identical frame counts.

Annotation type. Boxes tolerate a few pixels of drift. Segmentation does not. Per instance the gap is smaller than people assume: COCO recorded 79 seconds per instance for polygon drawing against the 34.5 second box reference, so roughly 2 to 3x. Per frame it is far worse, because full semantic segmentation assigns a class to every pixel rather than to one object. That is where the 10x shows up.

Motion complexity. Interpolation only pays when motion is predictable. Erratic movement, occlusion and camera cuts force keyframes closer together, and leverage collapses toward 1x.

QA depth. This is a multiplier, not a line item. A single review pass adds 15 to 25%. Consensus labeling multiplies the whole job: the SBD dataset merged five annotations per instance and recorded 315 seconds per instance in total. If your contract specifies three-way consensus, you have bought roughly 3x the annotation, not a QA surcharge.

Turnaround. Rush work typically carries a 20 to 50% premium, because the vendor is pulling trained annotators off other accounts and paying for the disruption on both sides.

Video Annotation Qa Review Hidden Cost

Scope: a 10-minute video, 30 fps, 5 objects per frame, bounding boxes with persistent track IDs.

That is 18,000 frames and 90,000 object instances.

Scenario A: every frame annotated by hand

At 8 seconds per instance including track ID assignment, 90,000 instances need about 200 annotator hours. Eight seconds is faster than the 34.5 second research reference because a specialist working consecutive frames of one scene carries the ontology in their head and does almost no visual search. Push the assumption to 15 seconds and the total nearly doubles, which is why this input belongs in the contract. Add 25% for senior QA and consistency review: 250 hours.

At a loaded rate of $8 to $11, delivery cost lands at $2,000 to $2,750. That is $0.11 to $0.15 per frame, or $200 to $275 per minute.

Scenario B: keyframes plus interpolation

Keyframe every fifth frame, so 3,600 keyframes and 18,000 fresh boxes at 8 seconds each: 40 hours. Roughly 45% of the 72,000 interpolated instances need adjustment at 2 seconds each: 18 hours. Full-sequence consistency review: 3 hours.

Annotation subtotal is 61 hours, a 69% reduction, which sits inside the 50 to 70% range HabileData measures on predictable motion. Senior QA adds about 13 hours.

At 74 hours and $8 to $11 loaded, delivery cost is $600 to $820. Quoted price for a one-off, including project management and tooling, typically lands at $750 to $1,100. That is $0.04 to $0.06 per frame, or $75 to $110 per minute.

The useful part is the gap. Those $150-per-minute ceilings you see published assume interpolation. Nobody delivers true 30 fps frame-by-frame at that price. If your quote is near the ceiling and the vendor is promising frame-by-frame, one of those two things is wrong.

Factor In-house Outsourcing
Cost per frame Higher. Fully loaded onshore annotator cost is $25 to $60 an hour before tooling and management. $6 to $12 an hour offshore for equivalent trained output on standard work.
Ramp time 2 to 3 weeks for a new annotator to reach target pace, plus hiring. Pilot batch in 2 to 3 business days, production within 48 to 72 hours of scope sign-off.
Scalability Fixed headcount. Volume spikes mean overtime or delay. Elastic. Parallel workflows absorb spikes without renegotiating headcount.
Quality control Full visibility, but QA competes with annotation for the same people. Independent QA layer by design. Verify the vendor measures IAA and MOTA, not just spot checks.
Domain knowledge Deep and permanent. Best for genuinely novel ontologies. Requires guideline investment upfront. Pays back from batch two onward.
Data security Contained. Depends on ISO certification, NDA coverage and access controls. Non-negotiable for medical and surveillance footage.
Outsourcing Video Annotation Delivery Team

In-house wins when your ontology is still moving and volumes are small, because the feedback loop between your ML team and your annotators is short. Outsourcing wins once the schema settles. HabileData clients report up to 60% lower costs than building the same capability internally.

Start with frame rate, not with the rate card. Ask your ML team what temporal resolution the model actually needs. Going from 30 fps to 15 halves the bill, and on most detection and tracking work it costs nothing measurable in performance. This is the negotiation people skip because it happens inside their own team rather than with the vendor.

Then look at interpolation. Frame interpolation pre-labeling cuts annotation time by 50 to 70% where motion is predictable, because annotators correct boxes instead of drawing them. Keep frame-by-frame for high-action or low-frame-rate footage, where interpolation would inject errors you then pay to remove.

Model-assisted pre-labeling stacks on top of that. A baseline model generating proposals roughly doubles annotator throughput with no accuracy penalty, as long as a human reviews every output.

The cheapest QA you will ever buy is an afternoon spent writing guidelines. Occlusion rules, minimum box size, ID persistence through screen exit and re-entry. Ambiguity is what generates rework, and rework is the line item nobody budgets.

Two more, briefly. Run a pilot on your own footage before committing to volume, because seconds-per-instance is the input that varies most and the one you cannot borrow from someone else’s project. And match precision to purpose: segmentation costs around 10x what boxes cost, so use it only where boundary accuracy changes what the model learns.

We annotate 50,000+ frames a day across 300+ annotators, with a 99.5% first-time approval rate and 95%+ annotation accuracy.

What matters more for your budget is how we get there. Every project starts with a pilot batch measured for inter-annotator agreement, and production begins only after benchmarks are met. The full technique and format coverage sits on our video annotation services page. QA runs in three stages: primary annotation with track ID assignment, senior review including temporal consistency checks across sliding five-frame windows, and automated MOTA calculation across the full sequence.

We work inside your tooling, whether that is CVAT, Labelbox, Scale AI, SuperAnnotate or a proprietary platform, and deliver in MOT Challenge, COCO Video, nuScenes, BDD100K or Waymo format.

On a sports analytics project covering 2 million frames across 500 games, maintaining ID consistency through occlusions and camera cuts brought the client’s ID switch rate from 8% down to under 2%.

That 300x spread in published per-minute pricing is not a market failure. It is 300 different combinations of frame rate, object density and QA depth, all sold under one label.

Which makes “what does video annotation cost” the wrong question. The useful one is what you are being quoted for. Boxes at 5 fps and boxes at 30 fps are different products with the same name, and the cheaper one will quietly starve a tracking model of the temporal continuity it was supposed to learn.

Settle four things before you put quotes side by side:

Quotes built on the same four assumptions can be compared. Quotes built on different ones cannot, however close the headline numbers look. Get those four into writing and the pricing conversation turns into an engineering one, which is where it belonged from the start.

Is per-frame or per-minute pricing better?

Per-frame gives you more control and is preferable when object density varies or you need frame-by-frame precision. Per-minute is simpler for long continuous footage with steady density. Either works if the assumptions are written into the contract.

Why is video annotation more expensive than image annotation?

Volume and temporal identity. A single 30 fps minute contains 1,800 frames, and every object must carry a consistent track ID across all of them. That consistency requirement, not the drawing, is the cost driver.

How do vendors calculate video annotation pricing?

They estimate object instances, multiply by seconds per instance, divide by interpolation leverage, then apply a loaded hourly rate and QA overhead. Most run a pilot first because seconds per instance varies heavily with footage quality.

Can AI-assisted annotation reduce costs?

Yes, typically 50 to 70% on predictable motion through interpolation, and roughly 50% through model-assisted pre-labeling. It does not remove human review. Fully automated labels degrade quietly on the edge cases your model most needs to learn.

What frame rate should I annotate at?

Most object detection and tracking models train well on 5 to 10 fps. Reserve 30 fps for fast-motion applications like ball tracking or collision analysis. This is the single largest cost lever you control.

How much does segmentation cost compared to bounding boxes?

Roughly 5 to 10 times more. Annotators produce 200 to 400 boxes an hour versus 20 to 50 segmented frames. Use segmentation only where boundary precision changes model performance.

What should be in a video annotation quote?

Annotated frame rate, object density ceiling with tiering above it, QA stages and acceptance criteria, output format, rework policy, and how scope changes are priced. Missing any of these means the number is provisional.

We price your video annotation from a measured pilot, not an industry average.

Get a video annotation quote   »

Leave a Reply

Your email address will not be published.

Author Snehal Joshi

About Author

, Head of Business Process Management at HabileData, leads a 500-member team of data professionals, having successfully delivered 500+ projects across B2B data aggregation, real estate, ecommerce, and manufacturing. His expertise spans data hygiene strategy, workflow automation, database management, and process optimization - making him a trusted voice on data quality and operational excellence for enterprises worldwide. 🔗Connect with Snehal on LinkedIn