How Accurate Is AI Video Analytics? Test It, Don't Trust It

An industrial yard fence line at dusk in light rain, with a pole-mounted security camera angled along the fence and wet gravel reflecting sodium light.

A video analytics datasheet quotes an accuracy figure in the high nineties, and the figure cannot be wrong, because it never says what it was measured on. The buyer finds out what it means on their own fence line, on a wet night, when the alert queue fills with shadows or stays silent through the one event that mattered. The question “how accurate is AI video analytics?” has a real answer for a given site, but the percentage on the datasheet is not it.

The problem: how accurate is AI video analytics, measured against what?

An accuracy claim is only meaningful if four things come with it: the scenario being detected, the scene it was tested in, the conditions — light, weather, distance, camera movement — and which of two possible errors the number describes. Marketing figures for this category routinely arrive with none of the four. That is not necessarily dishonest. It is simply unfalsifiable, and a buyer cannot plan an operation around a number they cannot check.

The deeper problem is arithmetic. In a security or safety scene, genuine events are rare, so a system can score almost perfectly while being useless. Take one camera watching a yard fence for 30 days, which is 43,200 minutes, with two genuine intrusions in that month. Score each minute as right or wrong.

System What it does Minute-by-minute “accuracy”
A Never alarms. Misses both intrusions. 43,198 of 43,200 minutes correct: 99.995%
B Catches both intrusions, plus 90 false alarms over the month 43,110 of 43,200 minutes correct: 99.79%

The system that would have caught both break-ins scores lower than the one that watches nothing. That is not a trick of the example. Any measure dominated by the long stretches where nothing happens will reward a detector for staying quiet.

System B has its own problem, and the same table exposes it. Of its 92 alarms, two were real, so an operator who responds to every alert is right about one time in 46. A buyer who asks only “how accurate is it?” learns nothing about either system’s failure.

A detector that watches nothing can score 99.995%.

Why the usual approach falls short

Buyers are not naive about this. They have learned to discount datasheets, and they reach for three substitutes, each of which helps less than it appears to.

The vendor demo. It runs on footage the vendor chose, in conditions the vendor knows the system handles. It proves the detection exists. It says nothing about how often it fires on your site’s birds, headlights and rain.

The pilot on the best camera. A trial on the newest, best-placed camera in daylight is designed to succeed, and it usually does. The camera at the far corner of the yard, the one pointed into low sun at 5 p.m., is where the alerts will actually go wrong, and it was never in the test.

The request for a better number. Asking for “false positive rate” instead of “accuracy” is progress, but a rate without a denominator repeats the original problem. False alarms per what — per camera, per hour, per event, per thousand frames? Each gives a different number for the same system on the same night.

What is missing in all three is an external reference point: a way of stating a test that does not depend on trusting whoever ran it. That reference point exists, and it has existed for some time.

What good looks like: test video analytics the way the standards do

Video analytics accuracy is not one number but two, each stated for a named scenario and scene. Recall is the share of genuine events that produce an alarm. Precision is the share of alarms that correspond to a genuine event. A figure that does not say which it measures, or under what conditions, is a claim rather than a measurement.

Two errors, two costs, and a role that decides between them

A missed event and a false alarm are different failures, and which one hurts more depends on what happens to the alert.

The UK Home Office made this explicit in its i-LIDS programme. The i-LIDS User Guide, published by its Centre for Applied Science and Technology in February 2011, scores systems on a weighted combination of recall and precision. It evaluates them for one of two roles. In the operational alert role, a human controller must deal with each detection in real time. In the event recording role, detections only trigger recording for later review, so false alarms matter less. The guide sets a separate weighting for each role and each scenario, and for abandoned baggage the recording role’s weighting towards recall was a hundred times the live-alert role’s.

That is the most useful idea a buyer can take from the programme. Before asking how good a system is, decide what a false alarm costs your operators and what a missed event costs your site.

If alarms go to The failure that costs most What to weight in your test
A person who must respond now False alarms, which spend attention and train people to ignore the queue Precision, stated as false alarms per camera per day
Recording or later review Missed events, which leave nothing to review Recall, stated as genuine events caught out of genuine events staged

What a government test library taught the category

i-LIDS paired that metric with discipline about the test itself. The i-LIDS brochure describes six scenarios, including abandoned baggage, doorway surveillance, parked vehicles and sterile zone monitoring, built from footage it says represents real operating conditions. Several design choices are worth copying directly.

  1. The certifying footage was private. Each scenario had a public training set and a public test set, and certification ran on a third, private evaluation set held by the Home Office.
  2. Every alarm event was defined in advance. An abandoned object, for example, counted only once its owner had been gone from the area for a set time.
  3. Timing was part of the answer. A system had ten seconds from the start of an event to raise an alarm; anything later counted as a false positive.
  4. Noise was catalogued, not ignored. For sterile-zone footage along fences, the records logged potential causes of false alarms — foxes, bats, insects on the lens, shadows through the fence, a camera switching from colour to monochrome — along with time of day, rain, snow and fog.
  5. The pass mark was never published. The guide states that the scores needed for a recommendation were not made public, so a manufacturer could not tune to a known threshold.

The programme dates from over a decade ago, and its hardware and formats show it. The method has not aged.

What IEC 62676-6:2026 adds

On 18 February 2026, the IEC published IEC 62676-6:2026, a 224-page international standard for performance testing and grading of real-time intelligent video content analysis in surveillance systems. Its published scope says it specifies functions, performance, test methods, performance evaluation and grading rules, and that it applies to live and forensic analysis.

The scope organises testing into three dimensions:

  • Core capability — classifying objects and detecting object activity such as stopping, starting and direction of movement.
  • Complex capability — detecting scenarios built from that activity, with examples including loitering, perimeter intrusion detection, person down, tailgating and abandoned object detection.
  • Degree of difficulty — testing under defined levels of operating stress, with examples such as sterile or non-sterile scenes, indoor or outdoor, target obscuration, extreme weather and camera shake.

It also names end users, installers, integrators and certification providers among the people it is meant to give measurement methods to.

A note on what this article can and cannot say. The full standard is sold, and everything above comes from its publicly readable scope. Nothing here describes its clauses, test procedures or grade boundaries, and no claim is made about any product’s grade. What matters for a buyer is that the vocabulary now exists in a published international standard: scenario, capability, and degree of difficulty. A vendor who claims performance can be asked to state it in those terms.

Five questions that turn a percentage into a test

Put these in the RFP, and ask for written answers.

  1. Which scenario, and what counted as a detection? Name the event, when it starts, and how long the system had to alarm.
  2. Recall and precision, as counts. How many genuine events were staged or present, how many were caught, and how many alarms fired in total, per camera per day.
  3. Under which conditions? Day and night, weather, distance to target, obscuration and camera movement — the degree of difficulty, stated rather than implied.
  4. Was the test footage separate from the training footage? A system scored on the footage it learned from has been marked on its own homework.
  5. Which role was it tuned for? A live-alert configuration and a record-for-review configuration are different settings, and one set of figures cannot describe both.

The video analytics RFP questions vendors cannot bluff cover the architecture side of the same procurement: streams, edge processing and data retention.

Writing an acceptance test for your own site

No external standard can say what a system will do on your cameras. It can only tell you how to find out. A site acceptance test borrows the structure above and applies it to your own footage.

Choose the two or three scenarios that justify the purchase. Include the worst cameras, not just the best: the long view, the backlit gate, the dock that floods with headlights at shift change. Run across nights and weather, not a single afternoon. Define each alarm event and the response window before the test starts, and count true alarms, false alarms and misses per camera per day.

Then set the bar by role. If a guard must respond to every perimeter alert, agree the number of false alarms per camera per night the team can genuinely work, because the cost of loitering and dwell-time false alarms is paid in attention long before it shows up anywhere else. If alarms only flag footage for review, weight the misses.

Where Nsightify fits

Nsightify does not publish an accuracy percentage for any detection, and after the arithmetic above it would be an odd thing to offer: a figure measured on someone else’s scene tells you little about yours.

What Nsightify does is run its detections on the IP and CCTV cameras a site already operates, with real-time alerts, which means the test described above can be run on the real estate rather than a demonstration rig. Some detections are ready to run from day one, such as intrusion into a restricted zone, loitering, PPE compliance and forklift–pedestrian proximity. Others, such as perimeter pacing, fence-climbing and abandoned object detection, are included and switched on during commissioning once tuned against the site’s own footage, because a busy yard and a quiet one produce different results.

The limits apply to Nsightify as to any system. Detection depends on camera placement, sightlines, lighting and weather, and the only honest answer to “how accurate is it here?” comes from counting alarms, misses and false alarms on your own cameras, in your own conditions.

Frequently asked questions

How accurate is AI video analytics?

It depends on the scenario, the scene, the lighting and weather, and the alert threshold, which is why a single percentage cannot answer the question. Accuracy is two measurements: recall, the share of genuine events that trigger an alarm, and precision, the share of alarms that are genuine. Ask for both, as counts, for a named scenario under stated conditions, and confirm them on your own cameras.

What is a good false positive rate for video analytics?

There is no universal figure. The acceptable rate depends on what happens to an alarm. If a person must respond to every alert in real time, false alarms consume attention and should weigh heavily. If alarms only mark footage for later review, missed events cost more. Set the target as false alarms per camera per day, based on how many alerts your operators can genuinely work.

Is there a standard for testing video analytics?

Yes. IEC 62676-6:2026, published on 18 February 2026, specifies test methods and grading rules for real-time intelligent video analysis in surveillance systems. Its published scope covers object classification and activity, complex scenarios such as loitering and perimeter intrusion, and testing under defined operating stress. Before it, the UK Home Office’s i-LIDS programme set a public precedent for scenario-based testing against defined alarm events.

Why does my camera keep sending false alarms?

Usually because the detection is working in conditions it was never tested under. The UK government’s i-LIDS test footage for fence-line scenes logged false-alarm causes such as foxes, bats, insects on the lens, shadows through the fence and cameras switching between colour and monochrome, alongside rain, snow and fog. A system tuned on a clear daytime scene meets all of them at night. Test and tune on your own cameras.


What to change in your evaluation this week

Strike the word “accuracy” from the vendor questionnaire and replace it with the five questions above. Pick the three cameras on your estate most likely to embarrass a detector, and write down, before any trial starts, what counts as an alarm event and how many false alarms per camera per day your team can work. The same camera list does double duty as the stream inventory IT needs before approving video analytics, so it is not wasted if the trial moves forward.

If you want to run that test on your own cameras rather than read about it, apply to the Nsightify pilot program.

More on this from Nsightify: AI video analytics on existing IP and CCTV cameras.

See Nsightify in Action

We're onboarding a limited number of pilot partners. If you're an operations or security leader in construction, warehousing, or manufacturing — let's talk.

Nsightify team collaborating in the operations center