PPE Detection: What a Camera Verifies, and What It Cannot
A hard hat is either on a head or it is not, and a camera answers that question all shift without getting bored. The question a PPE policy actually asks is narrower and harder: whether the right equipment, rated for this hazard, was worn correctly by this person in this area. PPE detection answers the first question well and cannot answer the second at all, and knowing exactly where the line falls is the difference between a useful observation layer and a false sign-off.
The problem: PPE detection is asked to prove more than it can see
A PPE program is a documented chain, not a moment. Under OSHA’s general PPE standard at 29 CFR 1910.132, the employer has to assess the workplace for hazards, select equipment appropriate to those hazards, communicate the selection to affected employees, and ensure the equipment properly fits each employee. The employer then has to verify that the assessment happened through a written certification identifying the workplace evaluated, the person certifying, and the date.
Read that list against a camera. Hazard assessment is a judgment made in advance. Selection is a procurement decision. Fit is a physical property of a specific garment on a specific body. Certification is a document. Exactly one item on the list — whether the selected equipment is actually being worn, here, now — is a visual fact in a scene.
The gap matters most where inspection attention is already concentrated, which for distribution and storage sites is the territory covered by OSHA’s warehouse national emphasis program.
PPE detection is the automated observation of whether required protective equipment is present on a person in a camera’s field of view, reported with a timestamp and usually a zone. It answers presence and absence continuously, on every camera, on every shift. It is not an assessment, a selection, a fit check, or a certification.
That single sentence is the whole boundary, and almost nothing sold in this category draws it.
Why the usual approach falls short
Buyers are rarely naive about this. They arrive having tried the alternatives, and each one leaves the same gap.
The walk-through. A supervisor with a clipboard checks a floor. This is the only method on the list that can judge fit and read a garment label, which makes it genuinely irreplaceable — and it covers the few minutes it covers. The sample is small and the observed behave differently from the unobserved.
The toolbox talk and the sign-in sheet. Excellent evidence that training was delivered. No evidence at all about what happened on the floor afterwards.
The vendor datasheet. Page one of a search for PPE detection is product pages, and they are structured identically: what the product detects, a list of items, an accuracy percentage, a demo button. Not one of them states a limit. The percentages are the tell — they arrive with no test set, no metric definition and no conditions, which makes them unfalsifiable rather than dishonest.
Compare that with how the research literature reports the same task. A 2020 study in Automation in Construction by Nath, Behzadan and Paal built an annotated dataset of roughly 1,500 images and around 4,700 worker instances, and its best approach reached 72.3% mean average precision on real-world images at eleven frames per second. Models have moved on since 2020 and that figure is not a verdict on any current product. What it demonstrates is the form of an honest claim: a named dataset, a named metric, a stated operating condition.
The percentages are unfalsifiable rather than dishonest. That is worse, not better.
What good looks like: split the policy into what a camera can observe and what it cannot
The useful design is not “buy detection” or “keep the clipboard.” It is knowing which question goes to which instrument.
| What the policy asks | Can a camera answer it? | What answers it instead |
|---|---|---|
| Is the required item being worn, here, now? | Yes — continuously, with a timestamp and a zone | Nothing else does this at full coverage |
| Is it the right type and class for this hazard? | No | Procurement specification and a hazard assessment |
| Does it fit this employee properly? | No | A physical fit check by a person |
| Is it still serviceable, or cracked and expired? | No | Scheduled inspection and replacement records |
| Was the hazard assessment performed and certified? | No | The written certification the standard requires |
What a camera answers well
Presence and absence of a large, high-contrast item in a predictable position is close to the ideal computer vision task. A hard hat sits at the top of a person, roughly where a detector expects a head. A high-visibility vest covers a large area of the torso in a color chosen precisely because it is conspicuous.
Two further things come free with continuous observation and are easy to undervalue. The first is the timestamp: an observation that carries a time and a zone is a record, and a record is what an incident review needs weeks later. The second is coverage of the hours nobody is walking the floor. Neither requires the camera to make a judgment about the garment.
Where the answer stops: type, class, fit and condition
This is where the vendor pages go quiet, and it is not a small gap.
Head protection has two independent axes, and a camera sees neither. ANSI/ISEA Z89.1-2014 (R2019), the American National Standard for industrial head protection, sorts helmets by Type — Type I tested for impact to the top of the head, Type II also tested for lateral impact — and separately by electrical Class: Class G tested at 2,200 volts, Class E tested at 20,000 volts, and Class C offering no electrical protection at all, often because it is vented. A worker in a vented Class C helmet near energized equipment is wearing head protection and is wearing the wrong head protection. A detector reports a hard hat in both cases.
The distinction is live rather than theoretical. In a Safety and Health Information Bulletin published on 22 November 2023, OSHA set out the limits of the traditional hard hat — protection concentrated on the top of the head, and no chin strap, so it can come off in a slip or trip — and encouraged employers to consider modern safety helmets with side-impact protection and retention. The bulletin is advisory and creates no new legal obligation. But if a site decides to act on it, the change from hard hat to safety helmet is one a camera is poorly placed to confirm, because from an overhead aisle view the two occupy the same space on the same head.
High-visibility apparel is worse, because the class is defined by measurement. ANSI/ISEA 107 is the governing standard for high-visibility safety apparel, and it sorts garments by Type and performance Class. Those classes turn on the measured area of background and retroreflective material and on laboratory photometric performance — properties established in a test lab and attested on a label, not visible in a frame. A camera can see that something bright and yellow-green is being worn. It cannot see square meters of retroreflective tape, and it cannot read certification. The 2020 edition is the one in effect; ISEA put a sixth revision out for public review with a comment period that closed on 11 May 2026, and no new edition has been published since.
Fit and condition are not inferable. Fit is an explicit duty in 1910.132(d)(1). A chin strap hanging loose, a suspension adjusted wrong, a shell with a hairline crack, a vest washed until the retroreflective tape has stopped performing — all of these are compliant-looking and non-compliant, and all of them read as “present” to a detector.
The resolution floor, and the part nobody mentions
Even the question a camera answers well has a distance beyond which it stops answering.
Detection works from pixels. IEC 62676-4:2025 anchors operational tasks to pixel density on the target plane, and the ladder Axis Communications publishes from that standard runs from 20 px/m for overview, through 40 for outline, 80 to discern, 125 to perceive and 250 to characterize, up to 500 and 1500 px/m for the most demanding tasks. A camera placed to give you an overview of an aisle is, by that ladder, placed to tell you a person is there — not to tell you what is on their head.
Two honest caveats matter more than the numbers. First, those values assume a human operator is interpreting the image, which Axis states plainly; applying them to an algorithm is an extrapolation, not a specification. Second, no published standard states a pixel density for PPE detection specifically, so anyone quoting you one has derived it themselves and should say so. Axis is equally direct that the model is a simplification and that light direction, optical quality and compression all move the result.
The practical consequence is unglamorous. The camera that gives you good site overview is usually the wrong camera for PPE, and the fix is a camera position and a lens, decided per zone, not a better model.
Designing the split
A program that works treats detection as the continuous layer underneath periodic human verification, and writes down which is which.
Put the items whose presence carries the risk on cameras, in the zones where the rule applies, and accept that the output is an observation rather than a finding. Keep type and class on the procurement side, where a specification and a receiving check settle them once for everyone, rather than trying to police them per person per day. Keep fit and condition on a scheduled human inspection, because they are physical properties. Then use the detection record to aim the human effort — a zone generating repeated observations at shift change is telling you where to spend the walk-through you were going to do anyway.
Five questions to put to a PPE detection vendor
Ask for written answers.
- What exactly does each detection assert? “Hard hat present” and “approved head protection in use” are different claims and only one of them is true.
- What is the metric, the test set and the conditions? A percentage without all three is not a measurement.
- At what pixel density on target does performance fall off, and how was that established for this detection rather than borrowed from a human-viewing standard?
- How does it behave at night, in rain, backlit, and against a busy background — measured on footage like ours, not on a curated reel.
- What does it do when it cannot tell? A system that reports uncertainty is more useful than one that guesses and reports a clean answer.
Where Nsightify fits
PPE compliance is one of the detections Nsightify runs from day one on the IP and CCTV cameras a site already operates, with per-zone rules, so a requirement that applies on the dock and not in the office is expressed as a zone rather than as a blanket alert. Each observation carries a time and a zone, which is what turns monitoring into a record an incident review can use later.
What the detection asserts is deliberately narrow: whether required equipment appears to be present on a person in view. It does not assert that a helmet is Type II, that a vest meets Class 2, that a garment fits, or that anything has been certified. Those stay with procurement, with the hazard assessment, and with a human inspection — and a vendor telling you otherwise is selling past the capability.
The limits are real. Detection depends on camera placement, sightlines, lighting, weather and distance, and a camera mounted for aisle overview will often resolve a person without resolving what is on their head — a mounting decision rather than a software one. Occlusion defeats it outright: a worker behind a rack or facing away may not be assessable at all. Nsightify publishes no accuracy figure for PPE detection or any other detection, because a number measured on someone else’s scene tells you little about yours. The way to find out is to test it the way the standards test analytics, on your own cameras.
Frequently asked questions
Can AI detect if someone is wearing a hard hat?
Yes, presence and absence of head protection is one of the things camera-based detection does most reliably, because a hard hat is a large, high-contrast object in a predictable position. What detection cannot report is which hard hat. The governing standard sorts head protection by impact type and by electrical class, and those distinctions are not visible to a camera at working distance.
How accurate is PPE detection?
There is no single figure, and any vendor quoting one without a test set, a metric and the conditions it was measured under has given you a claim rather than a measurement. Published peer-reviewed work reports its numbers properly: a 2020 study in Automation in Construction reported 72.3% mean average precision on real-world images using its own annotated dataset. Ask for the same disclosure.
Can a camera tell what class my vest is?
No. ANSI/ISEA 107 sorts high-visibility apparel by Type and by performance Class, which are determined by the measured area of background and retroreflective material and by laboratory photometric testing. A camera sees a garment that looks high-visibility. It cannot measure retroreflective area or confirm certification, so garment class stays a procurement and inspection control.
Does PPE detection work at night?
It depends on the light the camera actually has, not on the time of day. Detection works from the image, so anything degrading the image degrades the result: low light, glare, backlighting, rain and heavy compression. A yard lit to allow safe movement is not necessarily lit well enough to resolve what is on a person’s head at forty meters. Test on your own cameras, at night, before relying on it.
What camera resolution do I need for PPE detection?
No published standard states a pixel density for PPE detection specifically. IEC 62676-4:2025 anchors operational tasks to pixel density on the target plane, ranging from 20 px/m for overview up to 1500 px/m for scrutinizing detail, but those figures assume a human operator is interpreting the image. Treat them as a starting point and confirm on your own cameras at the distances you actually need.
What to write down this week
Take your PPE policy and put a mark against every requirement in it: observable continuously by a camera, or not. The list usually splits about one to four, and the short column is the one worth automating.
Then check the cameras covering those zones against the distance question rather than the megapixel count, because a person-sized target and a hard-hat-sized target are not the same problem. The zones where the answer is “too far” are the ones to fix first, and they are the same zones worth raising when you ask the questions a video analytics vendor cannot bluff. If the requirement is about keeping people out of a space rather than what they are wearing, restricted-area detection around machinery is a different and often easier problem.
If you want to see what your own cameras can and cannot resolve before committing to anything, book a demo with the Nsightify team.
More on this from Nsightify: PPE detection and hazard-zone monitoring.
See Nsightify in Action
We're onboarding a limited number of pilot partners. If you're an operations or security leader in construction, warehousing, or manufacturing — let's talk.