Issue 019 - Surveillance - Computer vision throughput
How many facial-recognition alerts can one protest day create?
Reuters reported that activists are challenging the use of AI-powered surveillance at recent New Delhi youth protests. Police deployed an Ikshana mobile surveillance van with mast-mounted cameras and real-time facial-recognition capability while thousands gathered at Jantar Mantar.
The problem
Estimate, for one full protest day, the number of facial images or detections the system might process, the number of identity comparisons made against a police watchlist, and the number of false alerts that might require human review.
Optionally estimate the amount of video or cropped-face data generated.
Would computation, data storage, or human review probably be the main operational bottleneck?
Because Fermi problems target an order of magnitude, I normally use no more than two significant digits and write most calculations in scientific notation; the Fermi reference explains both conventions.
Before checking sources
Matt's first pass
I made a lot of assumptions on this one. When the reporting described thousands of people being present, I assumed about 10,000 unique protesters over the course of the day, but likely closer to an average of 4,000 protesters at any one time. I assumed a full day of protesting would be about 8 hours.
I assumed the camera setup would use one HD camera for every 30 degrees of ground coverage, or about 12 cameras. For storage, I assumed video would require about 1 GB per camera per hour.
video storage ~= 12 cameras x 8 hours x 1 GB/camera/hour
~= 96 GB
I assumed faces for only about 10% of the crowd would be distinguishable on camera at a given time, and that the crowd would move or shift enough to create unique captures about every minute, though many would repeat previously identified people.
average crowd ~= 4 x 10^3 people
visible face fraction ~= 10%
visible recognizable faces ~= 4 x 10^3 x 10%
~= 4 x 10^2 faces
protest duration ~= 8 hours
~= 480 minutes
face recognitions ~= 4 x 10^2 faces/minute x 480 minutes
~= 1.9 x 10^5 recognitions/day
I assumed the system would have a mechanism for preventing or at least limiting ID comparisons of already scanned faces, so only about 10% of all unique face recognitions would trigger new ID comparisons. I also assumed about a 5% error rate for ID comparisons.
ID checks ~= 10% x 1.9 x 10^5
~= 1.9 x 10^4 checks/day
false IDs ~= 5% x 1.9 x 10^4
~= 9.5 x 10^2 false IDs/day
I do not think 100 GB of storage for the video is a big deal. I also do not expect the computation would be particularly challenging. My understanding was that the software looks for unique identifiers like distances between facial features, then compares those against a database of name-labeled feature-distance measurements.
So the operational bottleneck would probably be human review, especially if humans need to review all comparison IDs. If a comparison ID could be human-checked once every 5 seconds, 190,000 checks would take about 25 person-hours. If that had to be done in real time, it would require at least four people.
Calibration Score
Matt's Calibration Score: 45 / 100
Higher is better: earn points for accurate pegs, sound models, correct math, and a result close to the sourced answer. The image shows percent full of it: 100 minus the Calibration Score.
Pegs: 10/30. Crowd, camera, and storage pegs were reasonable, but false-alert and detection-rate assumptions were weak.
Model: 15/30. The model needed a clearer distinction between unique people, frame-by-frame detections, comparisons, and human alerts.
Math: 10/10. The arithmetic was clean for the chosen assumptions.
Result: 10/30. The bottleneck conclusion was useful, but several intermediate counts were off by large factors.
Grounding facts
The crowd count can be only thousands while the computer sees millions of face boxes, because a visible person can be detected again and again across video frames. Then the database comparison count can jump again because each search is compared with many watchlist entries.
This is why the vocabulary matters: unique people, face detections, face tracks, watchlist searches, pairwise comparisons, and alerts are not interchangeable.
After checking sources
Check and recalibrate
Matt's storage estimate and camera count were in the right neighborhood. Indian Express reports Ikshana has eight fixed cameras, not far from Matt's 12-camera estimate. Full HD H.264 streams are often a few Mbps; Milestone gives a 4 Mbps example for a 1080p stream, which is about 1.8 GB/hour.
video storage range ~= 8 to 12 cameras
x 8 hours
x roughly 1 to 2 GB/camera/hour
video storage ~= 6 x 10^1 to 2 x 10^2 GB/day
Storage is not the bottleneck for a one-day event. The harder part is separating raw detections, deduplicated face tracks, identity searches, and database comparisons.
If 400 visible faces are processed once per second, the raw face-box count is already much larger than Matt's minute-level count:
visible faces ~= 4 x 10^2
processed frames ~= 1 frame/second x 8 hours
~= 2.9 x 10^4 frames
raw detections ~= 4 x 10^2 faces/frame x 2.9 x 10^4 frames
~= 1.2 x 10^7 face boxes/day
If the detector runs at 5 frames per second, that becomes about 6 x 10^7 boxes. But those are not necessarily all watchlist searches. A tracking system might turn many repeated boxes into fewer face tracks or face-search attempts. Matt's 1.9 x 10^5 minute-level estimate is better interpreted as face tracks, representative crops, or identity-search attempts.
The comparison count can jump again because one face-search attempt is compared against a watchlist. If there are 10,000 to 100,000 watchlist identities:
face-search attempts ~= 2 x 10^5/day
watchlist size ~= 1 x 10^4 to 1 x 10^5 people
similarity comparisons ~= (2 x 10^5) x (1 x 10^4 to 1 x 10^5)
~= 2 x 10^9 to 2 x 10^10 comparisons/day
That sounds enormous, but modern face-recognition systems convert images into numeric face templates or vectors, then use optimized vector search. The conceptual correction is not that the system compares hand-measured eye distances one at a time; it compares numeric embeddings and similarity scores.
For false alerts, the important rate is usually not "5% of all pairwise comparisons." In watchlist search, NIST reports false positive identification rate, or FPIR: the share of non-mated searches that produce one or more candidates above threshold. NIST's 1:N benchmark tables often report results at FPIR = 0.003, or 0.3%, as a standard operating point.
face-search attempts ~= 2 x 10^5/day
false-positive identification rate ~= 0.003
false alerts ~= (2 x 10^5) x (3 x 10^-3)
~= 6 x 10^2 alerts/day
A broad Fermi range is about 100 to 2,000 alerts, depending on threshold, image quality, watchlist size, masks, angle, lighting, and demographics. If the system is run loosely and returns candidate lists for many searches, the human burden can grow quickly. If it is run conservatively, it may miss more real watchlist matches.
Human review is probably the operational bottleneck, but with an important qualifier: humans normally review candidate alerts, not billions of vector comparisons. At 600 alerts and 10 seconds per alert:
review time ~= 600 alerts x 10 seconds/alert
~= 6 x 10^3 seconds
~= 1.7 person-hours
At 2,000 alerts, that is about 5 to 6 person-hours. That is manageable over a day, but it becomes real-time staffing, documentation, escalation, audit, and legal-risk work. If humans had to review every one of roughly 200,000 face-search attempts, the review burden would explode:
review every search ~= (2 x 10^5 searches) x (5 seconds/search)
~= 1 x 10^6 seconds
~= 2.8 x 10^2 person-hours
So the best answer is: video storage is modest, computation is substantial but tractable, and human review plus governance is the practical bottleneck. The bottleneck is not the hard drive; it is the meaning of each alert and what humans are required to do with it.
Post-check reflection
Matt's reflection
My video storage estimate and camera guesstimate are reasonable.
I had a conceptual shortcoming for facial detections. I was assuming a detection counts as each time a new facial box gets drawn, and not many times every second that a box is drawn, selecting some to do an identity comparison. I also had some confusion about how those checks actually operated. I had not considered the facial measurements being converted to vectors that allow for more efficient comparisons. Finally, the error rate was probably way too high, by an order of magnitude at least, maybe two.
The thing is, the error rate would not be something we could use to guide human review by itself. Humans would need to review alerts in order to identify the errors. That still means the same basic bottleneck I expected: human review. In any case, this does not significantly impact my impression of the news item. The operational limitation is going to be the same.
Recommended memory peg
Remember 8 hours is about 3 x 10^4 seconds, 1080p surveillance video is roughly 1 to 2 GB per camera-hour, and 1:N face search comparisons ~= face searches x watchlist size. For alert burden, use false alerts ~= searches x FPIR.
Reader results
Bars show how submitted estimates sort into the answer choices from the gut-check prompt.