[@ModernRogue] Flock Cameras Can Read Your Lips, Right Now.
· 4 min read
Link: https://youtu.be/Ks5Dzbt6EVI
Duration: 25 min
Transcript: Download plain text
Short Summary
Brian Brushwood and his co-host investigate whether AI can lip-read surveillance footage, using an Obsbot Tail 2 camera to simulate a Flock Safety ALPR/PTZ system. They demonstrate that, given enough compute, the AI can recover phrases from silent video, raising concerns that folk-grade surveillance now rivals state-grade capability.
Key Quotes
- "Yes, nationwide we got 120,000 Flock cameras. Here in Austin, we have 240-ish. There's a website, deflock.org, where people report where they are because concerned citizens want to know. I was thinking of 2001: A Space Odyssey when the HAL 9000 is reading the lips, and I thought to myself, can Flock cameras read our lips now? Is everything we do on the record now?" (00:01:16)
- "From what I understand, the more data you give it, for example, if it had a whole bunch of videos of Brian Brushwood talking, it would only get better and better, right? Again, this is one afternoon that I spent doing this so far." (00:10:42)
Detailed Summary
Episode Overview
Brian Brushwood and a co-host explore the intersection of AI lip-reading and mass surveillance, using consumer-grade cameras to simulate Flock Safety's Automatic License Plate Reader (ALPR) network.
Background on Flock and ALPR
- ALPR (Automatic License Plate Reader) cameras, deployed by Flock, track license plates, take photos, and record vehicle video.
- Approximately 120,000 Flock cameras are deployed nationwide, with about 240 in Austin; locations are crowdsourced at deflock.org.
- New Flock PTZ (pan-tilt-zoom) cameras feature 32x optical zoom capable of tracking people and capturing crisp video.
- On Lavaca Street, the hosts spotted two cameras: a fixed Flock camera pointing east and an adjacent city-operated domed PTZ camera.
Simulation Setup
- The hosts used an Obsbot Tail 2 (a 12x hybrid camera from the studio) alongside a GH5 to simulate a Flock-style setup.
- They filmed intentionally difficult sentences at close, medium, and long range, both indoors (hotel) and outdoors (swan boat).
- Flock cameras record at 24 fps, which the host hypothesized might be insufficient for accurate automated lip-reading.
Initial Lip-Reading Attempts
- An initial AI model couldn't read lips because it had no video module and merely sliced footage into a contact sheet.
- Two open-source projects were tested: "vallr" (wouldn't work) and "open alter ego" (produced spotty results).
- The first breakthrough was recovering "Irish folklore in Mobile, Alabama" from an indoor clip using a two-pass approach (confirm words, then infer context).
Outdoor Challenges and Breakthrough
- Outdoor filming was hindered by water glare, rocking swan neck blocking lips, wind, and speakers turning to profile.
- A wide, pixelated outdoor shot resembling surveillance footage initially yielded only five or six words.
- A prompt asking the model to be "more effortful" caused it to run the analysis about 64 times and statistically aggregate the samples into a usable read.
- Successfully recovered phrases included "very illegal drugs and or bombs" and "ever since I joined the mob, my life has really turned around," matching the actual audio.
Ensuring a Fair Test
- The presenter stripped the audio track to prevent the AI from cheating via transcription.
- ChatGPT's harness now runs baked-in transcription on incoming audio, which caused a HAL 9000 test clip to appear successful through audio rather than visual lip-reading.
The Underlying Mechanism
- The model performs frame-by-frame mouth-movement analysis mapped to likely phonemes.
- A "layer of intelligence" aggregates many guesses, making it more robust than brittle prior machine-learning systems.
- Frontier AI labs spend $7–8 million on 8–9 hour runs for math problems, sometimes running models "a million times" to statistically select the best guess — the same brute-force approach enabled their lip-reading success.
Surveillance Implications
- Citing Perry Carpenter's framework, the hosts argue folk-grade surveillance now matches state-grade because neither needs to be profitable.
- Flock's platform model is a risk: vendors accessing its 32x PTZ HD camera data could perform comparable lip-reading at scale.
- Flock's tech is not proprietary and has "broken containment" — everyone has cameras and access to similar tools.
- A hypothetical scenario imagined a festival surveillance camera transcribing every conversation and counting mentions of lemonade.
Closing Notes
- Brief NFL update: a Vikings-Packers game was in progress, with a Vikings pass to the 6-yard line and a facemask foul making it first and goal.
