Skip to main content

[@ModernRogue] Flock Cameras Can Read Your Lips, Right Now.

· 4 min read

@ModernRogue - "Flock Cameras Can Read Your Lips, Right Now."

Link: https://youtu.be/Ks5Dzbt6EVI

Duration: 25 min

Transcript: Download plain text

Short Summary

Brian Brushwood and his co-host investigate whether AI can lip-read surveillance footage, using an Obsbot Tail 2 camera to simulate a Flock Safety ALPR/PTZ system. They demonstrate that, given enough compute, the AI can recover phrases from silent video, raising concerns that folk-grade surveillance now rivals state-grade capability.

Key Quotes

  1. "Yes, nationwide we got 120,000 Flock cameras. Here in Austin, we have 240-ish. There's a website, deflock.org, where people report where they are because concerned citizens want to know. I was thinking of 2001: A Space Odyssey when the HAL 9000 is reading the lips, and I thought to myself, can Flock cameras read our lips now? Is everything we do on the record now?" (00:01:16)
  2. "From what I understand, the more data you give it, for example, if it had a whole bunch of videos of Brian Brushwood talking, it would only get better and better, right? Again, this is one afternoon that I spent doing this so far." (00:10:42)

Detailed Summary

Episode Overview

Brian Brushwood and a co-host explore the intersection of AI lip-reading and mass surveillance, using consumer-grade cameras to simulate Flock Safety's Automatic License Plate Reader (ALPR) network.

Background on Flock and ALPR

  • ALPR (Automatic License Plate Reader) cameras, deployed by Flock, track license plates, take photos, and record vehicle video.
  • Approximately 120,000 Flock cameras are deployed nationwide, with about 240 in Austin; locations are crowdsourced at deflock.org.
  • New Flock PTZ (pan-tilt-zoom) cameras feature 32x optical zoom capable of tracking people and capturing crisp video.
  • On Lavaca Street, the hosts spotted two cameras: a fixed Flock camera pointing east and an adjacent city-operated domed PTZ camera.

Simulation Setup

  • The hosts used an Obsbot Tail 2 (a 12x hybrid camera from the studio) alongside a GH5 to simulate a Flock-style setup.
  • They filmed intentionally difficult sentences at close, medium, and long range, both indoors (hotel) and outdoors (swan boat).
  • Flock cameras record at 24 fps, which the host hypothesized might be insufficient for accurate automated lip-reading.

Initial Lip-Reading Attempts

  • An initial AI model couldn't read lips because it had no video module and merely sliced footage into a contact sheet.
  • Two open-source projects were tested: "vallr" (wouldn't work) and "open alter ego" (produced spotty results).
  • The first breakthrough was recovering "Irish folklore in Mobile, Alabama" from an indoor clip using a two-pass approach (confirm words, then infer context).

Outdoor Challenges and Breakthrough

  • Outdoor filming was hindered by water glare, rocking swan neck blocking lips, wind, and speakers turning to profile.
  • A wide, pixelated outdoor shot resembling surveillance footage initially yielded only five or six words.
  • A prompt asking the model to be "more effortful" caused it to run the analysis about 64 times and statistically aggregate the samples into a usable read.
  • Successfully recovered phrases included "very illegal drugs and or bombs" and "ever since I joined the mob, my life has really turned around," matching the actual audio.

Ensuring a Fair Test

  • The presenter stripped the audio track to prevent the AI from cheating via transcription.
  • ChatGPT's harness now runs baked-in transcription on incoming audio, which caused a HAL 9000 test clip to appear successful through audio rather than visual lip-reading.

The Underlying Mechanism

  • The model performs frame-by-frame mouth-movement analysis mapped to likely phonemes.
  • A "layer of intelligence" aggregates many guesses, making it more robust than brittle prior machine-learning systems.
  • Frontier AI labs spend $7–8 million on 8–9 hour runs for math problems, sometimes running models "a million times" to statistically select the best guess — the same brute-force approach enabled their lip-reading success.

Surveillance Implications

  • Citing Perry Carpenter's framework, the hosts argue folk-grade surveillance now matches state-grade because neither needs to be profitable.
  • Flock's platform model is a risk: vendors accessing its 32x PTZ HD camera data could perform comparable lip-reading at scale.
  • Flock's tech is not proprietary and has "broken containment" — everyone has cameras and access to similar tools.
  • A hypothetical scenario imagined a festival surveillance camera transcribing every conversation and counting mentions of lemonade.

Closing Notes

  • Brief NFL update: a Vikings-Packers game was in progress, with a Vikings pass to the 6-yard line and a facemask foul making it first and goal.