
Latest insights in
Human Computing
Explore our blogs, research, and updates on Human Computing.
spotlight
Sparrow-2: Beyond Turn-Taking to Whole-Scene Conversational Understanding
Sparrow-2 is a real-time conversational understanding model that jointly models turn-taking, interruptions, backchannels, and the full acoustic scene as a single, unified system. It is an audio-native, streaming-first engine, rebuilt from the ground up, that goes beyond endpoint detection to transform the entire audio stream into decisions about when to listen, wait, or speak at a native 10 ms frame rate.