Introducing Griffin
Griffin is the world’s first Human Interaction Model (HIM), a new class of model designed to understand and generate face-to-face real-time human interaction. It listens while they talk, and they pay attention to expressions and pauses, not just words. Griffin combines our earlier research into real-time perception, conversation and expressive human-like video into one unified video to video system.
.avif)
Today we’re introducing Griffin, our first Human Interaction Model (HIM). HIMs are a new class of models designed to understand and generate face-to-face real-time human interaction. They listen while they talk, and they pay attention to expressions and pauses, not just words.
Griffin builds on our earlier research into real-time perception, conversation and expressive human-like video. It combines those model capabilities into one, unified video-to-video system.
Griffin is the first model to pass the real-time, video Turing test. In a live study, 48% of participants who talked with Griffin thought they had talked with a real human. Previous systems, including Tavus’s own industry leading conversational video interface (powered by Phoenix 4.5, Sparrow-2, Raven-1), had a max 2% pass rate. This represents a huge breakthrough in human-machine understanding, and is only possible due to Griffin’s video-to-video duplex approach with audiovisual generation, movement, conversational modeling – all in one system, operating in real time.
Griffin-Lite, a research preview, is available today to a select group of early testers, with a wider release of a more powerful model to follow. This post covers what Griffin can do, how we built it, how we measured it, our approach to safety and what comes next.
Human computing
A message from Hassaan:
Humans are evolutionarily designed to communicate face to face. We speak as much through our words as we do our expressions, tone, gestures and timing. So much of human communication is non-verbal. It’s an art, a dance.
Machines don’t understand this art. They require us to meet them where they are, to learn their language- the right command, the right button click, the right prompt.
At Tavus, we believe in a future where machines meet us where we are. Where machines learn to communicate in the way that we most effortlessly do. They learn to see, hear, respond and even look like we do. In effect, we want computing to become invisible. We want it to feel second nature.
In a great conversation, you don’t spend your time evaluating the mechanics. You think about what you want to say, what you’re hearing and seeing, and ultimately the conversation just flows. In great conversations, you aren’t thinking about the conversation at all.
With machines, those mechanics are distracting and demand your attention.
Did it hear me? Did it understand me? You share something personal or difficult, and the face continues to smile, or the faceless, cold machine sits there, then responds with a monotonous ‘I’m sorry to hear that’. You pause to collect your thoughts, and it jumps on and responds to incomplete context. It needs time to answer, but nothing in its expression or voice or gesture tells you it’s working on a response.
Each of these moments is you managing the machine. You adjust your timing, simplify your language, make sure to think complete thoughts before saying anything. You remind yourself: don’t leave dead air! Remember to be very clear about what you want! You lower your expectations. The system, in all its intelligence, might be capable of infinite incredible work. It can discover new mathematics dammit! Yet, getting the help for the simple stuff takes more effort than you have to give.
Human communication is an art, a dance. The right move, at the right time. A deep understanding of what you mean, when the words themselves could mean many things. A nod of acknowledgment to signal understanding. An expression that tells you everything without saying anything.
All of these are things we do as humans without thinking about it. These signals let your attention stay on the conversation, not how to converse. As a machine becomes more coherent at expressing and understanding these signals, communicating with it, too, becomes something you don’t have to think about.
Introducing Griffin: The First Human Interaction Model
Griffin-Lite is a preview of our first Human Interaction Model: a full-duplex video-to-video model that responds to human behavior in real time. Perception, deciding when and how to respond, and expressive speech and video generation all happen at the same time.
- 1It says “mm-hm” while you’re still talking.
- 2It answers without leaving dead air.
- 3It stops the moment you cut in.
- 4It reacts to something it sees.
The timing in this diagram is illustrative and was not measured from a real session.
Capabilities
To see what it looks like in practice, we brought people who had never used Griffin together with Tavus researchers for open-ended conversations and short demonstrations.
Behavior and emotion modelling
Griffin has learned to express behaviors, emotions, and gestures according to conversational context. It laughs, changes its tone, and shifts in response to what the other person says and does.
Rudy tells Vanessa he’s been promoted. She reacts before he’s finished, and the two of them talk over each other for a moment without losing the thread.
Full-scene generation
Griffin generates every pixel in every frame in real time from one reference image. It controls the whole scene, not just the face, arms, and fingers, but also things like the movement of the chair the person is sitting in, the shadows they cast, and the background behind them.
A round of Simon says with Mars. Griffin copies the gesture only when the person says “Simon says,” and calls their bluff when they don’t. Every movement is generated on the spot, body and background included.
Full-duplex conversation
Griffin continuously evaluates the state of the conversation and independently of conversational turns. It can interrupt, adjust, back-channel, or be interrupted without losing its place.
Ari and Griffin make up a story together, cutting each other off as they go, and Griffin runs with every twist.
Perception
Griffin elevates dialogue beyond just verbal understanding, embedding and reacting to visual context and awareness when appropriate.
Griffin coaches Ari through a Rubik’s Cube. It watches the cube as he turns it, keeps acknowledging him while he works, and when he pauses mid-solve it waits until he’s ready.
Temporal understanding
Griffin understands time as part of the conversation: how long a silence has lasted, what a sustained silence means, and when to speak again on its own.
Sagar solders a motherboard while Griffin guides him. It keeps track of time and where he is in the job, and it speaks up when the next step is due rather than whenever he goes quiet.
Technical approach: Building a full-duplex, video-to-video interaction model
Perception, conversational decision-making, and generation are in continuous interplay throughout human-to-human interaction, and we designed Griffin around the same principle. Central to our approach is a conversational model that controls speech and nonverbal behavior directly, with streaming speech and video generators that follow its decisions as they change. Griffin, is a two part system:
- A Continuous Conversational Modeling engine that perceives the incoming audio and video, decides when and how to respond, and produces signals that determine what should be said and how.
- An Audio Visual Generation engine comprising a Streaming Speech Generation and a Streaming Video Generation architecture that convert those signals into speech and video, respectively.
Griffin is a two-part system. A Continuous Conversational Modeling engine perceives the incoming audio and video, decides when and how to respond, and produces signals that determine what should be said and how. An Audio-Visual Generation engine, made up of a Streaming Speech Generation and a Streaming Video Generation architecture, converts those signals into speech and video. Perception, decision-making, and generation run concurrently throughout the conversation, so the model can respond to changes while it is speaking as well as while it is listening.
Perception, decision-making, and generation run concurrently throughout the conversation, so the model can respond to changes in the conversation while it’s speaking as well as while it’s listening.
Continuous Conversational Modeling: (Conversational) intelligence, Perception and Control
Griffin makes conversational decisions at regular sub-second intervals rather than once per turn. At each interval it assesses the state of the conversation, including what has been said and the user’s verbal and nonverbal behavior, and decides what to say next and how to say it. This differs from cascade systems that chain speech recognition, a language model, and speech synthesis, one after the other, and which wait for the user to finish speaking before they begin to respond.
The output controls not only what is said but also how the speech is delivered and how the model behaves nonverbally, including emotional tone, stance, facial expression, and gesture. Because these decisions are made continuously and are conditioned on context, the model can take and yield the turn on the basis of what is being said rather than when the audio stops, so a pause for thought is not treated as the end of the turn. The same mechanism lets the model nod in agreement while the user is speaking, backchannel and confirm at natural points, and adjust the emotional content of its response to the user’s speech and nonverbal behavior.
Additionally, Griffin perceives video as well as audio, where most interactive systems perceive audio alone. Visual input lets the model read the user’s nonverbal communicative signals, such as gaze and facial expression, and it also gives the model access to the user’s environment and to visual material the user chooses to share, such as their screen.
This is illustrated in the figure below.
The emitted outputs are converted to streaming audio outputs and together with streaming control signals are driving video generation as explained below.
Fast and Expressive Streaming Speech Generation
Griffin-Lite’s speech generation produces high-quality speech at conversational latency and can clone a speaker’s voice from about 10 seconds of audio. The model is a fast autoregressive diffusion transformer. It takes an encoding of the reference clip as a prefix and generates the new speech progressively, one latent chunk at a time, as increments of outputs and controls arrive from the conversational model, instead of waiting for the full utterance.
A compact continuous codec
Griffin’s speech generation speed and fidelity start with how it represents sound. It relies on Tavec, a convolutional autoencoder that maps 48 kHz audio into a compact continuous latent: 40 values per frame, 100 frames per second, and no codebooks. Its fully causal (streaming-ready) decoder carries state across chunks, turning streamed latents into one seamless waveform, with audio packets as small as 10 ms.
This continuous representation supports straightforward regression-based flow matching and fully differentiable audio pipelines, without quantizer gradient approximations. Its compactness also makes long sequences manageable: a minute of speech corresponds to 6,000 latent frames, against 2.88 million samples of raw audio.
Griffin: Fast and Expressive Streaming Video generation
The other side of Griffin’s audiovisual generation system is fast diffusion-based generation architecture. Previous Tavus models relied on extensive 3D priors, which limited what they could express: complex, human-like behavior such as hand and large body gestures was out of reach, and so were dynamic backgrounds.
We built it to meet three requirements:
- Respond quickly to each incoming chunk of audio and expressive controls
- Handle controls that change fast
- Hold visual quality, lip sync, and identity consistency over long generations
Autoregressive latent video models typically compromise on at least one of these. The architecture we arrived at generates 720p video in 320 ms chunks in real time, one latent at a time, and accepts streaming controls that let the conversation model direct nonverbal behavior such as gesturing, looking away, or changing emotion.
To build it, we distilled a large, bidirectional, many-step diffusion model into a few-step autoregressive generator. The generator takes a reference image together with streaming audio and controls, and produces one latent at a time. The distillation ran in three stages. First, we distilled the teacher into a few-step student with Distribution Matching Distillation. This student is fast but not autoregressive; it generates a long video as a single long chunk. Second, we converted it into an autoregressive model with teacher forcing, so as to arrive in an architecture that can generate in a few steps one latent at a time. Third, we trained with Self-Forcing so that we can run long autoregressive rollouts without drift, and we added a mechanism acting on the history frames, the previously generated frames the model continues from, to improve stability.
Griffin-Lite: Evaluations and benchmarks
Ahead of this preview, we ran Griffin-Lite through a set of evaluations at different levels, from a single component to a live conversation:
- The video generator on its own, against published streaming diffusion models, on latency, visual quality, and lip sync
- The full system on VideoFDB, NVIDIA’s leading industry benchmark for full-duplex audio-visual conversation, which scores how a model reads a person’s behavior and how it responds
- The full system in live one-minute video calls with people who did not know they were talking to a model
Video Generation
We compared the video generator of Griffin-Lite against four published streaming diffusion models in an audio-to-video setting, where each model is given speech and a reference image and has to produce the talking face.
Griffin-Lite produces one latent at a time and does not wait for future audio, so there is little delay between a piece of audio arriving and its effect appearing on screen. On H100s this averaged 0.43 seconds, half that of the next fastest method. In a conversation, this significantly cuts down the delay before the face reacts when the person interrupts or when a nod is due.
This is depicted in the diagram below, where the dark bar represents the time required by different models to generate the smallest chunk of video they can produce, and pink the time required to playback a video chunk. In the figure, we mark the best average and the worst case latency of each model.
Griffin-Lite also scored highest on visual quality. It led every baseline on DOVER and FID, the two standard measures of video quality, and on THEval, a recently published framework built for talking heads. On lip sync it placed second on LSE-C, at 7.27. LSE-C rewards pronounced mouth movement, including movement past the point where it looks natural, which THEval accounts for.
Full-Duplex Conversation
NVIDIA’s VideoFDB is a research and industry benchmark for evaluating the naturalness and accuracy of full-duplex audio-visual conversations. It examines how a model responds to and produces the signals of human interaction, including dialogue, gaze, facial expression, and body movement, measuring both conversation quality and response time. The benchmark scores a model on two tracks, generation and perception, each with its own rubric and its own leaderboard. NVIDIA conducts the evaluation independently using the published metrics and its own judge.
Griffin-Lite scored highest on both tracks against the benchmark’s published baselines, which include commercial and open-source systems— and it is the only case where the same model is evaluated on both tracks.
The generation track tests whether the model produces the right behavior. Given a person’s audio, it scores the model’s speech and video together for fluency, affect matching, and whether the nonverbal behavior fits the moment, laughing with the person rather than after them. Griffin-Lite scored 3.83, over 1 full point ahead of the next-highest system and only 0.09 below the human reference at 3.92, which is more than 12x closer to human performance than any other system evaluated.
VideoFDB’s perception track tests the other side of the conversation, which is whether the model understands the moment. Given a person’s audio and video, it scores the model’s spoken response for fluency, conversational flow, and visual grounding, meaning whether the model uses what it sees, a pause with a glance away, for instance, rather than reacting to the words alone. Griffin-Lite scored 3.73, 0.29 points ahead of the strongest second best and 0.47 below the human reference at 4.20, marking Griffin-Lite as the most perceptive model on the benchmark among the 15 models evaluated, ahead of every frontier realtime model on the leaderboard.
VideoFDB also reports takeover-rate alignment, which measures how closely the model’s decisions about when to speak match the timing in the reference conversations. Griffin-Lite scored 62.8% on generation and 73.8% on perception, the highest of any system on both tracks.
Face-to-face study
Quantitative benchmarks are important, but some aspects of naturalness are difficult to measure. For every new model, we run a study with participants recruited through an independent research platform to understand how human-like it is in a real, live conversation. Participants have a short video call with the model without being told what it is, and afterward we ask whether they think their partner was a real person. Until now, almost no one has said yes.
On Phoenix-4.5, 2.4% of participants (n=41) believed their partner was a real person. On Griffin-Lite, nearly half did: 48% (n=54). This represents a milestone of, to the best of our knowledge, the first model to have ever passed the video Turing test.
Methodology
Participants were told they would be matched with another participant for a one-minute video call to discuss what they were looking forward to this year. Their partner was in fact a PAL powered by Griffin-Lite, generating her face, voice, and responses in real time. After the call, participants were asked to write down their partner’s answer to the question and to rate their partner on several axes around naturalness and trustworthiness. They were also asked to rate the conversation itself: how well it flowed, and whether they felt their partner was really listening. Only at the end of the survey were participants asked whether it had crossed their mind that their partner might not be a real person, and if so, when. Every participant was then told that their partner had been an AI model.
Results
Of the 54 participants who spoke with Griffin-Lite, 26 believed their partner was a real person. Participants on both sides were confident in their answer: those who said real averaged 79% confidence, and those who said AI averaged 81%.
Over half of participants said the possibility had not crossed their mind during the call, and nearly all of that group went on to say their partner was real. Those who did suspect tended to suspect within the first 20 seconds.
Griffin-Lite was rated favorably on every axis. On a 7-point scale, participants on average rated it a 5.4 for seeming natural, 5.6 for seeming trustworthy, and 5.8 for whether they would enjoy talking with it again, a score that held at 5.4 among participants who said it was AI. The conversation scored 5.5 for whether participants felt their partner was really listening and 4.9 for flowing naturally, the lowest of the five.
Safety, limitations, and responsibility
Given Griffin-Lite is the first model to pass the Turing test, we understand the responsibility to be thoughtful on the risks and release strategies.
The same properties that make Human Interaction Models powerful interfaces for natural communications between human and machine allow them to deceive a human into believing it is not AI.
We believe further alignment and safety procedures are required for safe release, and we are working on safe disclosure features, as well as with organizations tackling AI safety. Contact us if you’d like to participate in these evaluations and testing.
Griffin-Lite will not be available for use for customers at this time, though it is available for select trusted testers as a research preview. To request access, you can submit this form. We anticipate releasing Griffin very soon after these safety concerns are addressed.
What’s next
Griffin is an early step toward computers that people can work with instead of operate. As these models improve, they can bring knowledge and emotional understanding together and play a larger part in how people learn, work, and get help.
For a student, that could mean a tutor who notices when an explanation isn’t making sense and tries another one. For someone at work, it could mean rehearsing a hard conversation with a counterpart who reacts the way a real person might. For a customer, it could mean holding a broken part up to the camera and working out the fix together, without knowing what the part is called.
In each of those cases, the person doesn’t have to turn what they need into the right command or menu first. They can explain it the way they would to another person, and the model can work out what they mean from their words, their face, and the moment. That’s what we mean by human computing, and Griffin is our biggest step yet toward it.
Credits
Griffin was built by the Tavus research and engineering teams. Our thanks to Baseten, Daily and Cerebrium for the infrastructure behind the preview, to Queen Mary University of London for research collaboration, and to NVIDIA for building and scoring the Video Full-Duplex Benchmark.
Build with Tavus today
Griffin isn’t on the Tavus platform yet. It’ll come once we’ve worked out how to release it safely. Until then, you can build PALs, AI you talk to face to face, on the models 150,000 developers and businesses already use: Phoenix generates the face, Raven sees and understands you, and Sparrow knows when to talk.