Introducing Griffin

The First Human Interaction Model

Griffin is the world’s first Human Interaction Model (HIM), a new class of model designed to understand and generate face-to-face real-time human interaction. It listens while they talk, and they pay attention to expressions and pauses, not just words. Griffin combines our earlier research into real-time perception, conversation and expressive human-like video into one unified video to video system.

48%
of people believed Griffin was a real person after a one-minute video call.
#1
on NVIDIA’s independent test of face-to-face AI
37%
ahead of the next best AI at reacting in the moment
Illustration of a griffin perched on a mountain ledge overlooking clouds at sunset
By: Hassaan Raza, Co-founder & CEO
Ioannis Patras, Head of Research
Tavus Research Team

Today we’re introducing Griffin, our first Human Interaction Model (HIM). HIMs are a new class of models designed to understand and generate face-to-face real-time human interaction. They listen while they talk, and they pay attention to expressions and pauses, not just words.

Griffin builds on our earlier research into real-time perception, conversation and expressive human-like video. It combines those model capabilities into one, unified video-to-video system.

Griffin is the first model to pass the real-time, video Turing test. In a live study, 48% of participants who talked with Griffin thought they had talked with a real human. Previous systems, including Tavus’s own industry leading conversational video interface (powered by Phoenix 4.5, Sparrow-2, Raven-1), had a max 2% pass rate. This represents a huge breakthrough in human-machine understanding, and is only possible due to Griffin’s video-to-video duplex approach with audiovisual generation, movement, conversational modeling – all in one system, operating in real time.

Griffin-Lite, a research preview, is available today to a select group of early testers, with a wider release of a more powerful model to follow. This post covers what Griffin can do, how we built it, how we measured it, our approach to safety and what comes next.

01

Human computing

A message from Hassaan:

Humans are evolutionarily designed to communicate face to face. We speak as much through our words as we do our expressions, tone, gestures and timing. So much of human communication is non-verbal. It’s an art, a dance.

Machines don’t understand this art. They require us to meet them where they are, to learn their language- the right command, the right button click, the right prompt.

At Tavus, we believe in a future where machines meet us where we are. Where machines learn to communicate in the way that we most effortlessly do. They learn to see, hear, respond and even look like we do. In effect, we want computing to become invisible. We want it to feel second nature.

In a great conversation, you don’t spend your time evaluating the mechanics. You think about what you want to say, what you’re hearing and seeing, and ultimately the conversation just flows. In great conversations, you aren’t thinking about the conversation at all.

With machines, those mechanics are distracting and demand your attention.

Did it hear me? Did it understand me? You share something personal or difficult, and the face continues to smile, or the faceless, cold machine sits there, then responds with a monotonous ‘I’m sorry to hear that’. You pause to collect your thoughts, and it jumps on and responds to incomplete context. It needs time to answer, but nothing in its expression or voice or gesture tells you it’s working on a response.

Each of these moments is you managing the machine. You adjust your timing, simplify your language, make sure to think complete thoughts before saying anything. You remind yourself: don’t leave dead air! Remember to be very clear about what you want! You lower your expectations. The system, in all its intelligence, might be capable of infinite incredible work. It can discover new mathematics dammit! Yet, getting the help for the simple stuff takes more effort than you have to give.

Human communication is an art, a dance. The right move, at the right time. A deep understanding of what you mean, when the words themselves could mean many things. A nod of acknowledgment to signal understanding. An expression that tells you everything without saying anything.

All of these are things we do as humans without thinking about it. These signals let your attention stay on the conversation, not how to converse. As a machine becomes more coherent at expressing and understanding these signals, communicating with it, too, becomes something you don’t have to think about.

02

Introducing Griffin: The First Human Interaction Model

Griffin-Lite is a preview of our first Human Interaction Model: a full-duplex video-to-video model that responds to human behavior in real time. Perception, deciding when and how to respond, and expressive speech and video generation all happen at the same time.

One conversation, two systems
Play again
Skip
Turn-based system
It waits for you to finish, then thinks, then talks.
You
AI
Griffin
It never stops thinking. Every sub-second mini-turn it decides whether to speak up, react or wait, while it keeps listening and watching.
You
Speaks
Thinks
Sees
1
2
3
4
  1. 1
    It says “mm-hm” while you’re still talking.
  2. 2
    It answers without leaving dead air.
  3. 3
    It stops the moment you cut in.
  4. 4
    It reacts to something it sees.

The timing in this diagram is illustrative and was not measured from a real session.

03

Capabilities

To see what it looks like in practice, we brought people who had never used Griffin together with Tavus researchers for open-ended conversations and short demonstrations.

Hover a clip to play it. Click for sound.
Behavior and emotion modeling

Behavior and emotion modelling

Griffin has learned to express behaviors, emotions, and gestures according to conversational context. It laughs, changes its tone, and shifts in response to what the other person says and does.

Rudy gets a promotion
Brittney and the Buckeyes
Clip: Good news

Rudy tells Vanessa he’s been promoted. She reacts before he’s finished, and the two of them talk over each other for a moment without losing the thread.

Full-scene generation

Full-scene generation

Griffin generates every pixel in every frame in real time from one reference image. It controls the whole scene, not just the face, arms, and fingers, but also things like the movement of the chair the person is sitting in, the shadows they cast, and the background behind them.

Mars plays Simon Says
Brittney plays Simon Says
Clip: Simon says

A round of Simon says with Mars. Griffin copies the gesture only when the person says “Simon says,” and calls their bluff when they don’t. Every movement is generated on the spot, body and background included.

Full-duplex conversation

Full-duplex conversation

Griffin continuously evaluates the state of the conversation and independently of conversational turns. It can interrupt, adjust, back-channel, or be interrupted without losing its place.

Ari tells a story
Mars talks about a date
Clip: Story time

Ari and Griffin make up a story together, cutting each other off as they go, and Griffin runs with every twist.

Perception

Perception

Griffin elevates dialogue beyond just verbal understanding, embedding and reacting to visual context and awareness when appropriate.

Ari solves a Rubik’s cube
Mars shows a peacock
Brittney looks cozy
Clip: Cube coach

Griffin coaches Ari through a Rubik’s Cube. It watches the cube as he turns it, keeps acknowledging him while he works, and when he pauses mid-solve it waits until he’s ready.

Temporal understanding

Temporal understanding

Griffin understands time as part of the conversation: how long a silence has lasted, what a sustained silence means, and when to speak again on its own.

Sagar solders a board
Clip: Soldering

Sagar solders a motherboard while Griffin guides him. It keeps track of time and where he is in the job, and it speaks up when the next step is due rather than whenever he goes quiet.

04

Technical approach: Building a full-duplex, video-to-video interaction model

Perception, conversational decision-making, and generation are in continuous interplay throughout human-to-human interaction, and we designed Griffin around the same principle. Central to our approach is a conversational model that controls speech and nonverbal behavior directly, with streaming speech and video generators that follow its decisions as they change. Griffin, is a two part system:

  1. A Continuous Conversational Modeling engine that perceives the incoming audio and video, decides when and how to respond, and produces signals that determine what should be said and how.
  2. An Audio Visual Generation engine comprising a Streaming Speech Generation and a Streaming Video Generation architecture that convert those signals into speech and video, respectively.

Griffin is a two-part system. A Continuous Conversational Modeling engine perceives the incoming audio and video, decides when and how to respond, and produces signals that determine what should be said and how. An Audio-Visual Generation engine, made up of a Streaming Speech Generation and a Streaming Video Generation architecture, converts those signals into speech and video. Perception, decision-making, and generation run concurrently throughout the conversation, so the model can respond to changes while it is speaking as well as while it is listening.

Perception, decision-making, and generation run concurrently throughout the conversation, so the model can respond to changes in the conversation while it’s speaking as well as while it’s listening.

04.1 / System overview
Conversational modeling ingests the person's audio and video and emits expressive controls as it goes. One streaming generator turns those controls into voice, face, gaze and gesture together, as they arrive. Perception never stops while the PAL speaks.
AudioVideoControls
You
Your audio
Your video
audio + video
Griffin
Continuous
Conversational modeling
Reassesses the exchange every sub-second mini-turn
listendecideact
controls
Streaming
Speech + video generation
Voice and face, driven by the same controls
voicefacegaze
audio + video
PAL
PAL audio
PAL video
↺ Perception is continuous, even while the PAL speaks
Controls carry timing, stance, expression and gesture; nothing waits for a turn to end.
Why it matters. Because the conversational model drives speech and video directly, a decision made mid-sentence (pause, look away, smile, interrupt) shows up in the voice and the face within the same mini-turn instead of after a hand-off between separate systems.
System overview. Continuous Conversational Modeling takes in audio and video and emits expressive controls; one streaming generator turns those controls into voice, facial expression, gaze and gesture together, as they arrive.

Continuous Conversational Modeling: (Conversational) intelligence, Perception and Control

Griffin makes conversational decisions at regular sub-second intervals rather than once per turn. At each interval it assesses the state of the conversation, including what has been said and the user’s verbal and nonverbal behavior, and decides what to say next and how to say it. This differs from cascade systems that chain speech recognition, a language model, and speech synthesis, one after the other, and which wait for the user to finish speaking before they begin to respond.

The output controls not only what is said but also how the speech is delivered and how the model behaves nonverbally, including emotional tone, stance, facial expression, and gesture. Because these decisions are made continuously and are conditioned on context, the model can take and yield the turn on the basis of what is being said rather than when the audio stops, so a pause for thought is not treated as the end of the turn. The same mechanism lets the model nod in agreement while the user is speaking, backchannel and confirm at natural points, and adjust the emotional content of its response to the user’s speech and nonverbal behavior.

Additionally, Griffin perceives video as well as audio, where most interactive systems perceive audio alone. Visual input lets the model read the user’s nonverbal communicative signals, such as gaze and facial expression, and it also gives the model access to the user’s environment and to visual material the user chooses to share, such as their screen.

This is illustrated in the figure below.

04.2 / Turn-based vs full-duplex
A cascaded system waits for you to stop, then hands your words through separate models before anything reaches the face. Griffin assesses the state of the conversation at regular sub-second intervals, from audio and video, and can nod, backchannel or begin its reply while you are still talking.
SpeechControlWaiting
Reading the diagram. Ticks on the Griffin lane are mini-turns. At each one the model decides whether to stay quiet, signal (a nod or expression), backchannel, or take the turn, using what it hears and sees. Timings are illustrative; the interactive walkthrough below shows the decisions in detail.
Continuous Conversational Modeling. At regular sub-second intervals Griffin assesses the state of the conversation, including what has been said and the person’s verbal and nonverbal behavior, and decides what will be said and how. It perceives both audio and video, and emits control signals for emotional stance, facial expression and gesture that drive the speech and video models.
04.2 / Full-duplex conversational model
What the model decides, at sub-second intervals
T+0.0s Listening
You
Speech / silence
Control signals
Illustrative exchange. At each sub-second interval the model reads what you said, how you said it and what it can see, then decides what to say (or to stay silent) and sends a control signal that shapes the face and voice.
An illustrative walk through the same loop. At regular sub-second intervals the conversational model reassesses the exchange while the person is still talking, decides whether to speak, stay quiet or act, and produces the words along with control signals for emotion, expression and gesture that the speech and video models act on immediately.

The emitted outputs are converted to streaming audio outputs and together with streaming control signals are driving video generation as explained below.

Fast and Expressive Streaming Speech Generation

04.3 / Streaming speech generation
A fast autoregressive diffusion transformer (VDiT) conditions on an encoded speaker prefix and generates latent speech progressively as control signals arrive from the conversational model. A causal decoder emits audio packets as small as 10 ms, so the PAL starts talking before the sentence is finished.
Audio / latentControlsNot yet generated
In
Speaker prefix
10 s clip, encoded once
Controls
One mini-turn at a time
Generate
Autoregressive diffusion transformer
VDiT
Next latent frame, conditioned on the prefix and everything so far
Latent speech · 100 frames / s
Decode
Tavec decoder
Fully causal
No lookahead, no codebooks. Packets as small as 10 ms.
10 ms packets
Playback
Runs while generation continues
Generate
Play
The first packet plays before the second frame is generated.
Why it matters. Because the model never waits for a full sentence, a control signal that arrives mid-utterance (a change of tone, a hesitation) is reflected in the next frames rather than the next reply. The speaker prefix lets a voice be cloned from a clip as short as ten seconds.
Streaming speech generation. A fast autoregressive diffusion transformer (VDiT) conditions on the encoded speaker prefix and progressively generates the latent representation of new speech as control signals arrive from the conversational model. Generation and playback happen concurrently: each new chunk of latents goes to the decoder as soon as it is available, so speech begins before the complete utterance exists.

Griffin-Lite’s speech generation produces high-quality speech at conversational latency and can clone a speaker’s voice from about 10 seconds of audio. The model is a fast autoregressive diffusion transformer. It takes an encoding of the reference clip as a prefix and generates the new speech progressively, one latent chunk at a time, as increments of outputs and controls arrive from the conversational model, instead of waiting for the full utterance.

A compact continuous codec

04.3 / Tavec codec
Tavec is a convolutional autoencoder that maps 48 kHz audio into a continuous latent: 40 values per frame, 100 frames per second, no codebooks. Its decoder is fully causal, so audio can be streamed the instant a frame exists.
Audio / latentModel
In
Waveform · 48 kHz
48,000 samples per second
Encoder
Convolutional
Compresses time and detail into 40 numbers per 10 ms frame
Latent
40 values per frame × 100 frames / s →
Continuous values, no codebook lookup
Decoder
Fully causal
Streams each frame out as 48 kHz audio, no lookahead
48 kHzSample rate in and out
40 × 100Values per frame × frames per second
6,000Latent frames per minute
Why it matters. A small, continuous latent is cheap for the speech transformer to predict and needs no discrete-token bottleneck, which is what lets VDiT generate frames faster than real time while the causal decoder plays them straight out.
A compact continuous codec. Tavec, a convolutional autoencoder, maps 48 kHz audio into a compact continuous latent: 40 values per frame, 100 frames per second, and no codebooks. Its fully causal decoder carries state across chunks and turns streamed latents into one seamless waveform, in audio packets as small as 10 ms.

Griffin’s speech generation speed and fidelity start with how it represents sound. It relies on Tavec, a convolutional autoencoder that maps 48 kHz audio into a compact continuous latent: 40 values per frame, 100 frames per second, and no codebooks. Its fully causal (streaming-ready) decoder carries state across chunks, turning streamed latents into one seamless waveform, with audio packets as small as 10 ms.

This continuous representation supports straightforward regression-based flow matching and fully differentiable audio pipelines, without quantizer gradient approximations. Its compactness also makes long sequences manageable: a minute of speech corresponds to 6,000 latent frames, against 2.88 million samples of raw audio.

Griffin: Fast and Expressive Streaming Video generation

The other side of Griffin’s audiovisual generation system is fast diffusion-based generation architecture. Previous Tavus models relied on extensive 3D priors, which limited what they could express: complex, human-like behavior such as hand and large body gestures was out of reach, and so were dynamic backgrounds.

We built it to meet three requirements:

  1. Respond quickly to each incoming chunk of audio and expressive controls
  2. Handle controls that change fast
  3. Hold visual quality, lip sync, and identity consistency over long generations

Autoregressive latent video models typically compromise on at least one of these. The architecture we arrived at generates 720p video in 320 ms chunks in real time, one latent at a time, and accepts streaming controls that let the conversation model direct nonverbal behavior such as gesturing, looking away, or changing emotion.

04.4 / Streaming video generation
A few-step autoregressive generator takes a reference photo plus streaming audio and controls and produces a single video latent per step. The VAE compresses time 8×, so at 25 fps one latent is 320 ms of 720p video, which keeps the token count small enough to run in real time.
Video / latentAudioControls
In
Reference image
A single user photo
Streaming audio
From speech generation
Streaming controls
Gesture · gaze · emotion
Generate
Few-step autoregressive generator
3 diffusion steps per latent
Attends to compressed history, predicts the next latent
← older, more compressed · now · next →
Older latents kept at lower resolution, so long rollouts stay cheap
Decode
VAE decoder
8× temporal compression
One latent → eight frames
720p · 320 ms chunk
Played as soon as it decodes
One latent = 8 frames at 25 fps = 320 ms. Fewer tokens per second is what makes it real time.
Why it matters. Strong temporal compression is what makes a diffusion model fast enough to keep up with a conversation, and the streaming controls are what let the conversational model steer a gesture, a glance or a change of emotion into the very next chunk.
Streaming video generation. A few-step autoregressive diffusion generator takes a reference image, streaming audio and streaming controls, and produces a single latent at a time. A VAE compresses the video 8x in time, so at 25 fps each latent covers 320 ms; Griffin-Lite generates 720p video in 320 ms chunks in real time.

To build it, we distilled a large, bidirectional, many-step diffusion model into a few-step autoregressive generator. The generator takes a reference image together with streaming audio and controls, and produces one latent at a time. The distillation ran in three stages. First, we distilled the teacher into a few-step student with Distribution Matching Distillation. This student is fast but not autoregressive; it generates a long video as a single long chunk. Second, we converted it into an autoregressive model with teacher forcing, so as to arrive in an architecture that can generate in a few steps one latent at a time. Third, we trained with Self-Forcing so that we can run long autoregressive rollouts without drift, and we added a mechanism acting on the history frames, the previously generated frames the model continues from, to improve stability.

04.4 / Three-stage distillation
Griffin-Lite starts life as a large, bidirectional, many-step diffusion model. Three stages turn it into a few-step autoregressive generator that takes a reference image plus streaming audio and controls and produces one latent at a time, without drifting over long conversations.
Diffusion step / latentBeing generated
Start
Teacher
Bidirectional, many-step diffusion
Dozens of steps
Sees the whole clip at once. High quality, far too slow.
1
After stage 1
Few-step student, whole video in one chunk
Few steps
Fast, but not streaming: still needs the full clip.
2
After stage 2
Autoregressive, chunk by chunk
Few steps per latent
Continues from the frames it has already generated.
3
Ship
After stage 3
Long rollouts, no drift
Few steps · real time
Trained on its own history frames.
Stage 1
Distribution Matching Distillation
The teacher is distilled into a few-step student. It is fast, but it still generates the whole video in a single chunk.
Stage 2
Teacher forcing → autoregressive
The few-step model is converted to generate one latent at a time, each conditioned on the frames generated before it, so it can stream.
Stage 3
Self-forcing
The model is trained on its own generated history so long autoregressive rollouts stay stable, with an added mechanism acting on the history frames, the frames it continues from, to improve stability.
Result. A generator that produces 720p in 320 ms chunks in real time and accepts multimodal streaming controls, giving the conversational model fine control over non-verbal behaviour instead of audio-led behaviour alone.
Three-stage distillation. A large, bidirectional, many-step diffusion model is distilled into a few-step student with Distribution Matching Distillation, converted into an autoregressive model with teacher forcing, and finally trained with Self-Forcing to perform long autoregressive rollouts without drift.
05

Griffin-Lite: Evaluations and benchmarks

Ahead of this preview, we ran Griffin-Lite through a set of evaluations at different levels, from a single component to a live conversation:

  • The video generator on its own, against published streaming diffusion models, on latency, visual quality, and lip sync
  • The full system on VideoFDB, NVIDIA’s leading industry benchmark for full-duplex audio-visual conversation, which scores how a model reads a person’s behavior and how it responds
  • The full system in live one-minute video calls with people who did not know they were talking to a model

Video Generation

We compared the video generator of Griffin-Lite against four published streaming diffusion models in an audio-to-video setting, where each model is given speech and a reference image and has to produce the talking face.

Griffin-Lite produces one latent at a time and does not wait for future audio, so there is little delay between a piece of audio arriving and its effect appearing on screen. On H100s this averaged 0.43 seconds, half that of the next fastest method. In a conversation, this significantly cuts down the delay before the face reacts when the person interrupts or when a nod is due.

This is depicted in the diagram below, where the dark bar represents the time required by different models to generate the smallest chunk of video they can produce, and pink the time required to playback a video chunk. In the figure, we mark the best average and the worst case latency of each model.

05.1 / Audio-to-video latency
True audio-to-video latency is the time from audio being received to the video showing the effect of that audio. Griffin-Lite generates a single latent with no lookahead audio, which is where most of the win comes from: on H100s it averages 0.43 seconds, half that of the next fastest method.
Griffin-LitePublished streaming diffusion models
Griffin-Litesingle latent · no lookahead0.50×
Strongest published baselinestreaming diffusion · state of the art1.00×
What it buys. Faster responses, quicker interruptions, better-timed backchannels and more natural nonverbal behaviour. Measured in audio-to-video generation mode against state-of-the-art published streaming diffusion models.
True audio-to-video latency, the time from audio being received to the video showing its effect, for Griffin-Lite against four published streaming diffusion models. Griffin-Lite generates a single latent with no lookahead audio and averages 0.43 seconds on H100s, half that of the next fastest method.

Griffin-Lite also scored highest on visual quality. It led every baseline on DOVER and FID, the two standard measures of video quality, and on THEval, a recently published framework built for talking heads. On lip sync it placed second on LSE-C, at 7.27. LSE-C rewards pronounced mouth movement, including movement past the point where it looks natural, which THEval accounts for.

05.1 / Visual quality
Against four state-of-the-art published streaming diffusion models in audio-to-video mode, Griffin-Lite is first on perceptual quality (DOVER), distribution fidelity (FID) and the Talking Head Evaluation framework (THEval), and second on lip-sync confidence (LSE-C), at 7.27.
Griffin-Lite ranks firstRanks second
DOVER ↑
Perceptual video quality
#1
of 5 · best
FID ↓
Frame distribution fidelity
#1
of 5 · best
THEval ↑
Talking Head Evaluation framework
#1
of 5 · clearly best
LSE-C ↑
Lip-sync confidence
#2
of 5 · beats 3 of 4 baselines
LSE-C · 2ndGriffin-Lite, at 7.27, ahead of three of the four baselines.Why not 1stLSE-C rewards pronounced mouth movement, including movement past the point where it looks natural, which THEval accounts for.
How to read it. Arrows show the direction that is better for each metric. Ranks are against the same four published baselines used for the latency comparison; DOVER, FID and THEval are established talking-head video benchmarks, LSE-C measures audio–lip alignment only.
Visual quality and lip sync. Griffin-Lite leads every baseline on DOVER and FID and on the Talking Head Evaluation (THEval) framework, and places second on LSE-C lip sync at 7.27. LSE-C rewards pronounced mouth movement, including movement past the point where it looks natural, which THEval accounts for.

Full-Duplex Conversation

NVIDIA’s VideoFDB is a research and industry benchmark for evaluating the naturalness and accuracy of full-duplex audio-visual conversations. It examines how a model responds to and produces the signals of human interaction, including dialogue, gaze, facial expression, and body movement, measuring both conversation quality and response time. The benchmark scores a model on two tracks, generation and perception, each with its own rubric and its own leaderboard. NVIDIA conducts the evaluation independently using the published metrics and its own judge.

Griffin-Lite scored highest on both tracks against the benchmark’s published baselines, which include commercial and open-source systems— and it is the only case where the same model is evaluated on both tracks.

The generation track tests whether the model produces the right behavior. Given a person’s audio, it scores the model’s speech and video together for fluency, affect matching, and whether the nonverbal behavior fits the moment, laughing with the person rather than after them. Griffin-Lite scored 3.83, over 1 full point ahead of the next-highest system and only 0.09 below the human reference at 3.92, which is more than 12x closer to human performance than any other system evaluated.

05.2 / VideoFDB · Generation
The Generation track scores the system's own voice and face: how natural, fluent and expressive the response is. Griffin lands 0.09 below human ground truth and well ahead of every compared system.
GriffinStrongest baselineHuman ground truth
VideoFDB · Generation track
Human ground truthreference recordings3.92
Griffin3.83
Gemini 2.5 + Anamnext-best system2.80
+1.03
vs next-best system
−0.09
vs human ground truth
62.8%
Takeover-rate alignment · best in class
NVIDIA scored these results in September 2026. A language-model judge rated each response from 0 to 5. Baseline and human-reference scores are reported on NVIDIA's VideoFDB leaderboard. Takeover-rate (TOR) alignment measures how closely a model's response timing matches expected conversational behaviour.
VideoFDB generation track. Griffin scores 3.83 out of 5, 1.03 points ahead of the next-highest model, Gemini 2.5 + Anam at 2.80, and 0.09 points below the human ground truth at 3.92. Baseline and human-reference scores are reported on NVIDIA’s leaderboard.

VideoFDB’s perception track tests the other side of the conversation, which is whether the model understands the moment. Given a person’s audio and video, it scores the model’s spoken response for fluency, conversational flow, and visual grounding, meaning whether the model uses what it sees, a pause with a glance away, for instance, rather than reacting to the words alone. Griffin-Lite scored 3.73, 0.29 points ahead of the strongest second best and 0.47 below the human reference at 4.20, marking Griffin-Lite as the most perceptive model on the benchmark among the 15 models evaluated, ahead of every frontier realtime model on the leaderboard.

05.2 / VideoFDB · Perception
The Perception track asks whether the model understands the moment: given a person's audio and video, it scores the spoken response for fluency, conversational flow and visual grounding, including whether the model uses what it sees.
GriffinBaselinesHuman ground truth
VideoFDB · Perception track
Human ground truthreference recordings4.20
Griffin3.73
MiniCPM-o 4.5strongest reported baseline · audio-only3.44
Gemini 2.5 Flash NativeGoogle3.17
OpenAI gpt-realtimeaudio-only2.97
+0.29
vs strongest reported baseline
−0.47
vs human ground truth
73.8%
Takeover-rate alignment · best in class
NVIDIA scored these results in September 2026. A language-model judge rated each response from 0 to 5; the compared systems include Gemini, OpenAI realtime and open-source models listed on NVIDIA's VideoFDB leaderboard. Griffin perceives audio and video; MiniCPM-o 4.5 and OpenAI gpt-realtime are scored in their audio-only configurations.
VideoFDB perception track. Griffin-Lite scores 3.73, 0.29 points ahead of the strongest second best and 0.47 below the human reference at 4.20: the most perceptive of the 15 models evaluated, ahead of every frontier realtime model on the leaderboard.

VideoFDB also reports takeover-rate alignment, which measures how closely the model’s decisions about when to speak match the timing in the reference conversations. Griffin-Lite scored 62.8% on generation and 73.8% on perception, the highest of any system on both tracks.

Face-to-face study

05.3 / Face-to-face study
Each square is one participant. After a one-minute live video call with a partner they were told was another participant, each was asked whether that partner had been a real person. Filled squares said yes.
Griffin-Lite · said real personPhoenix-4.5 + Sparrow-2 + Raven-1 · said real personSaid AI
Griffin-Lite 54 participants
48%
26 of 54 believed their partner was a real person.
Phoenix-4.5 + Sparrow-2 + Raven-1 previous system · 41 participants
2.4%
1 of 41 believed their partner was a real person.
79%
Confidence · said real person
81%
Confidence · said AI
<20 s
When doubters first suspected
Participants were recruited through an independent research platform and told they would be matched with another participant for a one-minute video call about what they were looking forward to this year. Their partner was a PAL running on Griffin-Lite, generating her face, voice and responses in real time; the Phoenix-4.5 + Sparrow-2 + Raven-1 run used the same protocol. Only at the end of the survey were participants asked whether it had crossed their mind that their partner might not be a real person, and every participant was then told it was an AI.
Face-to-face study. After a one-minute live video call, 26 of 54 participants (48%) believed a PAL powered by Griffin-Lite was a real person. On Phoenix-4.5, 1 of 41 (2.4%) did.

Quantitative benchmarks are important, but some aspects of naturalness are difficult to measure. For every new model, we run a study with participants recruited through an independent research platform to understand how human-like it is in a real, live conversation. Participants have a short video call with the model without being told what it is, and afterward we ask whether they think their partner was a real person. Until now, almost no one has said yes.

On Phoenix-4.5, 2.4% of participants (n=41) believed their partner was a real person. On Griffin-Lite, nearly half did: 48% (n=54). This represents a milestone of, to the best of our knowledge, the first model to have ever passed the video Turing test.

Methodology

Participants were told they would be matched with another participant for a one-minute video call to discuss what they were looking forward to this year. Their partner was in fact a PAL powered by Griffin-Lite, generating her face, voice, and responses in real time. After the call, participants were asked to write down their partner’s answer to the question and to rate their partner on several axes around naturalness and trustworthiness. They were also asked to rate the conversation itself: how well it flowed, and whether they felt their partner was really listening. Only at the end of the survey were participants asked whether it had crossed their mind that their partner might not be a real person, and if so, when. Every participant was then told that their partner had been an AI model.

Results

Of the 54 participants who spoke with Griffin-Lite, 26 believed their partner was a real person. Participants on both sides were confident in their answer: those who said real averaged 79% confidence, and those who said AI averaged 81%.

Over half of participants said the possibility had not crossed their mind during the call, and nearly all of that group went on to say their partner was real. Those who did suspect tended to suspect within the first 20 seconds.

Griffin-Lite was rated favorably on every axis. On a 7-point scale, participants on average rated it a 5.4 for seeming natural, 5.6 for seeming trustworthy, and 5.8 for whether they would enjoy talking with it again, a score that held at 5.4 among participants who said it was AI. The conversation scored 5.5 for whether participants felt their partner was really listening and 4.9 for flowing naturally, the lowest of the five.

06

Safety, limitations, and responsibility

Given Griffin-Lite is the first model to pass the Turing test, we understand the responsibility to be thoughtful on the risks and release strategies.

The same properties that make Human Interaction Models powerful interfaces for natural communications between human and machine allow them to deceive a human into believing it is not AI.

We believe further alignment and safety procedures are required for safe release, and we are working on safe disclosure features, as well as with organizations tackling AI safety. Contact us if you’d like to participate in these evaluations and testing.

Griffin-Lite will not be available for use for customers at this time, though it is available for select trusted testers as a research preview. To request access, you can submit this form. We anticipate releasing Griffin very soon after these safety concerns are addressed.

07

What’s next

Griffin is an early step toward computers that people can work with instead of operate. As these models improve, they can bring knowledge and emotional understanding together and play a larger part in how people learn, work, and get help.

For a student, that could mean a tutor who notices when an explanation isn’t making sense and tries another one. For someone at work, it could mean rehearsing a hard conversation with a counterpart who reacts the way a real person might. For a customer, it could mean holding a broken part up to the camera and working out the fix together, without knowing what the part is called.

In each of those cases, the person doesn’t have to turn what they need into the right command or menu first. They can explain it the way they would to another person, and the model can work out what they mean from their words, their face, and the moment. That’s what we mean by human computing, and Griffin is our biggest step yet toward it.

Credits

Griffin was built by the Tavus research and engineering teams. Our thanks to Baseten, Daily and Cerebrium for the infrastructure behind the preview, to Queen Mary University of London for research collaboration, and to NVIDIA for building and scoring the Video Full-Duplex Benchmark.

Citation
Tavus Research, “Griffin: The First Human Interaction Model”, Tavus, October 2026.