Written by: Minh Anh Nguyễn and Haiyao Xaio

[This is a research and technical overview of Phoenix-4.5. To jump to what the model unlocks, click here]

See the official announcement on X and LinkedIn.


Realism is not just how an individual looks. It is how they move.

Humans communicate with more than our faces. Our shoulders shift, our posture changes, our heads follow the rhythm of our speech, and we react while someone else is talking. These small, mostly unconscious movements are much of what makes someone feel present rather than rendered, and they are what conversation is built on.

For the last three years, our Phoenix model has been closing that gap on naturalness:

  • Phoenix-1 made it possible to model a person in 3D and drive them in conversation.
  • Phoenix-2 broke the real-time barrier and made the Conversational Video Interface possible.
  • Phoenix-3 expanded generation from the mouth to the entire face.
  • Phoenix-4 made that face emotionally intelligent, with controllable emotion, active listening, and continuous facial motion in a single unified system.

Today, we’re pushing the boundary of what’s possible again.

Introducing Phoenix-4.5, our biggest step forward in real-time human rendering yet — moving Phoenix from rendering a face to rendering a robust-bodied PAL.

The difference is immediately visible. Facial movement is more natural. Expressions are richer. Emotion feels more alive. And movement now extends through the head, shoulders, posture, and torso — so the whole PAL moves together while speaking and listening.

PALs don’t just look more real. They feel dramatically more natural to talk to.

And Phoenix-4.5 does it in real time, at 134ms from audio to video — 25% faster than any other model on the market.

Three things define Phoenix-4.5

Phoenix-4.5 improves the three things that matter most for making a PAL feel real: how naturally they behave, how much they look like themselves, and how quickly they can respond.

1. A huge jump in naturalness.

Phoenix-4.5 is the most natural real-time human rendering model on the market. Phoenix-4 — and other models — still feel rigid. Phoenix-4.5 changes that. Facial animation is richer, emotion is more expressive, micro-movements feel more natural, and importantly, motion now extends through the head, shoulders, posture, and torso. Because Phoenix-4.5 generates the full frame rather than animating a facial region in isolation, upper-body movement becomes part of expression — and you can feel it.

2. The PAL actually looks like themselves, and far more faces work.

Phoenix-4.5 has the strongest identity preservation we’ve seen in a real-time human rendering model. The details that make someone look like themselves carry through, even as the PAL becomes dramatically more expressive.

And it’s a far more tolerant model when it comes to creating custom faces. Long hair, glasses, earrings, headbands, and other characteristics that historically caused training failures now work far more reliably. In our tests, 64% of Faces that failed on Phoenix-4 trained successfully on Phoenix-4.5.

3. All of it still happens faster than ever.

Phoenix-4.5 goes from audio to rendered video in 134ms, making it the fastest real-time human rendering model on the market — while generating the full frame rather than just the face. Milliseconds matter when it comes to immersion.

And there’s another major change under the hood: Phoenix-4.5 is the first zero-shot Phoenix model. A new identity can render from a single image or video almost instantly, and then automatically refine afterward, unlocking faster iteration and building.

See Phoenix-4.5 in action
Try the demo

From face to PAL: the whole frame is generated now

When we speak, our bodies move with us. Our heads follow the cadence of our words, our posture shifts with emphasis, and our shoulders and torso move subtly with the rhythm of speech. These movements aren’t separate from expression — they’re part of it.

This is the architectural change that makes Phoenix-4.5 feel so different.

Phoenix-4 generated a region of roughly 512×512 pixels around the face and placed it onto recorded footage of the body. That design worked inside the boundary of the face and constrained everything outside it: the body stayed largely static, and movements led to exposing artifacts at the edge of the composited region.

Phoenix-4.5 removes that boundary entirely. We built a new generative renderer from scratch that generates the whole frame together in a single pass — face, head, neck, shoulders, torso, clothing, and surroundings — rather than treating the face as a separate layer. There is no mask, no seam, and no face boundary for motion to expose. That is what buys the model its freedom to move.

Now the head, shoulders, posture, and torso can move with the cadence of speech, so expression flows through the whole PAL rather than stopping at the face. Lips follow speech, facial expression follows meaning and intonation, the head follows cadence, and subtle torso movement follows the rhythm of speaking. We move with the words we say, and now PALs do too.

The breakthrough isn’t simply more movement. It’s coordinated movement that follows the conversation. A face can look incredibly real, but if the head, shoulders, posture, and torso don’t move with the way that PAL speaks and listens, something still feels off. Phoenix-4.5 lets expression belong to the whole PAL.

Human behavior, not just animation

A PAL doesn’t become still when they stop talking. In a real conversation, we keep reacting while we listen — our expression changes, our head and posture shift, and small movements signal that we’re still present in the exchange.

Phoenix-4.5 makes those behaviors noticeably more natural. Facial movement has more rhythm, expressions are richer, emotions feel more alive, micro-expressions are more natural, and lip sync is tighter. It’s not just a larger area moving; the animation itself is better.

Phoenix-4 introduced active listening and continuous facial motion so a PAL could remain expressive across the conversation, not just while generating speech. Phoenix-4.5 carries that behavior forward while improving the expression itself.

Underneath the model, the animator listens to both sides of the conversation — the PAL’s speech and the user’s — and continuously answers a simple question: how should this PAL be moving right now?

The result is expression that carries through speaking, listening, reacting, and the transitions between them, rather than switching between “talking” and “idle.”

Every metric that should move, moved

The difference is visible, but it also shows up in the numbers. We evaluated Phoenix-4 against the fine-tuned Phoenix-4.5 model on the same paired internal test set.

ModelAVSR ↑CSIM ↑ΔPFID ↓FVD ↓
Phoenix-40.2610.9140.46428.36160.81
Phoenix-4.50.3040.9120.60525.42148.48

Lip sync improves 16% over Phoenix-4. Motion vividness increases roughly 1.3×, making the increase in head movement measurable. Identity fidelity holds essentially level with Phoenix-4, while frame and video quality both improve.

Metrics only tell part of the story. Ultimately, what matters most is how natural the model feels to a person.

To measure that, we use Elo scoring based on human preference: evaluators compare outputs head-to-head and choose which feels more natural. Phoenix-4.5 shows a clear step forward. Even zero-shot Phoenix-4.5 scores 1002 Elo, ahead of fine-tuned Phoenix-4 at 929. Once fine-tuned, Phoenix-4.5 climbs to 1069.

In other words, the gains aren’t just visible in individual benchmarks. People can feel the difference.

Creation got a lot easier

Phoenix-4.5 changes both sides of creating a Face: you can see a new PAL almost immediately, and far more people can become one in the first place.

The wait before you can see anything is gone

Phoenix-4.5 is the first Phoenix model that can render a new identity zero-shot, with no per-identity training required upfront. Previously, a bespoke model had to finish training before you could see the PAL at all. Now, a preview is ready in about a minute while fine-tuning runs in the background. Roughly two hours later — half the time Phoenix-4 required — the fine-tuned Face takes over.

Zero-shot is a preview, not the finished product. Fine-tuning still closes the remaining gap in fine detail and identity fidelity. What changes is the workflow: from wait, then see to see, then refine.

Far more people and characters can become PALs now

Real human appearance is messy. Hair falls in front of shoulders, glasses cross facial features, earrings move, and accessories overlap the face. With a standard face-isolated approach, those details created difficult training cases — and could even require someone to change how they looked just to create a high-quality Face.

Because Phoenix-4.5 generates those regions together, many of those failure cases disappear. We retrained a set of identities that had previously failed on Phoenix-4: 64% trained successfully on Phoenix-4.5. Most of the remaining failures came down to source-footage issues like full-body framing or moving backgrounds rather than the individual’s appearance.

Now, creating a PAL is faster, less restrictive, and works for far more people than before.

Phoenix-4Phoenix-4.5
Generation areaA roughly 512×512 region around the face and headA full frame
What movesThe face and head, composited onto recorded footage of the bodyFace, head, neck, shoulders, torso, and background, generated together
ExpressionFull-face emotion and head pose, controllable in real timeFull-face emotion, head pose, and upper-body motion, with richer micro-expression
Motion between wordsMotion lives in the face; the body holds stillHead, shoulders, and posture keep moving with the rhythm of speech
ListeningActive listening in the faceListening carries through the face and the upper body
Rendering approach3D Gaussian splatting plus a neural clean-up passA learned generative (GAN-based) half-body network
Motion sourceExpression latents that drive the faceRicher motion latents that drive the face, head, and shoulders together
A new identityNeeds a bespoke trained model before it can renderRenders zero-shot in about a minute, then refines
Refinement timeAbout 4 hoursAbout 2 hours
Identity constraintsHair tucked behind the shoulders; accessories often failedLong hair, glasses, earrings, and headbands supported

Presence changes what you can build

Realism isn’t the end goal. The interaction is.

A Face isn’t just a visual layer on top of an AI. Conversation is a flow, and we’re constantly reading signals beyond the words themselves — whether someone is listening, how they’re reacting, when they’re about to respond, and the emotion or emphasis behind what they say.

There’s a reason so many important interactions have historically happened face to face. In education, healthcare, coaching, sales, and support, a person isn’t just delivering information — how they listen, react, explain, and connect is part of what makes the interaction work.

Phoenix-4.5 puts more of that on screen. A PAL that moves with its words, listens visibly, and answers as the words begin gives the person on the other side the same cues they would read in a room, and those cues are what carry an interaction through a full lesson, a hard conversation, or a real sales call. The face stops being the thing you notice and becomes the reason the conversation works.

And we’re already seeing that show up in the outcomes. Builders using Tavus to replace audio-only interactions with face-to-face conversational AI are seeing:

  • 2–3× higher engagement across sales and healthcare
  • 40% higher knowledge retention in learning and development
  • 50% faster ramp + 3× higher rep conversion in sales training

That’s why we keep pushing Phoenix forward. Presence is part of the product. The standard for human rendering isn’t whether a Face looks real in a still frame. It’s whether the PAL continues to feel human and natural as the conversation unfolds.

Under the hood: the research

Phoenix-4.5 is two models working together, both rebuilt for this release.

The animator answers a deceptively hard question dozens of times a second: given what has just been said, how should this PAL be moving right now?

It listens to two audio streams, the PAL’s own speech and the user’s, and a streaming diffusion transformer turns them into compact motion codes, block by block, with emotion, style, and identity conditioning injected throughout. The trained model is distilled down to a couple of sampling steps, which is the difference between a research result and something that can hold a live call.

The renderer is where the face becomes a PAL. It takes one reference image and the animator’s motion, reads the change between reference and target rather than treating motion as an absolute instruction, and paints the full half-body frame jointly. Motion signals from the face region drive coherent movement across the head, neck, and shoulders, with no compositing anywhere in the pipeline.

Together they run comfortably in real time: audio-to-video latency under 130 milliseconds, around 35–40 frames per second end-to-end on slower machines.

Upgraded existing faces and new faces

Every existing Tavus stock Face has been upgraded to Phoenix-4.5 — and we’re introducing 50 brand-new stock Faces. Existing Faces get the full jump in naturalness, expression, and upper-body movement, while the new set expands the range of people, styles, and personalities available out of the box.

The goal isn’t just a bigger library. It’s to make it easier to find the right Face for what you’re building — whether that’s a tutor, coach, concierge, interviewer, salesperson, or support PAL — and start with a production-ready human immediately.

Choose from the expanded stock library, or create your own Face with Phoenix-4.5.

What Phoenix-4.5 unlocks for real-world use cases

Better rendering matters most in experiences where how someone shows up is part of the experience itself. Those are the experiences that break first when a face goes stiff between words, and they are where builders on Tavus already see face-to-face AI outperform text and audio. Phoenix-4.5 raises the ceiling on every one of them.

Sales agents that hold a prospect through the whole conversation

A PAL on the front line of sales has to hold a prospect’s attention for the length of a real conversation, react to an objection while it is still being spoken, and bring the same energy to the fiftieth call of the day as the first. Phoenix-4.5 gives it a body that leans into the pitch and a face that answers as the words begin. Teams running AI sales agents on Tavus already report 2× the conversation engagement of their previous chatbot tools and 6× more meetings booked.

Healthcare agents that visibly listen

Intake, patient education, and follow-up are conversations where stillness reads as indifference. Someone explaining something difficult needs a face that visibly listens, and Phoenix-4.5 carries that listening through the head and shoulders instead of freezing between replies. Patients already prefer it: 70% choose face-to-face AI over text or audio for health conversations, and in elder care 97% chose video over audio so decisively that audio was removed entirely. See healthcare agents on Tavus.

Interviewers that candidates treat as real

Screening and role-play only work when the PAL across from you feels real. Candidates mirror what they see, so a stiff interviewer gets stiff answers, and the evaluation data is only as good as the conversation that produced it. Phoenix-4.5 moves, reacts, and holds emotion the way an interviewer in the room would. On Tavus AI interviewers, 75% of candidates come back for a second session or more, and Colleva cut hiring costs 70% with AI video interviews at scale.

Coaches that make practice feel real

Practice only builds skill if it feels real. A PAL coaching a rep through objection handling or a nurse through a hard conversation has to react to hesitation, hold attention through a full session, and feel like a live interaction rather than narrated content. Phoenix-4.5 is the difference between a training video with a face on it and a scene partner. Teams using Tavus for L&D report 300% faster ramp-up with AI sales coaching and a 54% increase in confidence after practice sessions.

The same holds for concierge, support, and kiosks, and for anywhere else a face is the interface. These are the experiences that break first when rendering feels stiff, and Phoenix-4.5 is built for them.

The bar is face to face

Phoenix has always been built toward a simple standard: can an AI feel natural when you’re actually talking to it face to face?

Phoenix-4.5 is our biggest step toward that yet. Not just a Face that looks real, but a PAL that moves with the conversation — speaking, listening, reacting, and expressing itself as one continuous person.

Phoenix-4.5 is live today for every PAL through the PAL Maker and API.