
A photorealistic depiction of a live video conversation in which nothing visibly reveals whether the person on the screen is human or AI. AI-generated image via ChatGPT (OpenAI)
Tavus Griffin: The Human-Like AI That 48% of People Thought Was Real
Tavus has introduced Griffin, its first Human Interaction Model, a research-preview AI designed to see, hear and respond during live video conversation. In Tavus's one-minute study, 48% of participants believed Griffin was a real person.
Real-time conversational avatars already exist. What Griffin is trying to add is the ability to treat the full flow of a conversation—what someone says, when they pause, how they look, whether they interrupt, and the gestures or expressions they make—as information the AI can use to understand what is happening and decide how to respond.
That creates a new decision point for businesses, developers, policymakers and the people interacting with these systems. If AI can convincingly communicate through the same social cues people already use with one another, knowing whether someone is human may no longer be something we can reliably determine from the interaction itself. The question becomes where human-like AI belongs, how it should identify itself and which interactions should remain distinctly human.
Instead of simply listening to what someone says and generating a response, Griffin-Lite is designed to see and hear the other person continuously while deciding what to say, when to say it and how to react. A pause, a confused expression, an interruption or a gesture can therefore become part of what the AI understands about the conversation—and part of its response.
That could allow AI to occupy more roles that once required another person's presence. AI tutors, coaches, customer-service agents and other digital representatives could communicate through the same kinds of social cues people already understand, reducing the need for users to translate what they naturally mean into the prompts, commands or menu choices a computer expects.
The same behaviors that make these systems easier to communicate with are also behaviors people use to recognize, understand and trust one another. Human-like interaction may become another capability AI can reproduce rather than reliable evidence that another human is present. As that happens, identity, disclosure and boundaries around machine participation may have to become explicit instead of relying on people to recognize AI from the interaction itself.
For now, Griffin-Lite remains a research preview available only to select trusted testers and is not available to Tavus customers. Tavus itself says a system that people can mistake for a person has to be released carefully.
To understand why Griffin pushes this issue forward, however, we first have to separate it from something that can look similar at first glance: the real-time AI avatars that already exist.
Key Takeaways: Tavus Griffin, Human-Like AI and AI Identity
Tavus Griffin is a research-preview Human Interaction Model designed to continuously see, hear and respond during live video conversations, making human-like social cues part of the AI’s conversational intelligence rather than simply part of an animated avatar.
Griffin differs from traditional real-time AI avatars by using audio, video, timing, gestures and facial expressions as part of both understanding the conversation and deciding how to respond.
NVIDIA’s VideoFDB benchmark found Griffin-Lite led the other AI systems tested on overall perception and audiovisual generation, although human performance remained higher overall.
In Tavus’s one-minute study, 48% of 54 participants believed Griffin-Lite was a real person, but the result has not been independently replicated and does not show that people would mistake Griffin for a human in longer or different interactions.
Human-like AI could make digital assistants easier to use because people may be able to communicate through natural speech, expressions, gestures and timing instead of translating everything into prompts or commands.
Human-like AI could also affect jobs in areas such as tutoring, coaching, training and customer service, forcing businesses to decide whether AI should supplement workers, handle tasks alongside them or replace portions of human work.
Griffin-Lite is not currently available to Tavus customers because the company says additional safety, alignment and disclosure work is needed before broader release.
AI disclosure, identity and content authenticity are separate problems: a system can disclose that AI is present, verify who or what an AI agent represents, or prove that media is synthetic without necessarily solving the other two.
Governments and standards organizations are already developing rules and technical systems for AI disclosure, provenance and agent identity, including the EU AI Act, California disclosure laws, C2PA, NIST and emerging IETF work.
The larger implication of Griffin is that human-like interaction may become something machines can reproduce rather than reliable evidence that another human is present, making AI identity and the boundaries of machine participation explicit choices for businesses, platforms and policymakers.
How Griffin Differs From Real-Time AI Avatars Like HeyGen LiveAvatar
Griffin did not invent the real-time conversational AI avatar.
HeyGen's LiveAvatar, for example, already supports two-way conversations with lifelike AI avatars that can listen and respond through voice, video or text with low latency. The system generates synchronized speech, facial expressions and gestures, and developers can connect it to external large language models to power the conversation.
HeyGen's own process for creating a custom LiveAvatar helps show how that type of system works. Training footage is divided into separate Listening, Talking and Idling sections. During the listening portion, the person being recorded is instructed to behave as someone naturally would while another person is speaking—smiling, nodding, raising their eyebrows or otherwise appearing engaged. Those behaviors give the finished avatar realistic ways to look attentive while the user talks.
Griffin is attempting to move some of that behavior deeper into the conversation itself.
Instead of generating speech and then wrapping a convincing visual performance around it, Griffin receives the user's audio and video as conversational input. What the person says matters, but so can what Griffin sees happening while they say it. The system then continuously generates both its spoken and visual response.
In an independent benchmark of conversational video systems, NVIDIA describes Griffin as a full-duplex audiovisual-to-audiovisual system, meaning it can take in both what a person says and what it sees while continuously producing speech and video in return. NVIDIA contrasts that with systems that first decide what to say and then generate an avatar's movements around the speech.
That difference matters because vision is not simply being used to make Griffin look more realistic. What Griffin sees can influence what it thinks is happening in the conversation, while its visual behavior can become part of the response. A look, nod, gesture or reaction does not have to be added after the AI has already decided what to say.
Existing systems can already conduct convincing real-time conversations while presenting an AI through a realistic human avatar. Griffin is trying to make audiovisual behavior part of the conversational intelligence itself.
And once those signals become part of the conversation, the interaction can no longer be treated as a simple sequence of one person speaking and the other waiting for a turn.
How Griffin Uses Full-Duplex Conversation to React in Real Time
Most digital conversations still follow a familiar pattern: you say something, the system processes it, and then it responds. Griffin is designed to keep interpreting the interaction while the conversation is still happening.
Tavus describes Griffin as having two broad parts. Continuous Conversational Modeling takes in audio and video, interprets what is happening and decides whether Griffin should speak, stay quiet or respond another way. Audio-Visual Generation turns those decisions into streaming speech and video. Rather than running those steps one after another, perception, decision-making and generation operate at the same time.
That means Griffin does not make one decision at the end of each conversational turn. Tavus says the system continually reassesses what is happening at sub-second intervals, including while Griffin itself is speaking. It can listen while talking, stop when interrupted, interrupt when appropriate, distinguish a thinking pause from the end of someone's turn, offer small acknowledgments such as “mm-hm,” nod without speaking and react to visual information. It can also change its emotion, facial expression and gestures as the interaction develops.
That is what “full duplex” looks like in practice: the AI does not have to wait politely for a clean handoff before it can begin reacting.
Tavus demonstrates that difference with a task that has little to do with ordinary question-and-answer conversation. In one example, Griffin watches someone soldering a motherboard. Instead of waiting for the person to stop working and ask what comes next, Griffin tracks what is happening and how much time has passed, then speaks when the next step is due rather than simply responding whenever the person becomes silent.
Another demonstration shows Griffin playing Simon Says. When the person gives a gesture instruction after saying “Simon says,” Griffin follows it. When the person tries the same kind of instruction without those words, Griffin recognizes the trick instead of simply copying the movement. The body movement and surrounding scene are generated in real time as the interaction unfolds.
Human communication contains information while nobody is saying anything—and while somebody else is still speaking. Griffin is not only responding to the last sentence it heard. It is trying to keep track of what is happening between the people in the conversation, including information that may arrive while nobody is speaking or while Griffin is already producing a response.
The output is built to keep changing along with that understanding. Griffin's speech is streamed in very small pieces, while its video is generated in chunks of about 320 milliseconds. That allows its speech and visual behavior to continue developing as the conversation unfolds rather than waiting to generate a finished performance first.
And Griffin is not only animating a face. Tavus says the system generates every pixel of every frame in real time from a single reference image, including the body, arms, hands, fingers and background. Because the entire scene is being generated, gestures and physical movement can become part of the response rather than being limited to lip movements or a set of facial animations.
That changes what counts as an AI response. A response can still be words, but it can also be a nod while someone continues speaking, a facial reaction to something shown on camera, a pause because the other person appears to be thinking, or a gesture that happens at the right moment.
In other words, the difference is between AI responding to what you said and AI participating in what is happening between you.
That distinction sounds impressive in a company demonstration. The more important question is whether independent testing can detect a measurable difference between Griffin and other conversational systems.
NVIDIA’s VideoFDB Benchmark Finds Griffin Leads Conversational AI Systems but Still Trails Humans
Tavus's demonstrations show what Griffin is designed to do. An independent benchmark gives us another way to ask whether that architecture actually makes a measurable difference.
NVIDIA and David AI developed VideoFDB, which NVIDIA describes as the first benchmark designed specifically to evaluate full-duplex audio-and-video conversational agents. Instead of testing whether an AI can simply answer questions about a video, VideoFDB examines whether it can use visual and verbal signals continuously during a two-way conversation.
The benchmark contains 237 clips from real video calls covering 11 conversational dynamics, including pauses, interruptions, backchannels, laughter, gaze, facial expressions and body behavior. Those 11 dynamics are the types of conversational moments being tested; some test whether the AI correctly understands another person's behavior, some test the behavior the AI produces in response, and some are used for both.
VideoFDB separates those two sides into perception and generation. Perception measures whether the AI understands what is happening, using scores for fluency, conversational flow and visual grounding. Generation measures the audiovisual response the system itself produces, using fluency, emotional matching and whether its nonverbal behavior is appropriate and well-timed.
Each applicable category is scored from 0 to 5, with higher scores representing better conversational behavior. Because natural conversation does not always have one objectively correct response, NVIDIA uses an LLM to judge responses against detailed scoring guidelines rather than looking for one predetermined answer.
NVIDIA also tested the reliability of that approach using three different LLM judges. When the judges independently evaluated the same responses, their scores fell within one point of each other 77% to 89% of the time, depending on the category. That does not make an LLM judge infallible, but it suggests the scoring was reasonably consistent across different judges.
Perception: How Well the AI Reads the Interaction
On perception, Griffin-Lite scored 3.73 overall, ahead of the other audiovisual conversational models NVIDIA tested, although still below the 4.20 human reference.

The human reference did not score a perfect 5 because 5 is not a score for “being human.” NVIDIA took responses from real human conversations and graded them with the same scoring system used for the AI models. Real people do not always make the response that a benchmark considers ideal, and an LLM judge is still deciding the score. That is why the human reference can land below 5.
NVIDIA then ran a separate comparison to see what happened when some audiovisual models could no longer see the video. This was not a rerun of every model on the leaderboard. NVIDIA reported audio-only comparisons for four of the speech-and-vision systems:

Giving some systems eyes did not necessarily make them better conversationalists. In several cases, perception scores actually improved when the video was removed, while Gemini 2.5 Flash Native scored exactly the same either way.
NVIDIA's analysis found two recurring problems in how these speech-and-vision models used visual information.
In captioning collapse, the AI treats the incoming video more like something it has been asked to describe than part of an ongoing conversation. Instead of naturally adjusting its behavior based on what it sees, it may begin responding with observations about the person or what they appear to be doing.
In visual-stream ignorance, the opposite happens: the model receives the video but barely uses it. NVIDIA found that some audiovisual responses were close to paraphrases of the model's audio-only responses, meaning that seeing the person did little to change the conversation.
NVIDIA presents these as current failure modes, not solved problems. Part of the difficulty is that conversational visual signals can happen extremely quickly. The speech-and-vision models NVIDIA tested generally processed video at about one frame per second, while cues such as interruptions or gaze changes can unfold within one or two seconds. Simply feeding the model more frames did not solve the problem either; in NVIDIA's testing with MiniCPM-o 4.5, performance eventually declined as the amount of visual information increased.
Griffin's results suggest that integrating audiovisual information more deeply into the conversation can help, but the benchmark does not establish a universal solution to these problems.
That helps explain why Griffin's architecture matters. Tavus did not simply give an existing speech model access to a camera. Griffin was designed around audiovisual interaction from the beginning, allowing what it sees to affect the ongoing conversation and allowing its own visual behavior to become part of the response.
Generation: How Well the AI Produces an Audiovisual Response
The generation test asks a different question: not how well the AI reads another person's behavior, but how well it produces its own spoken and visual behavior in response.
That also explains why the model lineup changes. The standalone Gemini, OpenAI and MiniCPM systems in the perception test produce audio responses rather than video of a conversational participant, so NVIDIA cannot evaluate them on facial expressions, gestures or other nonverbal output. For generation, NVIDIA compared Griffin with systems that can actually produce both the spoken response and a visible avatar.
Human ground truth scored 3.92 overall, while Griffin-Lite scored 3.83. Two systems that generated an avatar from Gemini 2.5's speech scored considerably lower.

The human score is below 5 for the same reason as in the perception test: NVIDIA graded real human behavior using its benchmark rules rather than automatically giving humans a perfect score.
NVIDIA points to an architectural reason for the gap between Griffin and the speech-to-avatar systems in the generation test. When an avatar's movement follows speech that has already been generated, the avatar is effectively waiting for that speech before it knows how to move.
That makes those systems effectively turn-based in their motion generation. They cannot naturally insert an independently timed nod, facial reaction or other nonverbal cue while the user is still speaking because the avatar's movement is being created around speech that has already been generated.
In NVIDIA's generation timing evaluation, Gemini 2.5 + Anam had a median response timing of 2.84 seconds and Gemini 2.5 + Keyframe took 3.52 seconds, compared with 0.9 seconds for human ground truth.
That finding reinforces the distinction between Griffin and the real-time avatar architecture discussed earlier. A system can generate highly realistic avatar behavior and still be limited if that behavior has to follow generated speech. Griffin is instead designed so speech and visual behavior can both participate continuously in the interaction. NVIDIA did not test HeyGen LiveAvatar in VideoFDB, so the benchmark should not be read as a direct performance comparison between Griffin and HeyGen.
Those results help explain why Griffin's integrated design performed better than the speech-to-avatar systems NVIDIA tested, but they do not mean Griffin has matched human conversation. The benchmark's more detailed generation scores show where the remaining gaps are.
On dyadic affect, which measures how well the system's emotional response matches the person it is interacting with, Griffin scored 4.40 compared with 4.14 for human ground truth. But humans remained ahead on nonverbal-cue appropriateness, scoring 3.18 compared with Griffin's 2.83.
Timing remained another gap. In NVIDIA's generation timing evaluation, Griffin's median timing was 1.892 seconds, compared with 0.9 seconds for human ground truth.
NVIDIA also measures whether a system makes the expected timing choice—such as staying quiet during a thinking pause, continuing when appropriate or yielding when interrupted. On that measure, called Takeover-Rate Alignment, Griffin had the highest nonhuman score on both parts of the benchmark: 73.8% for perception and 62.8% for generation. Humans still scored higher.
The benchmark therefore supports a narrower—and more useful—conclusion than saying Griffin has reached human conversation. It has not. The stronger evidence is that combining perception, conversational decisions, speech and nonverbal behavior into one continuous interaction produces measurable differences from systems that add vision or an avatar around a primarily speech-based conversation.
But NVIDIA's benchmark was not designed to answer the question Tavus tested next: Would the person on the other side of the conversation realize they were talking to AI?
Tavus Study: 48% of Participants Thought Griffin Was Human
NVIDIA's benchmark found that Griffin outperformed the other AI systems it tested on overall perception and generation. Tavus then asked a much more direct question: Would people realize they were talking to AI?
Tavus recruited 54 participants in the United States and Europe through an independent research platform. Participants were told they had been paired with another participant for a one-minute video conversation about what they were looking forward to that year.
Their conversation partner was Griffin-Lite.
Afterward, 26 of the 54 participants—48%—said they believed they had spoken with a real person. Those who thought Griffin was human reported an average confidence level of 79%. Participants who correctly identified Griffin as AI were similarly confident, averaging 81%.
Tavus also found that people who became suspicious generally did so within the first 20 seconds of the conversation.
The company ran the same experiment with its earlier Phoenix-4.5 system. In that test, only 1 of 41 participants, or 2.4%, believed the AI was human. Tavus uses that contrast to argue that Griffin represents a significant jump in how convincingly an AI can participate in live video conversation.
Tavus also describes the Griffin result as the first successful real-time video Turing test, but that characterization belongs to the company and needs some qualification.
This was a Tavus study involving 54 Griffin participants and one-minute conversations. It does not show that 48% of the general public would mistake Griffin for a human in longer conversations, across different topics or under different conditions. We also have independent benchmark evidence from NVIDIA showing Griffin's conversational strengths, but we do not have an independent replication of Tavus's 48% human-identification result.
Even with those limits, the result is still significant under the conditions Tavus tested. Human appearance combined with human-like conversational behavior was convincing enough that nearly half of participants did not realize they were interacting with AI.
That changes what visual interaction can tell us. A video call has historically carried a simple assumption: if a person appears to be looking at you, reacting to you and speaking with you in real time, another person is probably there. Griffin suggests that assumption may no longer be reliable on its own.
That sounds unsettling until we ask the other half of the question: Why would we want AI capable of communicating this way in the first place?
Why Human-Like AI Interaction Could Make Digital Assistants More Useful
The value of Griffin is not that it can fool people into thinking it is human. The value is that people may no longer have to communicate with AI like a machine.
Tavus describes this broader idea as human computing: instead of forcing people to translate what they need into the right prompt, command or menu option, the computer adapts to the way people already communicate.
In education, that could mean an AI tutor noticing that an explanation is not landing and changing how it explains the concept. In workplace training, someone could rehearse a difficult conversation with a counterpart that reacts naturally rather than following a fixed script. In customer support, a person could simply hold a broken object up to the camera and talk about the problem without first needing to know the name of the part that failed.
Those examples show why human-like behavior is not merely cosmetic. Eye contact, pauses, facial expressions, tone and gestures can carry information that would otherwise have to be converted into words before the computer could use it.
Based on the capabilities Tavus has demonstrated, similar forms of interaction could potentially support:
Language practice, where a learner could practice face-to-face conversation and receive responses shaped by both what they say and how the interaction is going.
Coaching, where the system could respond to hesitation, tone or visible uncertainty instead of relying only on the words someone enters.
Role-play and simulation, where someone could practice a presentation, interview or difficult conversation against a counterpart that reacts as the exchange develops.
Some accessibility applications, where speech, visual information and nonverbal cues could give people additional ways to communicate with a system.
Customer support, where an AI could recognize visible or audible frustration and adjust the interaction without requiring the customer to explicitly describe how they are feeling.
Education, where a tutor could notice signs of confusion and try a different explanation before the student has to say, “I don't understand.”
These are potential applications based on Griffin's demonstrated capabilities, not claims that Griffin is currently deployed for all of them.
The larger benefit is reducing the amount of translation people currently perform between what they naturally mean and what a computer knows how to receive. Instead of stopping to formulate the perfect prompt, explain every visible detail or wait for the system to decide that it is finally your turn to speak, people could communicate more like they already do with one another.
But making AI capable of communicating naturally in these settings also raises an obvious workforce question. Tutors, coaches, trainers and customer-service representatives are jobs currently performed by people. If AI systems become capable of occupying more of those roles, businesses will have decisions to make about whether the technology supplements human workers, handles certain interactions alongside them or replaces some human work altogether.
For the people doing those jobs—and for customers who may prefer or need a human—the distinction is significant. A system being technically capable of performing an interaction does not by itself answer whether that interaction should be automated, when a human should remain available or which parts of a role depend on something beyond the ability to communicate convincingly.
Those questions become even harder because the technology's usefulness and its broader implications come from the same underlying capability. The same signals that help AI understand people are also signals people use to understand one another. If machines become good at reading and reproducing those signals, natural communication gets easier while those same signals become less reliable evidence that another human is actually there.
Making an AI deliberately stiff, awkward or obviously unnatural would make it easier to identify as a machine. But doing that could also remove part of what makes this kind of technology useful in the first place.
The challenge, then, is allowing AI to communicate naturally without leaving people to guess whether they are interacting with a human or a machine.
A nod could still mean, “I understand.” A pause could still signal that the system is waiting. A concerned expression could still communicate attention or recognition. But those behaviors may no longer tell us who—or what—is producing them.
As human-like AI improves, people may need something other than the interaction itself to know whether another human is present.
Tavus appears to recognize that problem. And it is one reason Griffin is not currently being released broadly.
Why Tavus Is Keeping Griffin in Research Preview While It Builds AI Disclosure Safeguards
Griffin-Lite is currently available only to select trusted testers as a research preview. Tavus has not made it available to its customers.
The company explicitly connects that decision to Griffin's ability to be mistaken for a person. Tavus says, “A model people can mistake for a person has to be released carefully.”
That is an important acknowledgment because the qualities that make Griffin useful—natural timing, emotional reactions, eye contact, gestures and other familiar conversational behavior—can also make deception easier if people do not know they are interacting with AI.
Tavus says Griffin needs additional alignment and safety work before a broader release. The company is developing safe-disclosure features and says it is working with AI-safety organizations as it determines how Griffin should identify itself during interactions. Tavus says it expects to release Griffin more broadly after those concerns have been addressed.
Its broader principle is straightforward: “The mechanics of talking to a machine should disappear. The fact that it is a machine should not.”
That distinction gets to the heart of the problem. Natural interaction does not require hidden identity. An AI can speak fluidly, react to facial expressions and communicate through familiar social cues while still making it clear that the participant on the other side is artificial.
The harder question is much simpler: How should the AI tell you that it is AI when you can no longer reliably tell from the interaction itself?
We are already seeing early versions of that transparency problem addressed with AI-generated content. Meta labels some AI-generated media, LinkedIn displays Content Credentials on supported images and video, Google is expanding tools that help people identify how digital content was created or edited, and OpenAI uses provenance signals such as Content Credentials and SynthID.
A live AI participant may need something more direct. A persistent label on the screen, an audible disclosure or another clear signal could tell the person that the participant is AI even when its behavior feels completely natural. Technical identity and provenance systems could provide additional verification behind the scenes.
But those approaches are solving different problems. A watermark or Content Credential can help establish whether media was artificially generated. An identity system can help establish who—or what—is participating. A disclosure has to make sure the person actually knows they are interacting with AI.
And Tavus is not alone in confronting those questions. Governments and standards organizations are already beginning to treat AI-generated content, AI identity and disclosure as distinct but related problems—separating the question of whether content is synthetic, who or what is participating, and whether the person interacting with it actually knows the difference.
How Laws and Standards Are Starting to Identify AI Content, Agents and Interactions
The question of how people know they are interacting with AI is no longer theoretical. Governments and technical standards organizations are already building different ways to disclose AI participation, identify synthetic content and establish the identity of AI agents.
The European Union's AI Act makes an important distinction between those problems. Under Article 50, providers of AI systems designed to interact directly with people must ensure users are informed that they are interacting with AI unless that fact is already obvious to a reasonably informed and observant person in the circumstances.
Separately, the law requires certain AI-generated or manipulated audio, image, video and text to be machine-readable and detectable as artificially generated or manipulated, subject to the law's exceptions.
That separation matters. Telling someone that the participant they are talking to is AI is not the same problem as identifying whether a piece of media was generated by AI.
California has begun drawing similar boundaries in more specific settings.
Under SB 896, California state agencies that use generative AI to communicate directly with people about government services or benefits must disclose that AI is being used. The required disclosure depends on the medium: written chatbot disclosures must remain visible throughout the interaction, audio systems must provide a verbal disclosure at the beginning and end, and video interactions must display the disclosure prominently throughout. The law also requires agencies to provide information about how a person can contact a human state employee.
California has also addressed AI-generated people in advertising. SB 1050, enacted in September 2026, requires disclosures on certain video or audio advertisements that use AI-generated performers to sell products or services.
That is particularly relevant as realistic AI people move beyond generated clips and toward interactive roles. Policymakers are already considering situations in which a synthetic person can appear in a commercial interaction even when no human performer is actually present.
Technical standards are also developing alongside those legal requirements.
The Coalition for Content Provenance and Authenticity, or C2PA, develops Content Credentials that can carry verifiable information about where digital media came from and how it was created or changed.
Version 2.3 extended that system to live video. Because a live stream arrives in small pieces rather than as one finished video file, C2PA created a way for individual segments of the stream to carry verification information. Software can check those segments as the video plays to confirm that they belong to the expected stream, remain in the correct sequence and have not been altered or substituted.
Version 2.4 added a separate AI Disclosure Assertion, giving software a standardized, machine-readable way to recognize information about AI involvement in digital content. That could allow a platform or other application to automatically detect AI transparency information and decide how to present it to users.
The important limitation is that machine-readable disclosure is not the same thing as telling a person they are looking at AI. The technical signal can provide the information behind the scenes, but a platform or interface still has to make that information understandable and visible to the person if disclosure is the goal.
Those systems primarily help answer questions about the content itself: where it came from, whether it has been altered and whether AI played a role in creating it.
But an interactive AI creates another problem: What is the identity of the participant?
In February 2026, the U.S. National Institute of Standards and Technology (NIST) launched its AI Agent Standards Initiative to help develop common standards and protocols for AI agents that increasingly act across outside systems and data.
The initiative includes three broad areas: supporting industry-led agent standards, encouraging open technical protocols that different agents and systems can use together, and advancing research into AI-agent security and identity.
NIST is also examining how existing identity and access-management practices could apply specifically to AI agents. That work asks practical questions such as how a system identifies which agent is acting, what that agent is authorized to access or do, how its actions can be audited, and how organizations can prevent an agent from exercising authority it was never given.
Work is also beginning within the Internet Engineering Task Force (IETF). A September 2026 Internet-Draft proposes an AI Identity Management System, or AIMS, that would use established identity technologies—including OAuth 2.0 and related standards—to authenticate AI agents and define what they are authorized to do.
The proposal is still a work in progress, not an established universal standard. But its approach is notable because it does not assume AI needs an entirely separate identity system. Instead, it asks how existing digital identity infrastructure can be extended so an AI agent can have verifiable credentials, defined permissions and a clear relationship to the person or organization it represents.
Taken together, these efforts are beginning to separate three questions that can easily get blurred together:
Authenticity: Is this media genuine, altered or AI-generated?
Identity: Who—or what—is participating in the interaction?
Disclosure: Does the person actually know they are interacting with AI?
Solving one does not automatically solve the others.
A Content Credential might help verify that a video was generated by AI without ensuring the person watching it notices that fact. An authenticated AI agent could prove which company or person it represents without clearly announcing itself to the person on the other side. And a visible “AI” label could provide disclosure without proving anything about the underlying identity or provenance of the system.
As human-like AI becomes harder to identify through appearance and behavior alone, identity may increasingly have to be communicated through something other than perception.
But even perfect disclosure would not resolve the larger question Griffin raises.
Because what happens when everyone knows they are talking to AI—and they are perfectly comfortable with it?
What Happens When People Know They’re Talking to AI — and Accept It?
Disclosure can solve one important problem: making sure people know when AI is participating.
But disclosure does not answer a different question: What happens when everyone knows its AI—and nobody feels deceived?
If human-like AI becomes useful enough, people may willingly choose it for some interactions. An AI tutor could be available whenever a student needs help. A coach could rehearse the same difficult conversation repeatedly without getting tired. A customer-service representative could be available immediately instead of placing someone on hold. A conversational partner could help someone practice a language or prepare for an interview.
In those situations, the issue is no longer whether the AI successfully passed as human. The person may know exactly what it is and still prefer the interaction.
Businesses could face similar choices. If AI systems can communicate naturally enough to handle some customer-facing conversations, companies may increasingly use AI representatives in roles that once required a person to be present. California’s new advertising disclosure law already reflects one version of that future by addressing AI-generated performers used to sell products and services.
As AI agents gain more context and delegated authority, the idea could extend beyond customer service. An AI might eventually represent a person or business in routine digital interactions—answering questions, handling administrative tasks or even participating in some meetings on someone’s behalf.
That raises a different set of questions:
Where do people actually want AI participation?
Which interactions are improved by an always-available AI tutor, coach, assistant or representative?
Which situations still benefit from—or require—a human?
Could an AI eventually represent a person in routine digital interactions, including attending some meetings on their behalf?
What authority should an AI participant have when it is acting for a person or business?
Would genuinely human interaction become something businesses actively distinguish or market?
Could “human-made” or “human-to-human” eventually become a meaningful service characteristic precisely because AI participation has become commonplace?
That last possibility may sound strange now, but it follows from the same pattern. If AI participation becomes common, human involvement itself could become a meaningful characteristic of a service—much the way businesses already distinguish handmade products, live support or locally provided services from automated alternatives.
None of that means human-like AI will replace human interaction altogether. It means the boundary may become something people, businesses and platforms have to choose more deliberately.
At that point, the central question is no longer simply how do we keep people from being fooled?
It becomes: Where do we want machine participation, what authority should AI participants have, and which experiences do we want to preserve as distinctly human?
Human beings may not be able—or even want—to keep AI out of every social space. The harder task may be establishing how humans and AI share those spaces, and where we decide the boundary belongs.
That brings us back to the larger question Griffin raises—not just about avatars or disclosure, but about what human presence means when human-like interaction itself can be reproduced by a machine.
What This Means: Human-Like AI Makes Disclosure, Work Boundaries and Machine Participation Explicit Decisions
Human-like AI turns machine participation into a decision businesses can no longer treat as merely a technology choice. If Griffin-like systems become common, businesses will not simply be choosing a more natural AI interface. They will be deciding when AI is allowed to stand in for a person. That affects customer service, education, training, coaching and other interactions where people have historically assumed that a responsive face, voice and set of social cues meant another human was present.
For businesses, that creates a practical tradeoff. Human-like AI could provide faster service, broader availability and interactions that require less effort from the user. But companies will also have to decide when AI should identify itself, when customers should be able to reach a human, which parts of a job should remain human-led and whether using an AI representative is appropriate for the relationship they have with their customers. As AI agents gain authority to act on behalf of people and organizations, the decision expands from who handles the conversation to what the AI is actually permitted to do.
Workers and customers are affected differently by the same capability. For workers, increasingly capable AI representatives could change job responsibilities, supplement some roles or replace portions of work that currently require a person. For customers, the issue is not only whether an AI can perform the task, but whether they know AI is performing it and whether they still have a meaningful choice to interact with a human. If AI participation becomes commonplace, businesses may eventually find that clearly offering “human-to-human” service becomes a differentiator rather than the assumed default.
That puts pressure on platforms, developers and policymakers as well. Once people cannot reliably identify AI from appearance or behavior alone, disclosure and identity have to become features of the system rather than something users are expected to figure out for themselves. A convincing interaction cannot be the proof of who is on the other side. Labels, authenticated agent identities, provenance systems and clear rules about delegated authority may increasingly become part of the infrastructure around human-like AI.
Griffin does not establish how quickly that future will arrive. It remains a research preview, and Tavus's 48% result came from 54 participants in one-minute conversations and has not been independently replicated. That makes it too early to conclude that people will routinely mistake human-like AI for people in longer, everyday interactions.
But the larger decision does not depend on AI becoming perfectly indistinguishable from humans. The important change is that human presence can no longer be assumed simply because an interaction looks and feels human. As that distinction becomes less obvious, businesses and society will have to decide where machine participation is useful, where human involvement should remain protected or available, and how people are told which one they are dealing with.
The goal does not have to be making AI less natural so humans can recognize it. The harder challenge is making AI identity unmistakable while preserving the benefits of natural interaction—and deciding which parts of human life we actually want machines to participate in.
Q&A: Tavus Griffin and Human-Like AI Interaction
Q: What is Tavus Griffin?
A: Griffin is Tavus’s first Human Interaction Model, designed to see, hear and respond continuously during live video conversations. Instead of waiting for someone to finish speaking before generating a response, Griffin can use speech, visual information, pauses, interruptions, facial expressions and gestures as part of the ongoing interaction. Griffin-Lite is currently a research preview and is not available to Tavus customers.
Q: Did 48% of people really think Griffin was human?
A: Yes, under the conditions Tavus tested. Twenty-six of 54 participants, or 48%, said they believed they had spoken with a real person after a one-minute video conversation with Griffin-Lite. However, this was a Tavus study using short conversations, and the result has not been independently replicated. It does not mean 48% of the general public would mistake Griffin for a human in longer or different interactions.
Q: How is Griffin different from AI avatars like HeyGen LiveAvatar?
A: Real-time conversational avatars already exist. Griffin’s main difference is that audiovisual behavior is built into the conversational intelligence itself. What Griffin sees can affect how it understands the conversation, while its facial expressions, gestures and other visual behavior can become part of its response rather than simply being generated around speech the AI has already decided to produce.
Q: What does “full duplex” mean in an AI conversation?
A: Full duplex means Griffin can keep listening, watching and reacting while the conversation is still happening—even while Griffin itself is speaking. It can stop when interrupted, recognize a thinking pause, respond nonverbally, offer small acknowledgments while someone continues talking and adjust its behavior as the interaction changes.
Q: Did Griffin outperform other conversational AI systems?
A: In NVIDIA and David AI’s VideoFDB benchmark, Griffin-Lite led the other AI systems tested on overall perception and audiovisual generation, although humans still performed better overall. Griffin scored 3.73 on overall perception compared with a 4.20 human reference, and 3.83 on audiovisual generation compared with 3.92 for human ground truth. It still lagged humans in areas including timing and nonverbal-cue appropriateness.
Q: Did Griffin really pass a video Turing test?
A: Tavus describes its experiment as the first successful real-time video Turing test, but that is the company’s characterization. The study involved 54 Griffin participants speaking with the system for one minute each, and the 48% result has not been independently replicated. The experiment shows that Griffin convinced nearly half of participants under those specific test conditions, not that it can routinely pass as human in any conversation.
Q: Why would anyone want AI that communicates this much like a person?
A: The practical benefit is that people may not have to communicate with AI like a machine. Instead of translating what they mean into the right prompt, command or menu option, they could communicate using speech, expressions, gestures and other familiar cues. Potential applications include tutoring, coaching, language practice, workplace role-play, customer support and some accessibility uses.
Q: Could human-like AI replace jobs?
A: It could affect some work. If AI systems become capable of handling interactions that currently depend on human communication, businesses will have to decide whether AI supplements workers, handles certain tasks alongside them or replaces portions of human work. Roles such as tutoring, coaching, training and customer service could be affected, while customers may also want or need the option to reach a human.
Q: Is Griffin available to customers now?
A: No. Griffin-Lite is currently available only to select trusted testers as a research preview. Tavus says additional alignment, safety and disclosure work is needed before a broader release because a model that people can mistake for a person has to be released carefully.
Q: How will people know they are talking to AI if it looks and acts human?
A: There may not be one solution. Visible or audible disclosures can tell a person that AI is participating, provenance systems can help establish how digital media was created, and identity systems can help verify who or what an AI agent represents. These solve different problems: authenticity asks whether content is genuine or AI-generated, identity asks who or what is participating, and disclosure asks whether the person actually knows AI is involved.
Q: Are there already laws requiring AI to identify itself?
A: In some situations, yes. The EU AI Act requires certain AI systems that interact directly with people to inform users they are interacting with AI unless that is already obvious. California also requires disclosures in certain government AI interactions and in some advertisements using AI-generated performers. Technical standards including C2PA, NIST’s AI-agent work and emerging IETF proposals are addressing related questions involving provenance, identity and authorization.
Q: What is the bigger significance of Griffin and human-like AI?
A: Human-like interaction may become something machines can reproduce rather than reliable evidence that another human is present. If that happens, businesses, platforms and policymakers will increasingly have to decide where AI participation belongs, how clearly AI must identify itself, what authority AI representatives should have and where people should still have access to—or expect—the presence of another human. The challenge is not necessarily making AI less natural; it is making AI identity unmistakable while deciding which parts of human life we actually want machines to participate in.
Sources:
Tavus: Griffin: The First Human Interaction Model
https://www.tavus.io/griffinTavus: 48% of People Thought This AI Was a Real Human | Introducing Griffin
https://youtu.be/lHw6yoyPkpoNVIDIA Research: VideoFDB — Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
https://research.nvidia.com/labs/amri/projects/video-fdb/HeyGen Help Center: Introducing LiveAvatar
https://help.heygen.com/en/articles/12758516-introducing-liveavatarHeyGen Help Center: LiveAvatar FAQ
https://help.heygen.com/en/articles/12758866-liveavatar-faqHeyGen Help Center: LiveAvatar: Custom LiveAvatar Creation Guide
https://help.heygen.com/en/articles/9612935-liveavatar-custom-liveavatar-creation-guideEuropean Commission AI Act Service Desk: Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems
https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-50California Legislative Information: SB-896 Generative Artificial Intelligence Accountability Act
https://leginfo.legislature.ca.gov/faces/billNavClient.xhtml?bill_id=202320240SB896Office of Governor Gavin Newsom: Governor Newsom Signs New Law to Protect Workers, Require Disclosures on AI-Generated Advertising
https://www.gov.ca.gov/2026/09/16/governor-newsom-signs-new-law-to-protect-workers-require-disclosures-on-ai-generated-advertising/C2PA: Content Credentials: C2PA Technical Specification (Version 2.3)
https://spec.c2pa.org/specifications/specifications/2.3/specs/C2PA_Specification.htmlC2PA: Content Credentials: C2PA Technical Specification (Version 2.4)
https://spec.c2pa.org/specifications/specifications/2.4/specs/C2PA_Specification.htmlNIST: Announcing the “AI Agent Standards Initiative” for Interoperable and Secure Innovation
https://www.nist.gov/news-events/news/2026/02/announcing-ai-agent-standards-initiative-interoperable-and-secureInternet Engineering Task Force (IETF): AI Identity Management System
https://datatracker.ietf.org/doc/html/draft-ietf-wimse-aims-00arXiv: VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
https://arxiv.org/html/2605.30256v2
Editor’s Note: This article was created by Alicia Shapiro, CMO of AiNews.com, with writing support, AEO/GEO/SEO optimization, image concept development, and editorial structuring support from ChatGPT, an AI assistant. All final editorial decisions, perspectives, and publishing choices were made by Alicia Shapiro.
