Skip to content
News

Nearly half of people on a video call with Tavus’s Griffin AI thought they were talking to a human

Tavus, a company that sells tools for building AI characters you talk to face to face, has unveiled a model called Griffin. In Tavus’s own study, 26 of 54 people who spent a minute on a live video call with it (48%) thought they had been talking to a real person. Tavus says that makes Griffin the first model to pass a video Turing test, the old idea that a machine passes when people cannot tell it from a human. A preview called Griffin-Lite is open to a few trusted testers and not to customers.

The question for most readers is whether you can still trust that the person you see and hear on a video call is real. How much the 48% tells you about that depends on how the study was run.

What the study did

Tavus recruited people through a research platform. Each of them was told they would have a one-minute video call with another participant, about what they were looking forward to this year. That “participant” was Griffin-Lite, an AI that generated a woman’s face, voice and answers live.

After the call, the participants rated their partner. Only at the very end were they asked whether it had crossed their mind that the partner might not be a real person. Then they were told it was an AI.

Of the 54 people who spoke to Griffin-Lite, 26 said their partner was a real person. That is the 48%. More than half said the thought of an AI never occurred to them during the call, and nearly all of that group believed their partner was real. The people who did get suspicious tended to get suspicious within the first 20 seconds.

For comparison, Tavus ran the same test on its previous system, which combines three of its own models: Phoenix-4.5 for the face, Sparrow-2 for timing and Raven-1 for reading the person. Only 1 of 41 people (2.4%) took that one for a real person.

48%?

Alan Turing proposed his test in 1950. Researchers usually run it by putting a judge in front of a person and a machine at the same time and asking which is which. A UC San Diego study published this year found that GPT-4.5, when prompted to act like a person, was picked as the human 73% of the time in text chats run that way.

Tavus’s test was looser, in four ways.

  1. First, nobody was looking for an AI. The participants thought they were talking to a person and were never asked to choose between a person and a machine.
  2. Second, there was no comparison group. The write-up has nobody who spoke to a real person, so there is no figure for how often an actual human on the same call would have been called real.
  3. Third, the sample was small. With 54 people, the true rate could sit anywhere from about 35% to 61%, by our calculation using a 95% Wilson interval.
  4. Fourth, Tavus ran and wrote up the study itself, and its page links no paper or data for it.

What is different about Griffin

Most real-time AI works as a relay. One system turns your speech into text, a language model writes a reply, and other systems turn that reply into a voice and a face. Each handoff adds delay and drops things the next system never sees, such as your tone or what is on camera.

Griffin does all of it in one system. It listens and watches while it talks, and every fraction of a second it decides whether to nod, say “mm-hm”, speak or wait. Tavus calls this “full duplex”: both sides can talk and react at the same time, as people do. So Griffin can interrupt you, be interrupted, and react to something it sees. It also generates every pixel of the picture from one reference photo, down to the chair and the shadows, and Tavus says it can clone a voice from about 10 seconds of audio.

In the clips Tavus posted, Griffin plays Simon Says and copies a gesture only when the person says “Simon says”. In another it watches a man solve a Rubik’s Cube and waits when he stops to think.

Tavus’s argument is that AI calls feel robotic when the face keeps smiling while you say something hard, or when there is a silence with no sign the machine is working on a reply. Griffin is built to remove both. A late answer and a face that does not react are also tells you would look for on a call.

What NVIDIA’s benchmark shows

Tavus also claims first place on VideoFDB, a benchmark from NVIDIA and David AI. It uses 237 clips of real video calls, and an AI judge scores each response from 0 to 5 on two tracks. The generation track marks the AI’s own speech and video. The perception track marks whether it reads what is happening in the conversation.

Griffin-LiteBest rivalPeople
Generation3.832.803.92
Perception3.733.444.20

Griffin-Lite does lead both tracks, with two caveats:

  1. On generation, only three AI systems were scored, and the other two are relays built on Google’s Gemini 2.5, so the 1.03-point lead is a lead over relays.
  2. On perception, the best rival is an audio-only run of MiniCPM-o 4.5, and Tavus’s “15 models” are really eight systems, seven of them scored twice, once with video and once with audio only.

The same board lists Griffin-Lite’s median response time on the generation track at 1,892 milliseconds, against 900 for people.

No release date

Tavus gives no release date and no price for Griffin. It does not describe the “safe disclosure features” it says it is building. It does not say where the 54 participants lived, how old they were or which platform recruited them. Access to Griffin-Lite is by a request form, with no criteria stated. Nothing on the page mentions Kenya.

Tavus says models like Griffin can deceive a person into believing they are not talking to an AI. It says more safety work is needed before release, and it expects to release Griffin soon after that.

For many readers the risk is fraud. In August we wrote that voice-cloning tools need only seconds of audio. A model like Griffin adds a moving face and live reactions to what a caller can fake, once anyone ships one.

The advice from that piece still stands. A clone can copy how someone looks and sounds, but it cannot know a word your family agreed in private. Hang up and call back on the number you have saved, and agree a code word before anyone sends money.

Our view is that 48% is not a Turing test result in the sense researchers use, and it matters anyway. Scams run on short calls with people who are not looking for a machine, which is the condition Tavus tested. For now we wait for the release date.

The Analyst

The Analyst delivers in-depth, data-driven insights on technology, industry trends, and digital innovation, breaking down complex topics for a clearer understanding. Reach out: Mail@Tech-ish.com

Join the discussion

0 comments
posting as Samaki Mwepesi

Anonymous by default — no sign-up or email needed. Prefer to be recognised? Add a name or email above, your call. We don't email you about replies, so do check back.

protected, no CAPTCHAs
Back to top button