Who Do You Think You're Talking To?
Same system, same output — the label alone changes how people experience it.
LLM
Study
Role
AI & Robotics Researcher
Timeline
IRB-approved · n=60 · 3 sessions over 2 weeks
team
With Vito Borio
platform
Web

The Real Problem
People increasingly talk to language models that are hard to tell apart from a person. Who Do You Think You're Talking To? grew out of a simple question: when the words on the screen are identical, does simply believing you are talking to a human — rather than an AI — change how the exchange feels? The study isolates that belief from the behavior it usually travels with.
And here's what makes it slippery: the label and the behavior almost always move together. In the real world an AI usually sounds different from a person, so you can never tell whether your reaction is to what was said or to who you thought said it. To separate the two, the partner has to stay the same while only the label changes.
The problem with most claims about talking to AI is easy to state:
They compare a real AI against a real human — two different partners — so any difference could just be a difference in what each one actually said.
People's expectations about AI are strong, but no one had held the conversation itself fixed to see how much of the effect is the expectation alone.
Nothing rules out the obvious confound: that a labeled human simply communicates differently from a labeled AI in the first place.
The underlying issue isn't a lack of opinion about AI — it's that no study had *held the partner constant*: same model, same replies, only the label flipped. Belief is not the same as behavior.

Finding the Fix
Who Do You Think You're Talking To? keeps the conversation partner fixed and moves only the belief. Every participant actually talked to the same language model each time, but was told before each session that the partner was either another person or an AI. Three ideas shape the design:
One partner, two labels. Behind every conversation was the same language model; only the label attached to it changed.
Belief measured, not assumed. Participants rated each interaction on an adapted Relational Communication Scale, so the effect of the label is a number, not an impression.
The same people, over time. Each of the sixty students returned for three sessions across two weeks, so we could see whether the effect faded as the partner became familiar.
The distinction that matters is between what is said and who you think is saying it. Two people can read the very same reply and experience it differently depending only on the label in front of it.
So the study is built around three commitments:
Same system throughout. One language model answered every participant every time, so any difference in experience cannot come from the partner behaving differently.
A repeated, longitudinal design. Sixty students held three conversations over two weeks, letting us test whether belief keeps mattering once the novelty wears off.
Measured, not argued. Ratings on an adapted Relational Communication Scale and a mixed MANOVA turn a claim about transparency into an effect you can actually test.
The aim is to measure belief on its own — the effect of the label, cleanly separated from the behavior.

What Actually Happened
Sixty college students each held three conversations over two weeks. Before each one they were told their partner was either another person or an AI, though in every case the partner was the same language model. The design was IRB-approved · n=60 · adapted Relational Communication Scale · mixed MANOVA.
The same replies, presented under a human label, were rated as more immediate, more equal, and more relational than when they were presented as coming from an AI.
Holding the model identical across every participant and every session was the decision that made the result interpretable — it is what lets the label, rather than the behavior, carry the difference.
The gap between the human-labeled and AI-labeled conditions held across all three sessions rather than shrinking as participants grew familiar with their partner.

What Changed
With the partner held constant, believing you were talking to a human made the interaction feel more immediate, more equal, and more relational — and with n=60 across three sessions, the effect held stable rather than fading with familiarity.
The more interesting signal is what this separates:
Belief and behavior come apart. Same output, different label, different experience — the reaction was to who people thought they were talking to.
Familiarity did not erase it. Three sessions in, the human label still felt warmer and more equal, so this is not just a first-impression novelty effect.
It names a real tradeoff: the transparency of disclosing an AI comes at a measurable cost to how engaging and relational the interaction feels.
The point is not a single headline number but a clean separation of belief from behavior — an IRB-approved, n=60 study whose effect held stable across three sessions.

What I Had to Work With
A controlled setup, not the wild. A controlled study gives a clean read on belief, but the conversations are still short lab interactions, so the setting had to be kept consistent enough that only the label varied. Every participant met the same model under the same conditions; only the human-or-AI label was allowed to differ.
Consistency matters. The comparison is only meaningful if every session is comparable — the same partner, the same rating instrument, and the same procedure each time. Responses were measured on an adapted Relational Communication Scale and analyzed with a mixed MANOVA across the three sessions.
Scope. The study measures how the interaction is experienced, not task performance or accuracy — it is about the effect of belief, not about which partner is better.
These constraints shaped the design: keep the partner constant, keep the procedure consistent, and measure belief rather than behavior.

What I'd Do Differently
Extend the design beyond a college-student sample and beyond three short sessions, to see how far the belief effect generalizes.
Vary the kind of conversation as well as the label, to map where disclosing an AI costs the most engagement and where it costs the least.
What I Learned
Isolating belief beats arguing about it. The value of Who Do You Think You're Talking To? is the clean comparison — same model, same output, only the label changing — rather than one more opinion about how people feel about AI.
Belief is not behavior. The words can be identical and still land differently depending only on who people think is speaking — the interesting signal lives in that gap.
The label is a design choice with a cost. Disclosing an AI is more transparent, but it measurably lowers how immediate, equal, and relational the exchange feels — the transparency-versus-engagement tradeoff, measured rather than argued.
Who Do You Think You're Talking To? asks how much of our experience of an AI comes from the AI itself — and how much from simply being told that is what it is.