The study, published on September 2 in Science Advances, found that AI-generated surrogates performed better than chance when predicting how specific individuals would respond to questions and social science experiments, but were still wrong about one-quarter of the time.
By Frank Ulom
·
Published on September 4, 2026
·
6 min read
Artificial intelligence digital twins designed to replicate individual human behaviour may not yet be reliable substitutes for real research participants, with a new study finding that the systems can distort the views and decision-making patterns of the people they are meant to represent.
The study, published on September 2 in Science Advances, found that AI-generated surrogates performed better than chance when predicting how specific individuals would respond to questions and social science experiments, but were still wrong about one-quarter of the time.
Researchers described the phenomenon as a “funhouse mirror” effect, in which artificial intelligence reflects some characteristics of a person while altering or exaggerating others.
The findings could have implications for social scientists who hope to use AI-generated human subjects, also known as digital twins, to reduce the cost and burden of behavioural research.
Human participants can be expensive to recruit, become tired during lengthy studies and may experience psychological distress when exposed to sensitive experiments. Digital twins, by contrast, can respond indefinitely without fatigue and could potentially allow researchers to conduct preliminary experiments without repeatedly involving human participants.
However, Olivier Toubia, a computational social scientist at Columbia Business School in New York City, said the results showed that the technology remains limited.
“There’s some promise,” Toubia said. “But [the twins’ performance] was overall a bit disappointing.”
How the AI twins were created
The research built on a dataset created by Toubia and colleagues in a previous project reported in Marketing Science.
More than 2,000 people from across the United States participated in that research, answering more than 500 questions covering a broad range of personal characteristics and behaviours.
The questions examined areas including age, ethnicity, income, education, religious practices, political preferences, personality traits, spending habits, mathematical abilities and vocabulary.
Participants also completed online tests designed to examine their thought patterns and cognitive biases.
The researchers subsequently made the dataset available as an open-source resource for other researchers.
According to Toubia, the dataset has been downloaded about 25,000 times.
For the new study, researchers supplied the information collected from each individual to a large language model (LLM), instructing the system to respond as though it were that particular person.
The resulting digital twins were then tested across 19 social science experiments.
The experiments examined a range of issues, including how people respond to individuals who donate to both Republican and Democratic political candidates and how they perceive algorithmic hiring.
Twins got individuals partly right
The researchers found that the digital twins could capture some differences between individuals better than AI systems provided only with demographic information.
For example, if one person rated themselves a two for self-control and another gave themselves a four, an LLM working with only demographic information might predict a middle value for both individuals.
The digital twins, which had access to more detailed personal information, were more likely to produce different predictions for the two people.
Although those predictions were not necessarily accurate, they provided a better indication of variation between individuals.
Despite this advantage, the twins were wrong on average about 25 per cent of the time and performed approximately as well as chatbots that had been given demographic information alone.
The researchers identified several factors behind the inaccuracies.
The digital twins tended to produce more homogeneous responses than the humans they represented. Their answers were also frequently influenced by demographic stereotypes.
The accuracy of their predictions improved among participants who were more affluent and better educated, suggesting that the technology did not perform equally well across different groups.
AI twins appeared more rational
The study also identified systematic differences between the digital twins and their human counterparts.
The AI surrogates tended to express greater trust in other people and showed less concern about technological threats. They also appeared to make decisions in a more rational manner than the humans they were designed to represent.
Hadi Hosseini, an AI researcher and economist at Pennsylvania State University, said he had observed a similar pattern in his own research examining how AI agents make healthcare decisions when resources are limited.
He said large language models tend to distort human judgement towards decisions that appear more rational or reasonable than those actually made by people.
According to Hosseini, the limitations may partly stem from the way the digital twins were trained.
The researchers relied on a relatively static set of questions to build profiles of participants. Hosseini suggested that more dynamic approaches could produce better representations.
These could include having an AI system observe an individual’s behaviour throughout the day or engage in regular conversations with the person to gather more comprehensive information.
Digital twins could still aid research
Despite the shortcomings, researchers said digital twins could have practical uses in social science research.
Toubia noted that AI surrogates could be useful when researchers require detailed responses to questions. While human participants may become tired and provide brief answers during lengthy studies, digital twins can generate much longer responses without fatigue.
Digital twins could also be used to pretest research experiments before real participants are recruited.
Such preliminary testing could help researchers identify problems with an experiment’s design without consuming the limited time and attention of human respondents.
Toubia said he intends to explore more sophisticated methods for training digital twins, but cautioned researchers against expecting synthetic data to perfectly reproduce human behaviour.
Social scientists, he argued, may sometimes assume that surveys and psychological scales adequately capture the full range of human experience.
The study suggests that this assumption becomes even more problematic when machines are asked to use such information to predict how individuals will behave.
The researchers therefore urged caution in treating AI-generated participants as replacements for humans.
For now, digital twins may provide a useful supplementary tool for research, but the evidence suggests they cannot yet reliably reproduce the complexity, variation and unpredictability of human behaviour.

