"Synthetic" is a word that makes real user researchers uncomfortable.
It sounds artificial. Manufactured. Like you're replacing something genuine with something fake and calling it progress. When a researcher who's spent years sitting across from real people - reading their body language, catching the pause before they answer, noticing the workaround they'd never mention unprompted - hears that an AI can now simulate all of that, the instinct is to reject it outright.
I understand that instinct. I've felt it myself.
But here's what I've come to believe after watching this space evolve, working with 3,000+ designers, and building research capabilities at Xperience Wave and Konfom: the question isn't whether synthetic users can replace real participants. They can't - not today, probably not fully ever. The question is whether they can provide 20%, 30%, 40% of the signal at a fraction of the cost and time - and whether that percentage keeps growing. Because the data says yes. And the designers and researchers who figure out how to use that signal well, alongside real research, will outperform both the dismissers and the over-reliers.
What Synthetic Users Actually Are
Let's strip away the mystique.
Synthetic user tools use large language models to generate responses that simulate how a hypothetical user with specified characteristics might react to interview questions, product concepts, survey prompts, or design stimuli [1]. You define a persona - demographics, behaviours, attitudes, context - and the model generates responses as if it were that person.
That's it. It's not a holographic replica. It's not a digital twin that feels emotions. It's a statistical model making predictions about how someone with those characteristics would likely respond, based on patterns in the data it was trained on.
The sophistication varies enormously. At one end, you're prompting ChatGPT with "pretend to be a 35-year-old Indian banking customer." The output is generic, shallow, and about as useful as asking your colleague to roleplay. At the other end, you have custom models fine-tuned on actual interview transcripts, support tickets, behavioural data, and survey responses from your specific user base - models that aren't guessing what your users might say but pattern-matching against what they actually have said.
That distinction - generic LLM versus custom-trained model - is the single most important thing to understand about synthetic users. Everything that follows depends on it.
What They Can Actually Do (The Data Is Clearer Than You'd Expect)
The industry's most respected research methodology survey, the GRIT report, found that research teams using synthetic data report high satisfaction with results. Specifically: in well-documented categories - established user behaviours, known segments, validated UX heuristics - synthetic approaches are proving nearly as accurate as real-participant methods, and dramatically faster [2].
That's not a fringe finding. That's the industry's own assessment.
Here's where synthetic users are delivering genuine, documented value today:
Hypothesis generation and concept screening. Before you invest weeks in recruitment and formal studies, synthetic users can help you explore the problem space, screen multiple directions, and identify which concepts are worth validating with real participants. The consensus across every serious source - NNGroup, ACM, User Vision, TheySaid - converges on this: synthetic users are strongest when the cost of being wrong is low and the value of speed is high [3].
Rapid iteration during design.When you're in an active design sprint and need directional feedback on three layout alternatives before tomorrow's review, synthetic users can provide structured input in hours rather than the days or weeks that participant recruitment requires. The signal isn't as deep as real testing, but it's infinitely better than no signal - which is what most teams have at this stage.
Scaling qualitative research. A 2024 study by Kapania et al. had 19 UX researchers recreate one of their recent projects using GPT-4-Turbo instead of real participants. The researchers were initially sceptical. They were then surprised to see similar narratives emerge in the LLM-generated data [4]. The themes were recognisable. The directional insights aligned. The surface-level patterns held.
Augmenting small sample sizes.When budget allows five real interviews but you need broader coverage, synthetic users can extend the exploration - testing additional persona variations, edge-case scenarios, or market segments that your five participants don't represent. Not replacing the five. Extending beyond them.
Cost efficiency that changes the calculus.Synthetic user platforms like Synthetic Users charge $2–27 per AI interview, with an additional ~$5 for RAG grounding against your own customer data [5]. Compare that to the cost of recruiting, scheduling, incentivising, and conducting interviews with real participants - often hundreds of dollars per session. This doesn't make synthetic users "better." It makes them available in situations where real research was previously impossible due to budget.
The Five Things They Cannot Do (Today)
Here's where I want to be equally direct - because the limitations are real, documented, and dangerous to ignore.
1. They validate too much.
This is the single most dangerous characteristic of synthetic users, and it's insufficiently discussed.
AI personas have a documented tendency to validate ideas that real users would reject [3]. LLMs are trained to be helpful, cooperative, and agreeable. When you ask a synthetic user "would you use this feature?" the model is statistically biased toward "yes" - because "yes" is more common in the training data than honest rejection.
Real users say no. Real users say "I don't understand why this exists." Real users ignore your feature entirely because it doesn't match their mental model. Synthetic users, by default, give you the answer you want to hear. If you're using them for validation rather than exploration, you're building a confirmation machine, not a research tool.
2. They can't simulate embodied, contextual experience.
A user completing a checkout flow while distracted, on a phone, during a commute, with a slow connection, after a frustrating experience with a different app earlier that day - that's a real usage context [1]. Synthetic users operate in clean, context-free environments. They process the interface as it's presented, not as it's experienced.
Real UX happens in messy conditions - interrupted attention, emotional carry-over, environmental constraints, physical limitations. Synthetic users can't simulate the frustration of a user who's already had three failed login attempts that morning. They can't replicate the cognitive load of someone who's navigating your enterprise dashboard while simultaneously on a call with their manager. Context shapes behaviour. Synthetic users don't have context.
3. They reflect training data biases, not real user populations.
This one matters especially for our audience.
CleverX's 2026 analysis was direct: "Underrepresented groups, non-Western markets, elderly users, users with disabilities, users with low digital literacy, and users from economic contexts that generate less online text are all represented less accurately in synthetic simulations" [1].
For designers building products for Indian users - particularly for Tier 2/3 cities, for users who navigate between English and regional languages, for demographics that don't generate much English-language internet content - generic synthetic users are producing statistically plausible fictions that have no reliable grounding in the actual population the product is designed for. The model doesn't know how a first-time UPI user in Indore thinks about digital payments. It knows how English-language internet text describes digital payment behaviour. Those are not the same thing.
This limitation is not permanent. It's solvable with custom training data from your actual user base. But it's the default state of every generic synthetic user tool today.
4. They don't model social dynamics.
Real UX is rarely one person alone with a screen. People influence each other's choices - social proof, peer pressure, family decision-making, group dynamics. The ACM's Interactions journal flagged this directly: "Simulated individuals typically don't model multi-user interactions such as collaborative tools, family device-sharing, or viral effects in social apps" [6].
If your product involves shared decision-making (a family choosing an insurance plan), collaborative use (a team using a project management tool), or social influence (a social commerce platform) - synthetic users will miss the interpersonal dynamics that drive real behaviour.
5. They follow optimal paths, not real paths.
This connects to what we found in the usability evaluation blog- AI navigates interfaces too successfully. Synthetic users find the correct button. They interpret labels correctly. They don't bring the wrong mental model to a task. Their very competence makes them blind to the confusion, the misinterpretation, and the creative misuse that real users consistently demonstrate.
The most valuable insights in user research often come from the unexpected - the user who does something you never anticipated, the workaround that reveals a latent need, the error that exposes a flawed assumption. Synthetic users stick to the most logical paths. Real insights live on the illogical ones.
The Custom Data Argument (Why Generic vs Trained Changes Everything)
Here's where the conversation shifts from "synthetic users don't work" to "synthetic users don't work like that."
Baymard Institute's heuristic evaluation system achieved 95% accuracy - but only through a RAG architecture trained on 170,000+ real, expert-curated UX examples accumulated over 15 years [7]. Generic tools hit 50-75%. The difference isn't the model. It's the data underneath it.
The same principle applies to synthetic research participants. A generic LLM prompted with a persona description is guessing based on internet text. A model fine-tuned on your actual customer interviews, support transcripts, behavioural data, and survey responses is pattern-matching against real signals from real people.
Companies in the research space - and there are a growing number of them - are already building on this principle. RAG-grounded personas, where the model's responses are constrained by and sourced from your own customer data, produce meaningfully different output than generic prompting. The responses are more specific. The objections are more realistic. The edge cases are more grounded.
Is it 100% equivalent to talking to a real person? No. Is it meaningfully better than generic simulation? The data says yes. Is it improving with each generation of model architecture and each increment of training data? Demonstrably. This is exactly the trajectory that matters: not whether synthetic users are perfect today, but whether 30% confidence at near-zero cost is better than 0% confidence because you couldn't afford to recruit. And whether that 30% is becoming 40%, then 50%, then higher as your proprietary data grows and your models improve.
Where Synthetic Users Fit in the Research Process
Let me map this concretely. Here's a typical research workflow and where synthetic users add value versus where they don't:
Market segmentation and hypothesis formation.You're identifying which user segments to study, what behavioural patterns to expect, what attitudes might exist. Synthetic users can help here - generating preliminary hypotheses, stress-testing segment assumptions, exploring adjacent segments you hadn't considered. Confidence level: moderate. Value: high, because the alternative at this stage is usually assumption.
Participant identification and recruitment.This is where most research budget goes. Synthetic users don't help with recruitment - but they can reduce how many real participants you need by handling the exploratory phase that would otherwise consume your first five interviews. Instead of using expensive real-participant sessions for hypothesis generation, use synthetic users for that, and reserve real participants for hypothesis validation.
Research execution - qualitative. Sitting with people, watching them, asking structured and unstructured questions, removing bias, collecting supportive data. This is where synthetic users are weakest and real research is irreplaceable. The nuance, the hesitation, the body language, the unexpected revelation - none of this transfers to simulation. Use real participants here. No substitution.
Research execution - quantitative. Surveys, task-based testing, metric collection. In well-documented categories with established user behaviours, synthetic users can approximate real responses with useful accuracy [2]. For novel categories, new markets, or populations underrepresented in training data - real participants remain essential.
Data analysis and synthesis. This is where synthetic users become interesting again - not as data sources, but as analysis tools. Feed your real research data into an AI system and use it to help identify patterns, generate preliminary themes, and surface connections you might have missed. The insight still needs human judgment. But the pattern-finding speed is genuinely useful.
Insight validation.Before you commit to a direction based on your synthesis, synthetic users can serve as a pressure test - "given what we know about this segment, does this recommendation hold up?" Not as definitive validation, but as a directional check that catches obvious blind spots before you present to stakeholders.
What the Future Actually Holds
The trajectory is clear, even if the timeline isn't.
Custom model training is becoming accessible.Fine-tuning on proprietary data is getting cheaper and easier. Within two to three years, mid-size product companies will be able to build synthetic user models trained on their own customer data - not just persona prompts, but grounded behavioural simulations. This narrows the gap between "generic guessing" and "informed prediction" significantly.
Multi-agent architectures are reducing flatness. Current synthetic users often feel one-dimensional - they give a single voice, a single perspective. Multi-agent systems, where several models collaborate on each response representing different cognitive styles and behavioural tendencies, are producing richer, more varied output. This is already shipping in platforms like Synthetic Users and will become standard.
Behavioural modelling is the real frontier.The future isn't about making synthetic users look more real (holograms, AR avatars). It's about making their behavioural predictions more accurate - better models of decision-making, emotional response, and contextual adaptation. This is hard. But it's where the research investment is going, and progress is measurable year over year.
The 30% will become 50%, then 60%. Not linearly, not uniformly across all use cases - but directionally. Synthetic users grounded in real customer data, operating within well-documented categories, with multi-agent architectures, will deliver confidence levels that make them indispensable as a research augmentation tool. Not a replacement for human research. An augmentation that makes human research more focused, more efficient, and more impactful.
The competitive risk is real. Teams that dismiss synthetic users entirely will find themselves out-iterated by teams that use them for rapid exploration and reserve human research for high-stakes validation. The speed advantage compounds: more iterations, more hypotheses tested, more dead ends identified early, more confident direction when they finally invest in real participant research.
The Honest Conclusion
Synthetic users are not the future of UX research. Real users are.
But synthetic users are a growing, improving, and increasingly essential part of the research toolkit. The researchers and designers who learn to use them well - who understand where they add genuine signal, where they mislead, and how custom data changes the equation - will practise better research than those who reject them and better research than those who over-rely on them.
The skill isn't learning the tool. The skill is judgment: knowing when 30% confidence at zero cost and instant speed is exactly what you need, and knowing when nothing less than sitting across from a real human being will give you the insight that matters.
That judgment - the ability to design research approaches that blend AI capability with human depth, that match method to question, that know when to explore synthetically and when to validate with real people - is one of the core design skills that AI can't automate. It's research leadership. And it's what separates designers who use tools from designers who shape outcomes.
If you're building this judgment - learning to integrate AI into your research practice without losing what makes research valuable in the first place - that's what we develop in our mentorship programs. And it's what we're building into Konfom - AI-assisted research grounded in real data, with a confidence engine that tells you how solid your direction is before you commit. Not synthetic users replacing real research. AI making real research sharper. Book a strategy call if you want to talk about where your research practice stands and where it needs to go.
Sources & References
- [1] CleverX. (2026). "Synthetic Users for Research: What They Are and Where They Fall Short." March 2026. https://cleverx.com/blog/synthetic-users-for-research-what-they-are-and-where-they-fall-short/
- [2] User Vision. (2026). "Synthetic Users vs Digital Clones: What Every UX Researcher Needs to Know." April 2026. Referencing GRIT report findings on synthetic data satisfaction. https://uservision.co.uk/thoughts/synthetic-users-and-digital-clones-a-ux-researcher-s-honest-take
- [3] TheySaid. (2026). "What is Synthetic User Testing? Benefits & Limitations." April 2026. https://www.theysaid.io/blog/synthetic-user-testing-guide
- [4] MeasuringU. (2026). "A Review of Experiments with Synthetic Users." Referencing Kapania et al. (2025) study with 19 UX researchers recreating projects with GPT-4-Turbo. https://measuringu.com/review-of-experiments-with-synthetic-users/
- [5] AI-CMO / Articos. (2026). "Synthetic Users Review 2026." Pricing analysis: $2-27/interview + ~$5 RAG grounding. https://ai-cmo.net/tools/synthetic-users
- [6] Russell, D.M. (2026). "The Challenges of Synthetic Users in UX Research." ACM Interactions, January-February 2026. https://interactions.acm.org/archive/view/january-february-2026/the-challenges-of-synthetic-users-in-ux-research
- [7] Baymard Institute. (2026). "AI Heuristic UX Evaluations with a 95% Accuracy Rate." May 2026. https://baymard.com/blog/ai-heuristic-evaluations
- [8] Lin, V. & D'Amour, A. (2026). "The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study." Carnegie Mellon / Google. https://arxiv.org/pdf/2605.20767
Further Reading on Xperience Wave
- AI Will Change Everything About Design. Except the Part That Actually Matters.
- How to Evaluate Usability - Traditionally and Using AI
- 5 Things Senior Designers Should Be Doing With AI - None of Them Involve Figma Plugins
- How to Build Your Personal AI Workflow as a Designer
About the Author
Shaik Murad is Co-founder and Head of Product & Design at Xperience Wave, a UX design studio based in Bangalore, working across designer mentorship, UX services for businesses, and a design community of 1,000+ designers.
- Murad, Co-founder & Head of Product & Design, Xperience Wave