Synthetic data has gone from research curiosity to boardroom line item in under two years. AI-generated "personas" now sit in agency decks alongside real fieldwork, promising to answer any question about any audience in seconds, for a fraction of the cost of recruiting actual people. For an industry under constant pressure to move faster and spend less, the appeal is obvious. The problem is what synthetic data quietly asks you to give up in return.
The synthetic data gold rush
By most industry estimates, a majority of research and insight teams have now experimented with AI-generated respondents in some part of their process, up from a niche practice just two years ago. The pitch is straightforward: large language models trained on vast amounts of survey, review and social data can generate plausible consumer responses instantly, with no recruitment lag, no scheduling, no incentive payments, and near-zero marginal cost per additional "respondent."
For certain jobs, that trade genuinely makes sense. Synthetic panels are useful for early-stage ideation, stress-testing question wording before fielding, and filling small gaps in hard-to-reach segments. Used well, as a supplement rather than a substitute, they can make research faster without making it worse.
What synthetic data can and can't do
The trouble starts when synthetic data is asked to do more than that. A language model has no stake in the outcome it's describing. It has never actually bought the product, never actually felt a price increase land on a tight monthly budget, never actually had the argument in the car park about which insurer to renew with. What it produces is, at best, a very well-read guess – a statistically plausible answer built from patterns in text other people have written, not a report of what one specific person actually thinks right now.
That distinction sounds academic until you're the brand making a launch decision on the back of it.
"Synthetic data can tell you what a plausible consumer might say. It can't tell you what an actual person, with an actual life, actually feels – and in research, that difference is the entire point."
The empathy gap
Real consumers contradict themselves. They hesitate before answering an awkward question, change their mind halfway through a sentence, and say things that don't fit the pattern the rest of the sample suggests. That mess is not noise to be cleaned up – it is very often where the actual insight lives. A model trained to predict the statistically likely response will, by design, regress toward the average and smooth away exactly the outliers that a good researcher is trained to chase.
Emotionally loaded categories make the gap widest: financial stress, health decisions, bereavement, major life changes. These are precisely the moments where lived experience cannot be convincingly simulated, and precisely the moments where getting the insight wrong carries the highest cost.
Where synthetic data quietly fails
Price sensitivity testing is a good example. A synthetic respondent doesn't feel a real household budget constrict when a price rises – it estimates how a person like that, in aggregate, tends to describe feeling it. Regulated categories carry a sharper version of the same risk: a wrong "insight" in financial services or healthcare research isn't just embarrassing, it can have real consequences for real customers downstream.
None of this makes synthetic data worthless. It makes it a tool with a job description – and that job description does not currently include "final word on what real customers will do."
The hybrid future – and why verification matters more, not less
The winning model isn't synthetic versus human. It's knowing, with certainty, which is which, and never letting the two blend without disclosure. As AI generates a growing share of the responses flowing through research pipelines, verifying that a panel described as "real" is actually real becomes the load-bearing wall of the entire industry's credibility.
That's why voice verification matters more now than it did five years ago, not less. Confirming a live human being before a single question is answered turns "are these real people?" from a matter of trust into a matter of record – exactly the guarantee that gets harder to give, and more valuable to have, the more synthetic data floods the market around it.