
Synthetic respondents are becoming increasingly common in market research. They promise faster answers, lower data-collection costs, and the ability to simulate consumer responses without recruiting a new human sample for every question.
But “synthetic respondents” is not one technology.
They can be created from historical survey data, statistical distributions, consumer segments, large language models, retrieval systems, generative machine-learning models, or simulated populations of AI agents. Each approach makes different assumptions about how well existing information can represent consumers in a new research situation.
That distinction matters because synthetic respondents are most defensible when they stay close to patterns already observed in high-quality human data. They become harder to defend when research requires genuinely new evidence from the market.
Understanding how synthetic respondents are generated is therefore essential to understanding where they can – and cannot – be trusted.
Synthetic respondents can be useful for renovation testing, hypothesis generation, research instrument design, and simulation when they remain close to patterns already established with high-quality human data.
The risks increase when research requires new evidence: breakthrough innovation testing, new-category entry, current brand tracking, market sizing, willingness-to-pay for novel offers, poorly represented populations, or discovery of genuinely new consumer needs.
The fundamental limitation is simple: synthetic respondents generate new observations from existing data, learned patterns, or model assumptions. They do not independently go into the market and observe what consumers think today.
Generating thousands of synthetic responses can therefore create more simulated observations without creating the same amount of new empirical information as recruiting thousands of independent human respondents.
Synthetic respondents are artificially generated representations of consumers designed to produce survey answers, opinions, choices, or behaviors without collecting each response from a newly recruited human participant.
The underlying technology can vary substantially.
Some approaches use statistical models trained directly on previous human survey responses. Others create personas from customer segments and ask an LLM to answer as those personas. More complex systems combine LLMs with retrieval, behavioral rules, or agent-based simulations.
That means the question “Are synthetic respondents reliable?” does not have one universal answer.
Reliability depends on how the respondents were generated, the quality and representativeness of the underlying human evidence, and how far the new research question moves beyond what that evidence already contains.
Here are seven major approaches.
This approach trains a supervised statistical or machine-learning model on historical respondent-level data.
Demographics, attitudes, ratings, behaviors, and brand perceptions can become predictors, while an outcome such as purchase intent, consideration, satisfaction, or choice becomes the target. Regression, random forests, gradient boosting, neural networks, and related methods can then estimate expected outcomes for new synthetic profiles.
The limitation is that the model learns from historical observations rather than directly observing the current population.
If a new concept, audience, or market condition falls outside the range represented in the training data, the model must extrapolate. That increases uncertainty.
Sampling problems, measurement errors, or selection bias in the original data can also propagate into the predictions.
Rather than predicting a particular outcome, statistical distribution modeling attempts to reproduce the statistical properties of an existing human dataset.
Simple approaches recreate individual distributions. More sophisticated methods use conditional distributions, copulas, Bayesian networks, or related techniques to preserve relationships among variables.
The goal is statistical similarity to the source data, not simulation of individual consumer reasoning.
This distinction is important.
The resulting synthetic dataset may resemble the original sample statistically, but it still inherits the boundaries of that sample. Relationships that were missing, poorly measured, or underrepresented in the original human data cannot reliably appear simply because additional synthetic records have been generated.
Synthetic similarity to a sample is not necessarily the same as representation of the underlying consumer population.
Segment-based systems begin with consumer groups identified through surveys, interviews, CRM records, behavioral data, or segmentation studies.
Each segment is described through characteristics such as attitudes, needs, demographics, behaviors, and preferences. An LLM or another generative system is then instructed to respond as a member of that segment.
This is intuitive for market researchers because personas and segments are already familiar tools.
The problem is that a segment is an aggregate description.
When an AI repeatedly generates people from that description, it can reproduce the characteristics that define the segment while compressing the variation that exists among real people within it.
Synthetic members of a segment may therefore appear more internally consistent than actual consumers.
Recent research into persona-conditioned LLM survey respondents has also found that adding demographic personas does not necessarily improve alignment with human responses and can disproportionately distort results for some underrepresented groups.
A straightforward approach is to provide a large language model with a respondent profile containing information such as age, income, behavior, attitudes, or brand preferences and ask it to complete a survey as that person.
Repeating this process across many profiles creates a synthetic panel.
The attraction is obvious: LLMs can generate large volumes of plausible, natural-language responses very quickly.
But plausible language and population-level accuracy are different objectives.
LLMs generate responses based on patterns learned during model training and information supplied in the prompt. They can underrepresent human variation, produce responses that are more internally consistent than real survey data, and reproduce stereotypes embedded in either the model or the persona definition.
Research comparing LLM-generated survey data with human surveys has found that synthetic responses can sometimes reproduce aggregate averages while still showing lower variation, different statistical relationships, and sensitivity to prompt wording. (cambridge.org)
Retrieval-augmented generation, or RAG, adds existing human evidence to an LLM before it produces a response.
That evidence might include survey verbatims, interview transcripts, CRM records, product reviews, customer feedback, or previous research studies.
When a question is asked, relevant evidence is retrieved and provided to the model as context.
This can make responses more grounded than relying on the LLM's pretrained knowledge alone.
But the system remains constrained by the evidence available for retrieval.
If the repository contains rich information about existing products but nothing about a genuinely new concept, behavior, or audience, the model eventually has to infer beyond the retrieved evidence.
More detailed grounding can improve alignment with human data, but recent participant-matched persona research suggests that richer conditioning does not eliminate the gap between simulated responses and individual human differences. (sciencedirect.com)
More advanced synthetic-data systems use generative models such as GANs, variational autoencoders, diffusion models, or specialized tabular-data generators.
Rather than modeling a small number of relationships individually, these approaches attempt to learn the joint statistical structure across many variables in an existing respondent dataset.
Once trained, they can generate new combinations of attributes and responses that statistically resemble the source data.
These methods can capture nonlinear relationships that simpler models may miss.
However, they remain models of existing evidence.
Patterns need to be present – and sufficiently represented – in the training data for the model to learn them reliably. Biases, missing relationships, and weak representation of minority groups can therefore carry into the generated population.
Greater generative sophistication does not remove dependence on the original human data.
Agent-based systems go beyond generating isolated survey records.
They create multiple simulated consumers with profiles, preferences, memory, goals, and behavioral rules. Modern systems often use LLMs as reasoning engines while adding demographic information, retrieval systems, behavioral models, or interaction rules.
Agents can take surveys, encounter products, make choices, communicate with other agents, and participate in simulated market environments.
This creates powerful opportunities for experimentation and simulation.
But the apparent size of an agent population can be misleading.
Thousands of agents may still be generated from the same underlying model architecture, assumptions, prompts, and source data. Their errors can therefore be correlated rather than statistically independent in the same way as separately recruited human respondents.
A larger simulated population does not automatically create proportionally more independent information.
The limitations above do not make synthetic respondents useless.
They can be valuable when the objective is simulation or exploration rather than collecting genuinely new market evidence.
Potential applications include:
The common characteristic is that these applications remain relatively close to information already available.
Synthetic respondents are generally on firmer ground when asked to recombine, simulate, or explore known patterns than when asked to reveal something the underlying evidence has never observed.
The risk increases substantially when a research question requires genuinely new information about the current market.
Brand tracking is intended to detect changes in awareness, consideration, perceptions, usage, and other measures over time.
Synthetic respondents generated from historical information are conditioned on previous patterns. They cannot independently observe a genuine change occurring in the market today.
A genuinely new product, benefit, or behavior may sit outside the distribution represented in the training data.
The further the concept moves from previously observed evidence, the more the model is extrapolating rather than measuring.
Consumer needs, competitors, usage occasions, price sensitivity, and decision criteria can change significantly between categories.
Relationships learned in one market should not automatically be assumed to transfer to another.
These estimates depend on defensible measurement of prevalence in a target population.
The frequency of an outcome among generated records reflects the synthetic model and its assumptions. It does not represent additional independent draws from that population.
Price response depends on trade-offs consumers make under current product, competitive, and economic conditions.
If those combinations have not been observed in the source data, the model is estimating outside the evidence from which the demand relationship was learned.
A model cannot reliably reconstruct heterogeneity that is missing from its source data.
Rare or underrepresented groups are particularly vulnerable to being pulled toward majority patterns.
Synthetic systems generate from learned distributions, retrieved evidence, or predefined structures.
They are therefore much better at recombining known patterns than discovering motivations, language, or behaviors absent from the information used to create them.
Generating 10,000 synthetic respondents from evidence collected from 500 human respondents does not create 10,000 independent empirical observations.
Recent statistical work on LLM-simulated surveys similarly warns that increasing the number of simulated responses can create confidence intervals that appear increasingly precise without resolving misalignment between the model and the human population. (papers.ssrn.com)
The number of synthetic records should therefore not automatically be treated as the inferential sample size.
The fundamental distinction is not that human research is perfect and synthetic research is flawed.
Human research has its own sampling, measurement, response-quality, and methodological challenges.
The distinction is where new empirical information enters the system.
A newly recruited human respondent provides another observation from the population being studied. A synthetic respondent generates an observation from patterns learned from existing information, model parameters, retrieved evidence, or assumptions.
This is why synthetic respondents can reproduce known patterns surprisingly well while still being unreliable for discovering distribution shifts, unexpected behaviors, or new consumer needs.
The closer the research question is to established evidence, the stronger the case for simulation can become.
The more the business decision depends on learning something genuinely new about the market, the stronger the case for asking real people.
Compeers AI uses AI to accelerate market research without replacing the source of consumer evidence with synthetic respondents.
Qualitative, quantitative, mixed-method, and advanced analytics studies are conducted with real human participants. AI assists with work across planning, fieldwork, analysis, visualization, and reporting, while researchers remain involved in methodological decisions, review, interpretation, and recommendations.
This distinction reflects a broader principle: AI can dramatically improve how efficiently research is conducted without requiring simulated consumers to stand in for new human evidence.
Synthetic respondents should not be treated as one technology or evaluated with one universal claim about reliability.
There are multiple ways to generate them, ranging from traditional predictive and statistical models to LLM personas, RAG systems, generative models, and agent-based populations.
Their usefulness depends on the relationship between the question being asked and the evidence used to create the synthetic population.
When the task involves simulation, hypothesis generation, instrument design, or exploration close to well-established human data, synthetic respondents can be useful.
When the objective is to discover something genuinely new about consumers or measure how the market is changing now, the limitations become much more important.
The question is therefore not simply whether synthetic respondents “work.”
It is whether the available evidence and generation method are appropriate for the decision the research is expected to support.
What are synthetic respondents?
Synthetic respondents are artificially generated representations of consumers designed to produce survey responses, opinions, choices, or behaviors without collecting each response from a newly recruited human participant. They can be created using statistical models, historical data, consumer segments, LLMs, retrieval systems, generative models, or agent-based simulations.
How are synthetic respondents generated?
Common approaches include predictive modeling from historical survey data, statistical distribution modeling, segment-based personas, LLM-generated respondents, retrieval-augmented generation, generative machine-learning models, and agent-based synthetic populations.
Are synthetic respondents reliable?
Their reliability depends on the generation method, quality and representativeness of the underlying evidence, and the research question. They are generally more defensible when operating close to patterns already observed in human data and less reliable when extrapolating to new concepts, audiences, categories, or market conditions.
Can synthetic respondents replace human respondents?
Synthetic respondents can supplement human research or support simulation and exploration, but they do not create independent empirical observations of the current consumer population in the same way as newly recruited human respondents.
When should synthetic respondents be used?
They can be useful for hypothesis generation, research instrument design, scenario simulation, renovation testing, and other applications grounded in well-established human evidence. Their use requires greater caution when the research objective is discovery, measurement of current market change, or statistical inference about a population.
Does Compeers AI use synthetic respondents?
Compeers AI conducts custom market research with real human respondents. AI is used to accelerate and support the research workflow – including planning, analysis, visualization, and reporting – rather than replacing human consumer evidence with synthetic respondents.
Adapted from an article by Vijay Rajan, Founder of Compeers AI