
Synthetic respondents are often promoted as a faster, lower-cost way to generate consumer insights. The outputs look clean, coherent, and presentation-ready, making them attractive for concept testing, messaging evaluation, and early-stage research.
But there is a fundamental methodological question that deserves far more attention:
Can synthetic respondents produce stable, dependable measurements?
Unlike traditional research instruments, LLM-based synthetic respondents may change their outputs when underlying model conditions change. In research, the quality of the measurement instrument matters as much as the quality of the analysis. If the instrument itself is unstable, polished outputs do not necessarily translate into reliable insights.
Every research project depends on one assumption:
The measurement instrument should produce results that are reasonably stable when nothing important has changed.
Researchers routinely evaluate whether a measure is:
If small procedural changes produce substantially different results, confidence in the findings quickly begins to erode.
That principle applies whether the research involves surveys, interviews, experiments, or AI-assisted methods.
Many synthetic respondent platforms rely on large language models (LLMs) or other machine learning models to generate responses.
These systems can be sensitive to factors that have little to do with actual consumer attitudes, including:
None of these factors necessarily reflect a meaningful change in consumer opinion.
Yet they can produce meaningfully different outputs.
When that happens, researchers may no longer be measuring a stable underlying construct.
Instead, they may be measuring a moving target.
Imagine interviewing the same research participant twice.
The question changes only slightly.
If the respondent's answer changes dramatically, researchers would immediately investigate:
The same level of scrutiny should apply to synthetic respondents.
Instead, instability is often treated as an expected characteristic of the technology.
That creates an important challenge.
When synthetic outputs change, what exactly are they representing?
If those questions cannot be answered confidently, the results may still be interesting.
But they become much harder to treat as dependable research measurements.
Neither approach should be judged by how polished the outputs appear. The real question is whether the measurement remains stable enough to support business decisions.
When evaluating AI-powered research tools, don't focus only on speed or cost.
Ask questions about measurement quality.
A useful evaluation framework includes:
These questions reveal far more about research quality than polished demonstrations.
Before adopting synthetic respondents in your research workflow, ask your vendor:
The answers to these questions can tell you far more about the quality of a platform than its speed or user interface.
Synthetic respondents are increasingly used to inform:
These decisions often involve significant investment.
If synthetic respondents change their "opinions" because the vendor modified a prompt stack or upgraded the underlying model, organizations may mistake technical variability for genuine market insight.
That creates unnecessary risk.
Speed is valuable.
But methodological robustness is even more valuable when the decisions affect brands, products, and markets.
Compeers AI does not advocate using AI personas as substitutes for real respondents because LLMs are optimized to generate plausible, coherent language, not to faithfully represent the inconsistency, contradiction, and edge-case behavior that real consumer insight depends on.
Their outputs can also shift with prompts, model settings, and version changes, which makes them an unstable measurement instrument for decisions about brands, products, and markets.
What is stable measurement in market research?
Stable measurement means that a research instrument produces reasonably consistent results when the underlying conditions have not materially changed. This is a fundamental requirement for reliable research.
Why can synthetic respondents produce different answers?
Their outputs may change because of prompt wording, system prompts, model updates, temperature settings, or other configuration changes that are unrelated to actual consumer preferences.
Are synthetic respondents reliable enough for strategic decisions?
They can be useful for exploration and idea generation, but organizations should understand how vendors test for reliability, measurement stability, and reproducibility before using them to support major business decisions.
What questions should you ask a synthetic respondent vendor?
Ask how they evaluate measurement stability, monitor output variability, document model changes, test prompt sensitivity, and ensure results remain consistent over time.