Show a trained researcher a synthetic dataset today and, at the headline level, they would struggle to pick it out from a real one. That is the good news, and it is also why the rest of this is worth reading.
We ran our synthetic data study for a second time this year, using the same survey we fielded in 2025 so we could compare like with like.
Synthetic data has not improved evenly. It performs well on straightforward headline measures but is less reliable as the analysis becomes more detailed and the resulting decisions more important. That is what CMOs need to understand.
In both tests, roughly half the dataset was synthetic, well above the 10 to 20% share we would recommend in practice. Read the findings in that light.
Why synthetic data is attracting attention
Research teams are dealing with tighter budgets, concerns about survey fraud and falling response rates, particularly among younger audiences.
Synthetic data promises a compelling alternative: additional respondents generated in minutes rather than weeks, often at a flat fee and without involving identifiable individuals.
The attraction is obvious. Faster and cheaper research could help teams answer more questions and reach audiences that are difficult or expensive to recruit. But the important issue is not simply whether synthetic data works. It is whether it is reliable enough for the decision being made.
What has improved
At the headline level, the progress is significant.
Both providers in our study tracked real brand awareness within about three percentage points on nearly every brand. Two visible problems from last year have also largely disappeared: synthetic respondents no longer bunch heavily around round prices or cluster around the middle of agreement scales.
Their answers have become more consistent too. When we checked whether a stated attitude matched an actual behaviour, synthetic respondents contradicted themselves at almost exactly the human rate.
For simple topline measures, the results can therefore look remarkably convincing.
Where the risks begin
The problems return when the analysis becomes more specific.
We asked whether the 2026 World Cup made people more likely to buy snacks. At total level, the real and synthetic results were within a couple of points. Among women aged 35 to 44, however, the real figure was 31%, while one provider reported 14%.
That is not simply a statistical difference. It could lead a marketing team to prioritise the wrong audience.
The same risk appears in more advanced analysis. In the real data, promotions and advertising were the strongest influences on purchase frequency. One provider identified those correctly; the other made store availability the main driver, even though it ranked sixth in the real data.
For a CMO, that could mean putting budget behind the wrong marketing lever.
Segmentation can also look reassuring while still being misleading. Segment sizes were broadly accurate, but the synthetic data overstated willingness to pay for sustainable products by more than 20 percentage points within one segment. A business could therefore identify the right-sized audience while misunderstanding what motivates it.
It cannot yet track change
The clearest warning came from tracking.
Across all measures, the provider we could test over time called the direction of change correctly 47% of the time. On brand awareness, it was right only 19% of the time, missing real movement in 13 of 16 brands.
For many brands, real awareness rose by eight to 12 points while the synthetic figures remained almost flat.
If a tracker says awareness has not moved when it actually has, decisions about campaign performance and future investment are being made from the wrong starting point. The same concern applies to market mix modelling, where small differences can change conclusions about which channels are working.
What should CMOs do?
The opportunity is real, but synthetic data should be used selectively.
- Use it to extend real data, not replace it. It can work as a small top-up for genuinely hard-to-reach audiences. We recommend keeping it to 10 to 20% of the total sample.
- Match the method to the decision. Simple topline measures may be suitable. We would not currently use it for tracking, market mix modelling or to build new segmentation and demand-space frameworks.
- Build in checks. Tell stakeholders where synthetic data has been used, compare providers carefully and validate anything decision-critical against another source.
Asked on our podcast for the one thing I would tell a CMO, I compared synthetic data with a sketch rather than a photograph. It can capture the broad picture, but not always the detail, movement and texture needed for foundational decisions.
That is the most useful way to think about it. Synthetic data may be ready to support some research decisions. It is not yet ready to carry the most expensive ones.
Read the full study for the five-pillar testing framework, the pricing and cost-modelling work, and the year-on-year results in detail.
Whitepaper
Synthetic data 2.0:
Is this as good as it gets?
One year on – what’s improved, what still falls short, and what it means for insight teams.
About the author
- Hasdeep Sethi
- Group AI Lead & Data Science Director
Hasdeep is a data science and AI leader with 10+ years’ experience building, delivering and setting strategic direction on data science, machine learning and AI projects. He helps lead STRAT7’s AI offer through Nucleus and the AI Innovation Lab, and speaks regularly at industry events including MRS, GRBN, IIEX and ASC. His particular interest is how AI and synthetic data are reshaping insight and research.