Synthetic survey data can now get close to real results at the headline level but still falls short on the deeper analysis that real decisions depend on, according to a new STRAT7 white paper.
The clearest finding from the study which you can download here was that synthetic data correctly tracked year-on-year change in brand awareness just 19% of the time, missing real movement in 13 of 16 brands tested (Provider 1, awareness metrics, 2025 vs 2026).
The study is a re-run of STRAT7’s 2025 benchmark, in which synthetic data providers were asked to recreate real survey respondents that STRAT7 had deliberately held back. This year, STRAT7 repeated the test using the same survey framework. Around 3,000 real, nationally representative UK respondents answered questions about snacking, confectionery and crisp buying. We then compared two synthetic data providers against that real sample: one using a statistical machine-learning method, and one using an LLM-based method.
STRAT7 set out to answer a simple question: A year on, has synthetic data actually got better, and is it ready for the decisions insight teams are asked to support?
Headline findings
- Topline accuracy has improved. Both providers tracked real brand awareness to within about three percentage points for nearly every brand, with average errors of 2.0 and 2.8 points (all brands, boosted sample vs real holdout sample).
- Tracking change over time does not work. Synthetic data read year-on-year brand awareness movement correctly only 19% of the time (Provider 1, 16 brands, 2025 vs 2026). Real awareness rose by eight to twelve points for many brands over the year. The synthetic data stayed almost flat.
- The logic that holds a respondent together still breaks. In the pricing exercise, pure synthetic respondents contradicted their own price ordering 68% of the time, something the survey design makes impossible for real respondents (Provider 1, willingness-to-pay questions). Blending with real data lowered the figure to 32.8%, masking the problem rather than fixing it.
- Subgroup analysis is where errors compound. On a question about whether the World Cup makes people more likely to buy snacks, synthetic and real data agreed at the total level. Among women aged 35 to 44, the real figure was 31% and one provider read it as 14%, a 17-point miss that would have pointed a brand at the wrong target audience.
- Driver analysis pointed at the wrong levers. One provider showed almost no correlation with the real drivers of purchase, placing most weight on store availability, which ranked sixth in the real data.
Hasdeep Sethi, Data Science & AI Lead at STRAT7 and co-author of the study, said:
"The improvements are real, but they are concentrated where they are easiest to see and easiest to validate. Ask a synthetic respondent one question and the answer looks convincing. Ask them forty questions and the person holding those answers together starts to fall apart. That is the finding that matters, because subgroups, drivers, segments and change over time are where research actually earns its money. On current evidence, synthetic data is trustworthy for the topline, but not safe enough for the decisions beneath it."
Hasdeep Sethi, Data Science & AI Lead, STRAT7
The study concludes that synthetic augmentation has a narrow, legitimate role: small top-ups of genuinely hard-to-reach audiences, kept to around 10 to 20% of the sample for any subgroup work, used transparently, and validated against other methods. It also draws a firm line under one use case. On current evidence, synthetic data is not a substitute for tracking how real awareness, purchase, and attitudes change over time.
Methodology: STRAT7 surveyed around 3,000 nationally representative UK adults about snacking, confectionery, and crisp buying, then compared synthetic boosts from two providers, one statistical and one LLM-based, against a holdout of real respondents never shown to them. Test datasets were roughly half synthetic, a deliberately demanding benchmark.
The full white paper Synthetic data: is this as good as it gets? is available to download here.
About STRAT7
STRAT7 is a global tech-enabled strategy, insights, and analytics group that leverages AI-powered technology and human expertise to help businesses stay ahead of change. With a fully integrated ecosystem of 400+ experts, we equip brands with the actionable strategies and foresight needed to navigate disruption, accelerate customer-centric growth, and achieve meaningful commercial advantage.
Our agencies include: STRAT7 Advisory, STRAT7 Audiences, STRAT7 Bonamy Finch, STRAT7 Crowd DNA, STRAT7 Incite, STRAT7 Jigsaw, and STRAT7 Researchbods.
STRAT7 is headquartered in London with offices in Amsterdam, Leeds, New York, and Sydney.
For media enquiries, please contact hello@strat7.com.