Your synthetic audience will agree with you

For the past year, nearly every conversation I have had about synthetic audiences has been about accuracy. How close does the simulation get to what real people would have said? It is the right question to ask second, because asking it first imports an assumption that keeps turning out to be wrong, which is that we already know who those people are.
Before anyone runs a simulation, somebody writes a brief, and every brief contains an audience. She is thirty-four, with a household income, a commute, and a stated concern about sustainability. She is time-poor and value-driven. Most of the time that description came out of a planning meeting, a category convention, or whatever the last campaign in the category assumed. Then we hand it to a model and ask the model to be her.
That model is doing exactly what it was built to do. Language models produce the most plausible continuation available to them, so ask one to be a thirty-four year-old sustainability-minded shopper and it will return the most statistically ordinary version of that person it has ever seen. If the archetype in your brief is the conventional one –and conventional archetypes are conventional for a reason – the simulation will reproduce it faithfully and hand it back with a confidence score attached. The assumption goes in and comes back out wearing a number.
Take time-starved working mums. A planner’s shorthand would be easy to write: she is overloaded, practical, family-first, in the market for anything that saves time, reduces friction and helps her manage the home. A synthetic audience prompted from that brief would confirm it and return a woman defined by efficiency.
Structurally the brief is right. In our own audience data, this group is full-time working parents, 70% of them aged 25 to 44, half of them suburban, and more likely than average to be interested in parenting content, home organisation and family shows.
Their passions run somewhere else. They are 22% more likely than the general population to be into tattoos, 18% more likely to be into horoscopes and astrology, 16% more likely to be into love and sex advice, and 15% more likely to be into dating and relationships. That is a woman holding on to identity, pleasure and adult selfhood while she runs the household. And it’s exactly the contradiction a simulation smooths away, because the brief it inherits starts from the socially acceptable version of her.
The reason this worries me more than an accuracy problem is that an inaccurate simulation eventually announces itself. It misses in a way somebody notices, and the miss starts an argument. A simulation that agrees with the brief produces a campaign that performs exactly as poorly as the original assumption deserved, and nothing anywhere in that process raises a hand. The post-mortem blames the creative.
A campaign built on that archetype would have gone after relief and utility: save time, simplify dinner, reduce the mess, make the school week easier. None of that is wrong. The cost is that it stops there, and you end up with work that is perfectly relevant to her responsibilities and underpowered against her desires, which shows up as safe creative, flatter emotional resonance, and a narrower view of where the brand can meet her. Generic data hands you the plausible version of her, and only observed behaviour surfaces the distinctive one.
“Generic data hands you the plausible version of her, and only observed behaviour surfaces the distinctive one.”
None of this is an argument against simulation. The best public example I know is The Times, which built a synthetic panel with Electric Twin off its own database of 642,000 subscribers, then tested it by holding out 10% of that audience and asking the simulation to predict what the group would say. It came back at .918, or about 92% accurate, against the 93% their team cites for conventional research. That worked because The Times was not guessing who its readers were. It had a genuine relationship with an audience it already understood, and used simulation to extend what it already knew rather than invent what it did not.
So, the question I would put ahead of the accuracy question is whether the audience in the brief survives contact with observed behaviour. What people say about themselves tends to confirm the archetype for the same reason the model does, so this has to mean what they actually do, measured recently enough to still be true. Establish that first, and simulation becomes genuinely useful, because you finally know what it is standing in for.
Every synthetic audience is downstream of a decision somebody made about who the audience is. That decision is usually made in a room, in about ten minutes, by people working from memory. It deserves considerably more scrutiny than the model that inherits it.
Leslie Walsh is head of strategy at RYA
We hope you enjoyed this article.
Research Live is published by MRS.
The Market Research Society (MRS) exists to promote and protect the research sector, showcasing how research delivers impact for businesses and government.
Members of MRS enjoy many benefits including tailoured policy guidance, discounts on training and conferences, and access to member-only content.
For example, there's an archive of winning case studies from over a decade of MRS Awards.
Find out more about the benefits of joining MRS here.







0 Comments