Every new research method goes through the same three phases. First it is dismissed. Then it is adopted faster than anyone expected. Then, at some point, someone figures out how to measure it, and that is the moment it stops being a trend and becomes infrastructure.
Synthetic user research is in the third phase right now. That is the story of 2026, and I do not think enough people have noticed.
The published validation work has converged on something remarkable: across academic studies and commercial pilots, synthetic responses correlate with real respondent data at roughly 80 to 95 percent on directional questions. Bain’s research on synthetic customers found digital twins replicating around 90 percent of the outcomes from a full human conjoint study. These are not marketing numbers. They are measured results against human baselines.
But the number is not the interesting part. The interesting part is that we now know precisely what makes it move.
We know the conditions for accuracy, and they are all buildable
This is the shift that should excite anyone in this space. Two years ago, whether a simulation was accurate felt like a black box. Today the drivers are identified, published, and engineerable.
The validation literature is consistent on when accuracy holds. It holds when the persona is calibrated on real prior data from the same audience, when the question rewards general reasoning rather than unique lived experience, and when the platform exposes uncertainty instead of presenting every output as confident.
Read that again as an engineering spec rather than a caveat. Every one of those three conditions is something a platform can be built to satisfy.
Calibration on real audience data is a data pipeline problem. Solvable, and the solution compounds.
Matching question types to method strengths is a product design problem. You route the questions simulation is strong on to simulation and flag the rest.
Exposing uncertainty is a reporting layer problem. Harder than it sounds, but very much tractable.
None of these require a breakthrough. They require building. That is an enormously bullish position for a category to be in, because it means progress is a function of execution rather than of waiting for the next model generation.
Grounding is the lever, and it works
If you want to know where the accuracy gains are coming from, it is not model choice. Bain reached this conclusion explicitly in their own research: proprietary data mattered more than which model you picked. Our experience says the same thing, emphatically.
Three grounding channels have moved the needle most for us:
First-party customer material. Sales call transcripts, discovery calls, support conversations. A persona grounded in a company’s own recorded buyer conversations behaves measurably differently from one built on public data alone. This is the single highest-leverage input available and it is sitting unused in almost every company’s CRM right now.
Voice interviews with real people. We shipped an AI interviewer this quarter. A real human who represents an archetype gets interviewed conversationally by an AI, and what they say folds back into the persona profile. This is the piece I am most excited about, because it collapses the false binary the whole category has been arguing about. Synthetic versus real was never the right frame. The winning architecture is synthetic scale grounded in real humans, continuously.
Occupational and demographic structure as the spine. Our persona library is built on US government job classification data spanning roughly 120,000 job titles, which gives B2B personas real task, skill, and work-context structure underneath the behavioral layer.
Here is a finding from this work I did not expect and that I think is genuinely good news for the economics of research: there appears to be a saturation point on interviews. Somewhere in the range of ten to twenty-five interviews per archetype, additional interviews stop meaningfully changing how that persona behaves.
That number matters. It means grounding a high-fidelity archetype is not a hundred-interview project. It is a two-week project. The cost of getting this right is far lower than the industry assumed, which is exactly why adoption is about to accelerate.
The measurement layer is the unlock
The work I am most invested in right now is not making simulations better. It is making the error measurable.
We have been building what we call a delta confidence framework. Run a simulated study, run the equivalent human study, and quantify where they diverge, question by question. Not as a one-time validation exercise for a marketing page, but as a persistent output that ships alongside the findings.
We are running exactly this comparison live with a design partner right now, with the same research questions going through a simulated panel and a real recruited panel in parallel. That work is in progress and I do not have final numbers yet. When I do, I will publish them.
I want to be clear about why I think this is the bullish move rather than the cautious one.
A category that can measure itself gets bought by serious buyers. A category that cannot, gets bought by early adopters and then stalls at the enterprise procurement conversation. Every methodology that ever became infrastructure went through this: it had to become auditable before it could become standard. Measurement is not a hedge against the technology working. It is the thing that lets it scale.
The vendors who build this will not be constrained by skepticism. They will be the reason the skepticism resolves.
Scale is arriving faster than the discourse
One more reason for optimism, and it is a practical one.
Six months ago, running a study meant a handful of simulated participants and a real question about whether the sample was large enough to mean anything. We are now running studies in the thousands of participants, working toward ten thousand, with the cost curve moving in the right direction as the orchestration gets more efficient.
That changes what is possible. At five participants you are getting anecdote. At five thousand you are getting distribution, variance, segment-level divergence, and the ability to detect the effects that only show up in tails. The statistical objection to synthetic research quietly dissolves once the sample size stops being a constraint, and the sample size stopped being a constraint this year.
Combine that with the grounding work and you get something that did not exist twelve months ago: research that is fast, wide, grounded in real behavioral data, and honest about its own error bars.
Where this goes
The consensus pattern emerging across the industry is hybrid. Synthetic for breadth, iteration, and narrowing the option space. Human research for the decisions with real money behind them. As Qualtrics framed it in their own writing on the category, human respondents remain the anchor that keeps synthetic models connected to how people actually think and behave.
I agree with that, and I think it is a far more bullish position than it first appears. It means synthetic research is not fighting for a slice of the existing research budget. It is expanding the total surface area of questions a team can afford to ask.
Think about what that unlocks. Teams currently ask the three questions they can afford to research and guess at the other thirty. When the marginal cost of asking approaches zero, you do not just answer the three faster. You answer all thirty-three, and you send the two that actually carry the decision to real humans with far better framing than you would have had otherwise.
That is not research getting cheaper. That is research getting a hundred times wider, with human validation aimed exactly where it earns its cost.
The methodology is here. The grounding techniques work. The scale arrived. The measurement layer is being built right now.
This is the year it becomes real.
About Marketrix
Marketrix is an AI-powered agentic user simulation platform. We let teams run simulated user research and surveys, A/B testing, UX discovery, and adversarial Q&A against realistic synthetic users. Our PersonaOS is grounded in over 3 million personas across roughly 120,000 job titles, enriched through first-party customer data and AI-conducted voice interviews with real people.
A million users before your real users.
Learn more at marketrix.ai.
If you are running research and want to pressure-test any of this against your own use case, I would like to hear from you. Reach me at irosha at marketrix dot ai or on LinkedIn.
Subscribe below for more on simulation, user research, and where this category is heading.


