A synthetic persona is a persistent AI model of an individual person, real or representative, that carries a consistent set of demographics, attitudes, and preferences across multiple research questions. Unlike a static persona slide in a strategy deck, a synthetic persona can be asked a new question at any time and will answer in a way that's consistent with everything it "knows" about itself. In pharma and HCP research specifically, that means a way to pressure-test a message, a claim, or a concept against a modeled physician or patient before committing to a fully fielded study.
A traditional persona is a static composite. A research team runs interviews or a survey, spots a pattern across a segment, and distills it into a single fictional profile, "Dr. Community Oncologist" who prioritizes tolerability over marginal efficacy gains, for instance. That persona lives in a slide. It can't answer a question it wasn't built to answer.
A synthetic persona works differently in three ways: it represents an individual, a panel of synthetic personas captures a spread of views rather than one central tendency; it's queryable at any time rather than fixed at the moment it was created; and it functions as an actual research instrument you can survey and analyze, not a reference slide.
It's worth separating synthetic personas from a related but different idea: asking a general AI model to generate a "typical physician's" answer to a single question. That kind of one-off AI-generated response has no persistent identity. Ask it a follow-up question and there's no memory of the first answer, no consistent preference profile behind it, and often no grounding in real research data at all.
A synthetic persona persists. It carries an identity and a preference profile built from real or representative data, and a new question gets answered in a way that's consistent with that established profile, not generated fresh with no continuity each time.
If you've seen the term "digital twin" used in a market research context, it's describing the same underlying concept as a synthetic persona, a persistent, queryable AI model of a person used across research studies. "Digital twin" comes from engineering, where it originally described virtual replicas of physical systems; "synthetic persona" is the term that emerged specifically from AI-driven market research. The two terms are used interchangeably across the industry.
There are two general approaches. The first generates personas purely from population-level data, useful when a team is researching a population it has no existing data on, but limited to what's statistically typical for a demographic segment rather than any specific individual's actual views. The second, and the approach that produces more precise results, seeds a persona with real data an organization has already collected, survey responses, conjoint results, or qualitative interview transcripts, so the persona's answers are grounded in what a real respondent has actually said, not just what's typical for their segment.
The second approach is what makes synthetic personas useful for augmentation: asking an already-surveyed audience new questions without fielding a brand-new study.
Synthetic personas are a genuinely useful tool, but they come with real limits worth naming. A persona is only as good as the data grounding it: one built purely from generic population data can sound plausible while being wrong about what a specific specialty or segment actually believes. (This is the same underlying risk we cover in more depth in why generic AI tools fall short for pharma market research.) A response also needs to be traceable back to the real data behind it to be useful in a regulated industry, not just a plausible-sounding answer with no way to check it. And synthetic personas are a way to move faster and pressure-test earlier, not a full substitute for fielding real HCPs and patients when a decision genuinely needs that depth. (We cover why keeping a person in the loop on data provenance matters in Human-in-the-Loop AI: How Sagan Agents Combines AI Speed With Expert Judgment.)
Inside Sagan Agents, synthetic personas show up in a few concrete places. The Synthetic Personas agent lets a team start a freeform conversation with live HCP personas, drilling into the source quotes behind every response rather than taking an answer at face value. The Message Testing agent builds segment-specific message portfolios and tests them against HCP personas grounded in real-world primary and secondary data, benchmarked against ZoomRx's 15+ years of effectiveness data, not generic population assumptions. And when a question calls for new fielded research rather than a persona-based read, Agentic Market Research can field to synthetic audiences directly, or to ZoomRx's proprietary HCP and patient panels when real-respondent depth is what the moment calls for.
Data Archive Intelligence's citation model, cited earlier, applies here too: a persona's response can be traced back to the real research behind it, rather than treated as a standalone, ungrounded answer.
The fastest way to see how a grounded synthetic persona differs from a generic AI guess is to run a message or concept past one. Visit the Sagan Agents page to see the full platform, or use the form below to start a conversation about a pilot.