An API for believable people
PersonaGen came from an idea I couldn't leave alone: an API that generates believable people. I built the first version out of curiosity, then kept going as I tried to make each life fit together and whole groups stay varied.
The main technical bet was to use LLMs to build conditional probability maps ahead of time, then sample those maps at runtime. That gave me millisecond responses and a result I could reproduce from a seed.
Getting usable numbers from LLMs
I built the probability maps one trait at a time. For education, that meant a base distribution plus different odds by age and parents' education. I didn't expect census-grade figures from an LLM. I wanted the relationships to hold up when I generated thousands of people.
That took roughly 94,000 calls and 290 million tokens across 25 models. I tried around 30 prompt variants and regenerated the whole corpus three times. Some models couldn't reliably return valid JSON. Others returned clean JSON with suspiciously flat probabilities, or hedged on sensitive questions.
Names needed their own pipeline. I enriched about 30,000 with ethnicity and age-band data so a British Pakistani profile might draw Muhammad, while an Irish American one might draw Siobhan.
StackNext.js (App Router), TypeScript, PostgreSQL, Tailwind, BetterAuth, Nextra, Bun, Elysia, Drizzle ORM, Railway, OpenRouter
Making 123 dimensions fit together
The generator has 123 dimensions. I put them in tiers: immutable traits, life circumstances, then choices. Each one depends on roughly 4 to 12 others; any more and the probability maps explode. Age and income matter for housing; eye colour doesn't.
The first versions made believable individuals, then drifted when I generated thousands. Multiplying the conditional probabilities made the problem worse. I tried around 10 to 15 ways of combining them against large batches before settling on an approach I could calibrate against population baselines. Early correction rules patched obvious oddities but skewed the population, so I switched them off once the sampler improved.
The first runtime took more than 200ms. Indexed lookups, cached maps and cheaper sampling brought it under 20ms, with no LLM call and repeatable output from the same seed.
What I use it for now
I could have shipped a smaller version sooner. Pushing to 123 dimensions exposed the population-level drift and sampling problems that became the most interesting part of the work. PersonaGen now serves UK and US profiles, and I use it in nxk and other experiments.