Why I made it
nxk is a skill for coding agents. It lets an agent ask simulated people a question, see who answers differently, and decide what to investigate next.
I built PersonaGen to generate detailed synthetic profiles whose jobs, finances and households fit together. I wanted to know whether an LLM could use those differences to give more nuanced answers than a generic persona prompt.
I spent months building a full research app around that idea, with guided setup, question checks, raw responses and reports. I got close to launching it, but the UI kept growing to cover new kinds of studies. An interesting result meant setting up another study. A coding agent, on the other hand, already has the task context and can ask again. I decided the app was getting in the way of the idea, so I canned it and made nxk a skill.
How I built it
I built nxk by distilling months of work with traditional LLMs into instructions an agent could use mid-task. I tried roughly 50 models and kept finding that who I asked and how I framed the question mattered more than reaching for a frontier model. Written answers often sounded like the same assistant with different names. Single votes hid whether someone leaned 51% or 90% towards an answer. Those findings shaped the skill.
nxk tells the agent to draw a relevant group, ask each profile a concrete question through an answering model, keep every option's probability and inspect who answers differently. If it's choosing library hours, it can first ask who's likely to visit. It weights later answers by each person's chance of visiting, then follows useful differences, like full-time workers leaning towards evenings. The next question might be which evening works. That loop runs inside the task the agent is already doing.
Core piecesPersonaGen, Answering model, SKILL.md, Probability analysis
The hard parts
Making that loop cheap enough for ordinary use was the hard part. With traditional LLMs, a study meant hundreds of concurrent requests, some taking 30 seconds, plus retries when structured answers failed. A run could cost about 50p. I got it working, but it felt too cumbersome. I wanted research to be a quick call an agent could make whenever its next decision might affect a real person.
Jev changed the economics and the answer shape. It returns typed option probabilities, where I'd been trying to coax useful probabilities out of ordinary LLMs. In one recorded run, 100 people answered two questions in five seconds for 1.3 US cents. nxk recommends Jev but doesn't require it.
Speed wasn't proof of accuracy. I checked Jev against 34 questions from published surveys. Some misses came from using the wrong audience; others exposed model biases, such as people sounding too positive or too cautious. A model can weigh a new offer, but it cannot know private facts missing from a profile. The skill tells the agent to check whether its audience can answer the question, then flag results that look suspiciously uniform.
Where it stands
nxk is open source now. I ended up with a smaller tool that fits the work better than the app did. I use it for a directional read, then check important decisions with real people. I'm interested in what others ask it, especially questions I'd never have built a study template for.