Skip to content
The AI Whisperer* Contact
AI Personality

AI agent personality without a script: how Jev decides for my villagers

How six AI villagers in Jevs Village get distinct personalities from a decision model: perception-style state, typed choices, sampled actions, and an LLM that only narrates.

10. October 2026 10 min read
AI agent personality without a script: how Jev decides for my villagers

Most writing about AI agent personality starts with a prompt. Give the model a name, a backstory and a few adjectives, and hope the character holds.

In Jevs Village, the six villagers work differently. None of their actions is written by a language model. Each decision is a choice from a closed list, made by Jev, TypeSafe’s decision model, and then checked by ordinary code. The language model only arrives afterwards, to put words to what already happened.

In my introduction to the village I described who lives there. This article is about the machinery: what Jev sees, what it is asked, how its answer becomes an action, and where personality sits in that loop. I also include the things that went wrong, because that is where I learned the most.

Personality as a distribution, not a sentence

Jev does not write text. It receives a state and a set of typed questions, and returns typed answers: a probability distribution over options I define (Choice), a level on a scale (Score), or the probability that a yes/no question is yes (Noul).

That changes what personality means. A villager’s character is not a paragraph the model tries to imitate. It is the shape of the distribution over what they might do next.

Take an illustrative case, not a logged one. When Mira, the miller, has grain waiting and a neighbour at the door, Jev might give “keep milling” 0.62, “talk to Bram” 0.27 and “rest” 0.11. Bram, in a similar spot, distributes differently, because his situation and character read differently. The difference between the two villagers is measurable before anything happens. I can log that Mira was 27% inclined to stop working. An LLM agent can only tell me afterwards, in prose, why it did what it did.

One decision, step by step

Here is what happens when a villager is consulted in the live village. Not every tick reaches Jev. A villager is only consulted when someone interacts with them, when they perceive an event, or when the scheduler says it is time. A code arbiter then decides whether Jev is needed at all or a routine is enough.

1. Code recalls memories, and Jev ranks them

Each villager has a personal long-term memory store, searched by meaning. Code shortlists up to 24 candidate memories. Jev gets a first request with one question: which of these matters now?

That ranking is a separate request on purpose. Questions inside one Jev request are evaluated in parallel against the same state. They never see each other’s answers. If memory selection and action choice shared a request, the action could not depend on the memories chosen.

Code then blends Jev’s probability with importance and recency, drops anything Jev rated at 5% or less, and keeps four memories.

2. Code writes the state as perception

This is the step I underestimated most, and I come back to it below. The state describes the villager the way a person standing in that place would experience it: who they are in a phrase, what they want, where they are, what they see, who is around, what they have been doing, how they feel, what they owe, what they have just heard, and the four memories from step 1.

3. Jev chooses the action

The second request carries the main question. In the code its instruction reads: “What does this person actually do next? Choose the one action that genuinely fits, weighing the alternatives against each other.”

The options are only actions that are legal right now. Code enumerates them from the world: you cannot sell flour you don’t have or talk to someone who isn’t there. Jev cannot invent an action. The same request carries smaller follow-up questions, such as which variant of an action to take, or whether to respond to someone straight away.

4. Code samples instead of taking the top answer

This is the step that makes personality visible. Code does not take the most probable option. It corrects the distribution so that actions with many variants are not over-weighted, then draws from it with a random number generator seeded per run, tick and villager.

A villager who is 0.62 / 0.27 / 0.11 inclined will sometimes do the less likely thing. That is what makes them feel like people with tendencies rather than vending machines with one output per input. Because the seed is fixed and Jev’s answers are cached, a whole run can be replayed exactly. In one live 24-tick run, replay reproduced every decision byte for byte with zero API requests.

5. Code validates and commits

Before anything changes, code checks the answer again. Is the chosen action still legal in the current world? If not, a labelled default runs. Does the villager owe someone an answer? Then the obligation takes priority. Only then is the action committed to the world and written to the event log.

6. The language model renders

Afterwards, an LLM turns the logged event into a thought, a line of dialogue or a chapter of the village story. It works from a public character card for each villager. Code checks every line it writes: plain English, no invented names or numbers, no promises the villager did not decide, nothing private.

The rule I designed the village around: the renderer is a leaf. It reads the log, and nothing reads the renderer. The character cards are only imported by the voice modules, so editing a card changes how a villager sounds, never what they decide.

The biggest lesson: framing is the design

My first version of the state looked sensible. It was a tidy summary: the villager’s current plan, and a vector of drives like hunger 0.7, duty 0.8, sociability 0.3.

Jev answered with P(top) = 1.00 and an entropy of 0.00. It was not deliberating. It was reading back a decision the state had already made. The judgement belonged to whoever wrote the state, which was me.

I tried to fix it with better question wording: neutral labels, explicit exclusions, instructions to weigh trade-offs. It changed nothing. P(top) stayed at 1.00. But when I added a direct instruction for a test (“Choose rest”), Jev obeyed immediately. Jev follows directives, not considerations. Whatever reads like a directive in the state will win.

Then I rewrote the same information as perception. Not “plan: mill grain”, but what she sees, what has been happening, what is owed. On the same villager, P(top) fell from 1.00 to 0.72, and the number of options with at least 10% went from one to three. Across a 20-tick live run with 117 judgements:

Measure Decision table Perception
Most common action (approach_person) 60% 39%
Median P(top) 0.82 0.72
Decisions with P(top) = 1.00 frequent 1 of 117
confide (an implausibly common action) 29% 15%

Single facts now move the distribution in ways that make sense. Adding “nobody is waiting on her” dropped Mira’s top option to 0.38: free of obligation, she had more real choices.

For AI agent personality, this is the point I would underline. The traits are not where personality lives. The way you describe the situation is. If the state already contains the answer, there is no character left to express.

One honest caveat: the live state still includes each villager’s dispositions as numbers alongside the perception text. I have not re-run the ablation on that version, so I can’t yet say how much those numbers flatten the distribution.

Two mistakes from parallel questions

Jev evaluates all questions in one request independently. I knew that and still fell for it twice.

The abandoned-plan question. I asked Jev whether the chosen action abandoned the villager’s current intention. But that question could not see the chosen action, because it was answered in parallel with it. It said “no” almost every time. The mechanism built on it did nothing for 72 agent-decisions before I noticed.

The fix was to split responsibility. Code compares the sampled action with the plan, which is an exact check. Jev answers only what it can judge from the state: is the current plan still the right thing to be doing?

The unanswered question. Jev rated “owes a reply” between 0.81 and 0.97 in all 18 cases where someone was waiting for an answer. Yet in 7 of them, the parallel action question chose something else. Now code discharges the obligation directly.

The rule I took away: a question may reference the shared state, never another question’s answer. If one judgement depends on another, it belongs in code or in a later request.

Memory that the villager chooses to keep

At night, villagers dream. In practice, Jev reviews the day’s experiences in batches. Each memory gets one Noul question: how worthwhile is this experience for lasting personal recall?

At 0.60 or above, the memory moves into the villager’s long-term store. At 0.40 or below, it is forgotten from the save file (the event log keeps it). Protected memories always stay.

This connects directly to what I wrote about memory in the AI personality engineering guide. Continuity does not come from storing everything. It comes from consistent choices about what matters, and here those choices are made per villager, from their own situation.

When a villager should stop and think

Recently I added a mind: an LLM that can reflect on a villager’s situation and write a short note to themselves. It does not get to decide anything.

Instead, code may add one more option to Jev’s list: stop a moment and think this over. It only appears when there is a reason (the villager is stuck, torn, unsure, or facing something new) and at most six times a day. If Jev chooses it, the mind writes a note, and that note becomes a memory. Jev never sees the mind’s words as an answer. It sees a memory, on the next decision, like anything else the villager remembers.

So even reflection stays inside the same loop: the LLM adds to perception, and Jev still chooses.

What it costs and what I haven’t measured

On the prototype tick, one villager decision used about 2,000 tokens and cost about $0.00009. Latency was around 300 ms at the median and 700 ms at p95. I ran 200 simultaneous requests without a single rate-limit error. The constraint is the number of requests, not tokens: one request carries exactly one villager’s state, so six villagers mean six requests per tick. The live village runs on a capped daily budget of €5.

What I have not measured: whether the behaviour is good. Plausibility, so far, is judged by reading the output. Accuracy and calibration are unmeasured. TypeSafe’s own confidence documentation is clear that confidence is derived from the distribution and is not a guarantee. And Jev can only ever choose from actions I wrote, so what the village can surprise me with is limited by my action vocabulary.

Five rules for AI agent personality with a decision model

If you want to try this pattern, these are the rules I would start from:

  1. Describe perception, never a decision table. If the state contains the answer, there is no personality left to express.
  2. Sample from the distribution. The top answer only gives you the average villager. Sampling with a seed gives you tendencies, and keeps runs reproducible.
  3. Keep exact checks in code. Legality, obligations and plan deviation are rules, not judgements.
  4. Never let one question depend on another in the same request. Split into requests, or move the dependency to code.
  5. Keep the LLM a leaf. Let it give the character a voice, check every line, and never let its words change what happened.

That is what I mean when I say Jev can act like a fast nervous system for agents. The personality isn’t a voice laid over the top. It is in the shape of every small choice, and you can measure it.

You can watch the villagers live at jevsvillage.com.

Cao Hung Nguyen
Cao Hung Nguyen

Cao Hung Nguyen writes about AI personality engineering, conversational AI, memory, companions, and responsible human–AI interaction.

About Cao
All articles

Related reading

Thinking through an AI personality?

Describe the interaction, memory, or safety question you are working through.

Get in touch