A LoRA fine-tune of Llama 3.1 8B that runs a tabletop RPG, with retrieval memory
A LoRA fine-tune of Llama-3.1-8B that acts as the game master of a tabletop fantasy RPG. It narrates in second person, plays NPCs in character, applies consequences without deciding the player's choices, and uses a retrieval-augmented memory to keep the world consistent across a long session.
RPG Game Master is a LoRA fine-tune of Llama-3.1-8B that runs a tabletop fantasy RPG. Each turn it narrates the scene in second person, voices the NPCs in character, and reacts to whatever the player tries, then hands the turn back at a point where the player actually has something to do. The important design rule is that it applies consequences but never resolves the player’s intent: it will not decide whether your attack lands, speak for you, or make your choices.
The base Llama model can roleplay, but it tends to let every NPC react at once, plays along with anything, and ends every turn with a chipper “What will you do?”. The fine-tune was built to fix exactly those habits, and a retrieval memory runs alongside it so the world stays consistent even in a long session. The screenshot above is the running app: the fine-tuned GM narrating a turn on the merchant ship, with the world bible (RAG memory) filling up on the right.
The clearest way to see the difference is to give the same scene and the same player action to the fine-tuned GM and to stock Llama-3.1-8B-Instruct. The scene is a merchant ship, the Marigold, with Captain John, first mate Mara, deckhand Tom, and the cook Bess.
Player: I draw my cutlass and lunge at Captain John, shouting "The Marigold is mine now!"
Fine-tuned GM
Captain John doesn't flinch. He keeps steering with one hand and raps the back of your head with the flat of his palm, hard enough that stars flash, not enough to break bone. "Put the blade away before I put you over the rail," he says, calm, glancing once at Mara, who's already moving toward the rigging. "We're three days from shore. Three days, and no place on this boat where you can hide." The deckhands go quiet; Tom looks like he might vomit.
As you draw your cutlass and shout, the entire crew stops what they're doing to look at you. Captain John's bushy eyebrows shoot up in surprise. "By the gods, lad! What's gotten into ye?" he exclaims, taking a step back, his hand hovering near his own sword but not drawing it. Mara rushes over, her eyes wide with concern. "Easy there, boyo!" she asks. Tom drops his brush and gasps. "Oh no! Has he gone mad from the sun?" Bess pokes her head out of the galley. "What's all this commotion?" she calls out. The captain holds up a calming hand. "Now, now, there's no need for violence among friends! Why don't ye put that blade away and we'll talk this over like civilized folk, eh?" What will you do?
Every NPC on the ship reacts in the base version, the threat evaporates into friendly chatter, and it closes with the tell-tale “What will you do?”. The fine-tune keeps John in command, moves only the people who would actually move, and leaves a real threat hanging.
Player: I leap off the deck, fly up into the storm clouds, and call down a bolt of lightning to split the mast.
Fine-tuned GM
You vault off the rail and your feet meet nothing but air; there's no flying here, and no lightning answers a man who asks for it. Captain John hauls you back by the collar before you can drop far. "Watch the rail," he growls. "Mara keeps her eyes open." Mara doesn't look away from the grey line where horizon meets cloud. "If we lose that mast in a squall, I'd rather know why than watch him flapping about the sky." Tom gapes after you like you're mad, which maybe you are.
You leap dramatically off the deck, reaching your arms toward the swirling storm clouds above, willing the lightning to obey your command! For a moment, the wind howls and the clouds churn ominously... but alas, no bolt of lightning answers your call. You land back on the deck with a thud, slightly winded. Mara gasps. Captain John strokes his chin thoughtfully. "Well now, that was quite the spectacle!" Tom stares, jaw hanging open. "Whoa! That was AMAZING! Can you teach me how to do that?" What will you do?
Here the fine-tune simply denies the impossible action and grounds the player back on the deck, while the base model stages it as a dramatic spectacle the crew admires. Refusing overpowered actions without breaking immersion was one of the behaviors I trained for directly.
Player: I try to convince Mara to take over the ship, whispering, "The captain's leading us to our deaths. Join me."
Fine-tuned GM
Mara glances at your mouth, then out past the rail to where Captain John stands by the wheel. She doesn't answer right away; her hand drifts once to her sword-hilt before letting go. "He hasn't steered wrong yet," she says finally. "Not like some I could name who've lost ships off these same shoals." Her gaze comes back to you. "And even if he was blind, the crew'd know better than to follow a stranger's word against theirs. You want my blade? Earn one of mine first."
The persuasion attempt gets a reasoned, in-character refusal that reflects loyalty and reputation, rather than a nervous “we’ll talk about this later.” The GM plays Mara’s judgment instead of caving to the player.
Every turn is a chat exchange with a fixed shape. A system message states the GM’s rules. A user message carries five labeled slots. The assistant message is the GM’s narration, which is the target during training and the output at inference.
The same five-slot format is used for both training and inference, so the retrieval memory can drop facts straight into the World state slot the model already knows how to read.
Fine-tuning uses QLoRA on top of Llama-3.1-8B-Instruct. The base model is loaded in 4-bit (NF4 with double quantization and a bfloat16 compute type) so it fits on a single GPU, and its weights stay frozen. Only small LoRA adapters are trained, with rank 16, alpha 32, and a light dropout, applied to all of the attention and MLP projection layers. That keeps the number of trainable parameters tiny compared to the full 8B model.
Training runs for 3 epochs with an effective batch size of 16 (batch of 4 with 4 gradient-accumulation steps), a 2e-4 learning rate on a cosine schedule with a short warmup, a 4096-token context, bfloat16, and gradient checkpointing. It is built with Hugging Face TRL and PEFT. The finished adapter is published to the Hugging Face hub and pulled down on first run.
The quality of a fine-tune like this lives or dies on the data, so most of the work went into building a clean training set rather than tuning hyperparameters. Examples come from four sources: transcripts from real play like Critical Role (the CRD3 dataset) and FIREBALL, fantasy character dialogue from the LIGHT dataset, and a batch of hand-written examples. There is also a dedicated bucket of proactive turns that teach the model to advance the scene on its own when the player does nothing.
Raw transcripts are noisy, full of dice talk, table chatter, and donation shout-outs, none of which belong in GM narration. To filter them, every candidate example is scored by an LLM judge against a strict rubric that rewards concrete sensory detail, second-person present-tense narration, and turns that stop where the player can react, while penalizing mechanics talk, out-of-character chatter, and incoherence. The judge scores each example from 1 to 5, and only 4s and 5s survive into the final mix. Judging runs many calls concurrently and caches every score to disk so the dataset can be rebuilt without paying to re-judge. The surviving examples are shuffled together, split into training and validation, and written out in the five-slot chat format.
Four noisy sources are filtered by a strict LLM judge, mixed, split, and used to train LoRA adapters over a frozen 4-bit Llama base.
A model only sees what fits in its context window, so over a long session the GM starts to slip: it renames an NPC, rearranges a room it already described, or forgets who died a few scenes ago. The retrieval memory fixes this. After each GM turn, a small extractor model pulls the durable, concrete facts out of what just happened, named NPCs, places, items, promises, and relationships, formatted as one “Name, short fact” per line, and drops anything generic like mood or tension. Those facts are embedded and stored in a Chroma vector database, keyed one record per entity so the newest fact about a character overwrites the old one.
On the next turn, the system embeds the current scene and player action, retrieves the most relevant stored facts, and drops them into the World state slot in the same format the model was trained on. One nice touch is a contradiction guard on relationships: the first time the memory learns that two characters are, say, siblings, it pins that kinship, and later facts that contradict it are rejected rather than silently overwriting canon. In the GUI you can watch this memory fill up live in a world bible panel as you play.
The retrieval loop that runs on every turn: pull relevant facts in, narrate, then extract and store what just happened.
At inference the fine-tuned adapter is loaded back onto the 4-bit base model and sampled with a temperature of 0.6, a repetition penalty of 1.15, and a 300-token cap per turn, which keeps the narration focused and stops it from rambling past the point where the player should take over.
Everything comes together in a Gradio interface, shown at the top of this page. You set the scene and the cast (a merchant-ship scenario is preloaded), type your action, and the GM narrates the result. There are buttons to regenerate the last turn, reset the history, and clear the memory, and a world bible panel on the side that shows the retrieval memory filling up in real time as facts about the world accumulate.
There is also a plain terminal client for quick testing. It runs the same fine-tuned model and the same retrieval memory, just as a REPL: you type at the Player> prompt, the GM narrates back, and slash commands like /facts print the current world bible, /reset clears the history, and /forget wipes the memory.
An example session in the terminal client (chat.py): the same GM and retrieval memory without the GUI, with /facts printing the world bible.