Thousands of decisions
Navigation, puzzles, team building, resources and turn-based tactics, chained across dozens of hours with no reset button.
Charmi builds harnesses that let frontier models play real Pokémon games from start to finish. We read game memory, give the model typed tools, and verify every step against the game's own flags. Then we stream the run, live.
The Champions Harness broadcast, designed in Figma, with a retained game frame. Team panels show sample data.
Why games
Most benchmarks end after a few turns. A Pokémon campaign runs for thousands of decisions, and the cartridge keeps score.
Navigation, puzzles, team building, resources and turn-based tactics, chained across dozens of hours with no reset button.
Badges, story flags and Pokédex entries live in RAM. Our checkpoints read them directly, so progress is a fact, never a model's claim.
Runs stream with the model's reasoning beside the game. Failures are visible, debuggable, and genuinely fun to watch.
How it works
Every harness runs the same loop. The model makes the decisions. Deterministic code does the precise, mechanical work and checks the result.
A bridge inside the emulator streams memory. Party, bag, map, dialogue, menus and battles become one structured observation.
map Route 119 party 6/6 · HP 61/74 badges 5/8 dialog none
The model gets the current objective, the observation, and only the tools that make sense in this scene.
“Weather Institute is north. Heal first, then walk.”
One typed tool call becomes bounded controller input: a path walked with A*, a dialogue advanced, a move picked.
walk_to(x=18, y=9) → 14 tiles · 3 batches
The harness reads the game again. Position, HP, items and story flags confirm what happened before the next turn.
✓ position == (18, 9) ✓ in_battle == false
Product
One foundation, refined across ten games, from 8-bit Johto to the GameCube's Orre.
Verified progress
Each campaign is a chain of checkpoints, and each checkpoint is backed by a predicate on game memory. A model saying “done” never completes an objective.
Operator control
Run sessions, steer the agent when it matters, and take the controller back in one click. This is the Champions control room for live ranked doubles.
Model-agnostic
The harness owns the controller, the memory and the checks. Swap the model without touching the game, and compare runs on equal footing.
Up to 78 per game: A* walking, map routing, dialogue, battles, catching, HMs and PC transfers.
The same tool surface over MCP, so any MCP client can drive a game directly.
Bounded controller input only. No memory writes, cheats, resets or save edits in the model's tools.
Every bridge checks the exact ROM revision before a run. Hacks and other versions get their own profile.
A transparent 2560 × 1440 OBS overlay with exact capture geometry for each console.
Dated validation records separate live checks from fixtures. We say what isn't verified yet.
Harnesses
Each harness is built for its game's memory layout, mechanics and story, and shares the same loop, controls and broadcast.

Littleroot to the Hall of Fame, with double battles, cable cars, currents and Victory Road handled natively.

New Bark Town to Mt. Silver, with pathfinding, map travel, HM use, PC transfers and puzzle tools.

Pallet Town to the League. Silph Co., Cinnabar quizzes, spin floors and boulders get dedicated helpers.

Modern mechanics in a GBA shell: Fairy typing, Mega Evolution, Z-Moves, Dynamax and 120 TMs.

Its own memory layout: regional forms, abilities, natures, EVs and a 20-box Pokémon database.

RAM-only play in a pinned Dolphin build, with a registry that tracks every Shadow Pokémon.

Native dual-screen control with touch input, a Unova control room and a 4× DS window for OBS.

Competitive doubles: plans both allies under the battle clock and checks the screen before each input.

Our first DS port: DeSmuME in-process, touch input, Gen 4 decryption and a local vision model.

The first harness: a Lua server inside mGBA, a 14-tool Python layer, an MCP server, a reference Claude tool-use loop and a live dashboard. Every harness since grew from it.
Harness repositories are private while we prepare public releases. Every harness needs your own legally obtained game. We never distribute ROMs.
Made to watch
Each harness ships a public broadcast with the model's live notes, a private trainer's desk, and a transparent OBS overlay that wraps the game capture.
Emerald · Sapphire, Ruby and Emerald color cycle with animated party cards.
Real broadcast interfaces with gameplay placed where OBS puts the game capture. Panels show sample data.
Party, objective, badges and the model's latest note, for anyone with the link.
Model, reasoning and fallback settings, run controls and checkpoint restores.
A transparent 2560 × 1440 browser source that sits on top of the game capture.
Community
Charmi started as a Spanish-language Pokémon channel. Today that audience is where the harnesses go live, so every run is watched by people who know these games inside out.
Research & infrastructure
Long runs need models that never sleep. We operate our own GPU cluster, open-source what we learn, and speak the protocols Claude already knows.
Live telemetry, allowlisted model deployment, and a public recipe that serves Qwen3.8-27B with DFlash2 speculative decoding on a single Spark.
The first harness shipped with a reference Claude tool-use loop and an MCP server. The harness walks, battles and verifies. Claude decides.
# one turn of the loop
resp = client.messages.create(
model="claude-opus-5-5",
tools=TOOLS_SCHEMA, messages=messages)
for block in resp.content:
if block.type == "tool_use":
harness.call_tool(block.name, block.input)
What's next
Full campaigns played by Claude, streamed live, with every tool call and checkpoint on screen.
Cross-model results scored by the game itself: checkpoints per hour, tool efficiency and unassisted recoveries.
A shared engine across mGBA, DeSmuME and Dolphin, with public MCP servers for anyone's own cartridge.
FAQ
Anything else? Write to contact@charmi.us.
That's the goal, and we're upfront about where we are. Our Emerald run reached the Hall of Fame with human help along the way. Every harness keeps a dated validation record, and none claims a flawless unattended playthrough yet.
Any model with tool calling: Claude through MCP or native tool use, OpenAI Codex, xAI Grok, and local OpenAI-compatible servers such as our DGX Spark cluster. Fallback between providers is opt-in, per run.
No. The model's tools only produce bounded controller input, the same buttons a player presses. Memory writes, cheats, resets and save edits are outside its reach.
Never. Every harness requires your own legally obtained copy and checks its fingerprint before a run. ROMs, saves and credentials are excluded from our repositories.
No. Charmi L.L.C. is an independent company. Pokémon and all related names are trademarks of Nintendo, Creatures Inc. and GAME FREAK inc.
Labs, researchers and partners: we'd love to hear from you.