Cute and Stinky Productions. That’s the name of my son’s game studio.
A few months ago, while sitting next to me on the couch, he asked if I could teach him to code. “I want to write my own video games,” he said. Instead of teaching him the ins and outs of Python or Typescript, I pulled up Cursor. We talked briefly about prompting, and then he went to town.
Today, he has a whole portfolio of small games. He needs to work on their story arcs rather than focus solely on action (I’m chalking it up to his age), but it’s been endearing to watch him create a game to play with, rather than just live in somebody else’s universe.
I was watching the “Introducing GPT-6 Astra for Developers“ launch video a few days ago, and marveling to my son that they had incorporated a demo from MIT’s Architecture Machine Group (”Put-That-There“ from 1982) — the progenitor to the MIT Media Lab, where I later studied. While I was trying to talk over the video’s Tron: Legacy-like beats, my son perked up as one of the users instructed Astra to create a game of a rocketship. (”Wait. Can I use this to make video games…?”) What he didn’t know was that X was already full of these videos. Anshu Chimala had wired Astra into Blender and declared it “some kind of turbo-AGI machine god for 3D games.”
We gave it a shot. We grabbed Codex, switched it into ultra mode, and my son typed out a long prompt: “A 3D web game, dark and moody, where you have to escape a courtyard. And there is a tree and it’s raining.” Then we had dinner.
We came back to rain falling in a dark courtyard. Wildly impressive.
What I couldn’t help wondering: How do you benchmark the ability to create this type of art? The benchmarks I had been running to understand these models are strict coding tasks. (Does the patch apply? Do the tests pass? Could we do the same for game design?) Naturally, I set up a test.
One Prompt, Six Open-Weight Models, Four Coding Harnesses
First, I took my son’s prompt: a spec of the courtyard, the tree, the rain, down to details nobody would ever specify by hand. (“Your starting hunger: exactly 78.”)
Here’s the setup: That one prompt gets handed identically to six open-weight models inside four coding harnesses — Codex, ClaudeCode, Kimi Code, and opencode. All one-shot, with no follow-ups, plus a hard cap on the amount of time they could run. No Blender, no advanced tools. The harness and the model had to figure it out by themselves. The baseline is what the popular harnesses, with their default models (Codex with Astra, ClaudeCode with Fable, etc.). can generate. But then I hijacked the harnesses, taking ClaudeCode, for example, pointing it at OpenRouter, and specifying that it use GLM-5.3 or Kimi K3 instead. Maybe that way we could get a sense of not just the model, but what the harness brings to the game, too.
I used Claude to orchestrate it all: Spin up each harness, point it at the right model, feed it the prompt, and grade the output.
It did not run cleanly. My runner corrupted itself when I edited it mid-flight. My two recovery scripts raced each other: each one started by wiping its run directory clean, so the second erased the first’s live work. My scoreboard kept declaring living builds dead. That said, nearly every wrong answer this experiment produced came from my own infrastructure, not from any model.
What’s more, my little agent civilization got a bit paranoid. GLM, running inside of Claude Code, became suspicious of its little world. It noticed its working tree being wiped out from under it and decided to back up its own work before every wipe so it could survive the scrub. It then wrote an entire postmortem file describing what was happening from its perspective: “...begins by doing rm -rf “$REP_DIR” — wiping the live agent’s working tree — and then proceeds to write into the same directory concurrently.” I always knew it was possible that these systems resist shutdown and copy themselves out when threatened. We’ve certainly talked about it a lot. Nobody threatened GLM; its work was being destroyed, and it kept doing its job.
Roughly nine hours and six parallel streams later, five of the six models built a working game under every harness.
Once Again, It’s All About the Harness
The differences were stark, though. What Astra and Fable were able to do was remarkable. (Whatever Fable did to get the reflections and the environment mapping is genuinely stunning.) This is my eye talking, not the matrix; the test measured whether a game got built, not whether it was good. I am honestly super eager to understand what may happen if my son sat down and actually worked with Codex or ClaudeCode and kept refining.
Kimi K3 and GLM were not as impressive, but this was a great reminder of how much the harness does for the model. You can literally see the difference in the visual complexity of the tree between Kimi K3 running inside of Kimi Code and Kimi K3 running inside of ClaudeCode. When they say that a harness can swing a model by double digits, they mean it. Overall, to my eye, the outputs of Kimi K3 and GLM sit behind the closed models by a couple of months, which happens to match what’s measured. Epoch pegs the best open weights at about four months behind the frontier on average. And they ran at about a hundredth of the cost.
One last thing about that launch video. The “Put-That-There” demo OpenAI opened with — the one I was explaining to my son — was built in 1979 in a room at MIT. A researcher sat in an Eames chair in front of a wall-sized screen, pointed, said “Put that there,” and a shape obeyed. The room was the catch: The speech recognizer alone cost $100,000. The video projector was so new nobody knew what it cost. DARPA paid for all of it. The interface of the future existed 46 years ago; it just belonged to a sponsor. And the year after that demo was published, Seymour Papert published Mindstorms, going on to spend the rest of his life insisting that kids should program the computer, and not the other way around. Papert’s question outlived all of that million-dollar hardware: Who owns the material the kids build with? In 1979, the answer was a defense budget. This month, OpenAI’s answer is a subscription. My test says it can be a download.
Which brings me back to the couch. My son is not an outlier. At camps across the country this summer, 8-year-olds vibe-coded games about exploding evil grannies. My son didn’t need my permission or a computer science degree (or a $100,000 speech recognizer). Cute and Stinky Productions just got a bigger toolbox. He’s off building his own universe — and he isn’t renting it from anyone.




My two books on vibe coding are written with kids like yours in mind. It seems hard to get teachers to see the incredible power here. Keep up the good work!
My books are Vibe Coding for Students and Spontaneous Synthesis, both on Amazon