Somewhere around week six of an eight-week obsession he had not planned for, PewDiePie found himself unable to stop running training loops at two in the morning, watching a tiny AI model improve, flatline, and improve again, knowing full well the video announcing it had already been late for two months. Ajax, his fine-tuned local AI model named as a reference to the cleaner spray, was almost done. Almost.
The question of why a creator builds his own AI model from scratch when far more powerful ones already exist is, in PewDiePie’s view, exactly the wrong question. Most people, he argues, are reaching for a trillion-parameter model just to sort their inbox, the rough equivalent of hiring a stadium full of servers to answer one email. His estimate: it would take around 27 versions of his own computer, and roughly 150 houses worth of electricity, to run a single instance of the kind of model OpenAI or Anthropic offers. That math drove the whole project.
Getting banned twice before the model was even finished
The path to Ajax ran directly through a terms-of-service wall. PewDiePie had identified a published study showing a method to decrypt the reasoning tokens that OpenAI’s models use for chain-of-thought processing, the hidden problem-solving layer the company strips from its outputs before anyone sees them. He attempted to use this via their own API to generate seed training data for Ajax. OpenAI banned his account. He appealed, was unbanned, ran the process again, and was banned a second time. His genuine bewilderment at the speed of that second ban came through clearly: ‘HOW DID THEY EVEN KNOW?’
The decensoring process proved equally strange to explain. Rather than a simple toggle, removing a model’s refusal behaviors requires identifying where those behaviors live across the model’s weights, which is not a single location but a pattern spread throughout. The method he used, an open-source program called Heretic, works by feeding the model dialogues it would normally decline, identifying the refusal patterns, and ablating them. The tradeoff is real: the more you remove, the more general capability can degrade. He drew his own line at content that causes harm to others or to the user.
The data problem nobody helped him solve
Before any of the training could happen at scale, PewDiePie needed data. Because his Odysius harness is built on a strict no-collection policy, he could not pull it from his own users. So he built a dedicated website, data.pdy.com, where people could voluntarily submit their data, spent several days on it, launched it, and then sat back waiting. Nobody submitted anything. ‘I was as close to in tears as I have been in a very very long time,’ he said, ‘because no one still gave me the data.’
He kept building anyway. Ajax version one runs locally, browses the web, manages calendars and to-do notes, handles email sorting, and completes tasks successfully nine out of ten times by his own benchmark. The GRPO reinforcement training he added ran for four weeks, hitting a wall on day four when he exhausted his task set, requiring him to synthesize new ones mid-run. The model improved, flatlined, then improved again.
The training loop that refused to end
The countdown timer he added to the project’s release window was, by his own admission, primarily there so no one could complain about further delays. He knew he would blow past it.
Ajax is available now through Odysius. PewDiePie framed its purpose plainly: getting information to you quickly, on your own hardware, without feeding data to a system built on other people’s data. ‘I really strongly believe in the small model trained for a specific harness.’ The countdown timer eventually ran out. He shipped it anyway, then immediately started running the next version.


