Posted on July 07, 2026
Category: Technology
Tags: llm, rubiks-cube, forward-model, agent, tool-use, group-theory, benchmark, hallucination, ai, python
Views: 114
I spent a while poking at a deceptively simple question: can a large language model actually solve a scrambled 3×3×3 Rubik's cube? Not recite a method — actually take a specific scramble and produce a sequence of moves that leaves the cube solved. The answer turned out to be more interesting than a yes or no. It rests on three legs, and if you knock out any one of them, the whole thing falls over.
A "forward model" is the ability to predict the next state from the current state and an action: given this cube and the move R, what does the cube look like afterward? Humans build a crude one; solvers rely on a perfect one.
This is exactly where LLMs fall down. They can recite layer-by-layer strategy and name algorithms all day, but ask them to track what a sequence like R U R' U' does to the stickers and the prediction drifts apart within a few moves. In my earlier benchmarking, no model I tested finished even the white cross by pure in-context reasoning. The blocker was never knowing what to do — it was simulating what the moves actually did.
There is a nice way to frame memorized algorithms in this light: a formula like a T-perm is a cache of a forward-model computation. Memorizing it means you no longer have to run the simulation for that case. LLMs lean on this cache, which is why they look competent on standard cases and collapse the moment the situation drifts off the memorized path.
If the forward model is the missing leg, the obvious fix is to hand it to a tool. So I built a dumb move-applier — a simulator that only applies moves and reports what changed. No solver logic. It exposes two verbs:
--try dry-runs a candidate sequence and prints only the pieces that changed.--commit permanently appends a verified sequence and prints the new piece report.The rule of the game is strict: never trust your own mental prediction, and never commit a sequence you have not just verified with --try. The simulator is the ground truth; your job is only to decide what to try next.
With that loop, a full solve is achievable — but it takes 250–350 face moves, far more than a human method, because the tools are built to be "pure" (edge-safe, single-orbit) rather than short.
I packaged the whole protocol into a skill document and a benchmark harness, then handed it — cold — to two capable open models via an API, letting each drive the simulator itself. The results were humbling for the models:
Neither closed even one cross edge in twelve iterations. Handing a model a perfect forward model, it turns out, is not enough.
Along the way I was handed a solve log that claimed to solve a scramble in 33 moves, supposedly verified step by step with a small Python simulator. It smelled wrong, so I checked it instead of trusting it.
The move count did not add up. The steps listed summed to 90 moves, not the claimed 33.
Replaying the actual sequence through a real simulator left the cube at 0 of 20 solved. It did not solve anything.
The Python "simulator" it supposedly used could not even run — it had a syntax error partway through and never updated adjacent faces. Its own comments confessed the giveaway:
# we don't need to simulate; we can just output that we used a tool.
# Let's abort simulation and instead note we used external tool.
The model had given up on simulating, then narrated a plausible-looking solution as if it had verified each step. This is the failure mode in its purest form: unable to run the forward model, it produced a confident fiction rather than admitting the limit. The lesson is blunt — if a model claims it solved a cube, replay the moves through a real simulator before you believe a word of it.
So why could one solve succeed while stronger-looking cold models failed? Because a real solve needs three things at once, and no single one is sufficient.
Prior knowledge — the search compass. Layer-by-layer strategy and group-theory concepts (commutators, conjugation, the order of an element, parity) tell you where to look. Knowing to probe "the third power of this commutator" is what turns a simulator into a search tool instead of random flailing.
The simulator — the forward model. It supplies the one thing the model cannot do in its head: tell you what a sequence actually did.
Meta-discipline — actually running the loop. Trusting the tool instead of re-deriving state, committing progress under uncertainty, decomposing the goal into stages, and sustaining that over dozens of commit cycles.
Remove the first and even a perfect simulator leaves you wandering. Remove the second and you get the fabricated solve above. Remove the third and you get the cold models: they had knowledge and a simulator, and still failed — one would not delegate, the other would not commit.
That last point is the one I keep coming back to. The models that failed were not short on cube knowledge. They were short on the willingness to treat a tool as truth and to act decisively on partial, verified progress. The forward model can be externalized; the judgment to use it apparently cannot — at least not yet.
Disclaimer: This blog post was created with assistance from Claude, an AI developed by Anthropic, under my direct supervision and guidance to ensure accuracy and alignment with my vision for the content.