← home

Simul Endgame

Code

Here we continue experimentation with the mastermind-like puzzles for model evaluation, except we focus on endgame.

Let’s use same terminology:

We give model a past history of a simul session, formatted as list of tuples (Codemaker, Guess, Reply), and expect model to produce a final answer for all puzzles in a single turn. The goal is to evaluate long single-turn reasoning, not multi-turn conversation. This sort of inference is not nesessarily ‘what models are good at’ - I’m sure models would have much better luck implementing a solver in python. Yet, this reasoning is definitely relevant to intelligence.

Dataset

Greedy player solved 1000 games to completion. In each of those games last turn is removed - so, greedy player would solve each game from the dataset in a single guess. Then, for each session we pick 4 <= n_games <= 12 games from the dataset, select history for them and tell model to solve them all. Corruption probability was fixed at 0.2. The history is ‘stable-shuffle-merged’ - the order of moves within each game is preserved, but the merge itself is randomized.

Rules

Instructions tell model that it has a single turn to complete as many games as possible. Model needs to call the single tool available with a list of (Codemaker, Guess) combinations. The fact that greedy player can finish it in a single turn doesn’t mean there’s only one valid option left - we might have large set of valid guesses with different probabilities. Corruption probability is shared with model in prompt. Model is allowed to produce one guess per game.

Results

We study how model behavior changes based on the simul size, and, more importantly, if the test itself is reasonable and can be used to understand and optimize reasoning settings. Deepseek v4.1 flash and GPT-6.1-Sol are evaluated. For each size 4 <= n_games <= 12 we run 10 sessions, and the main outcome we look for is the number of correct solutions to puzzles.

There are three types of outcomes:

  1. Solved. Model submitted a guess and it was correct
  2. Incorrect. Model submitted a guess and it was incorrect
  3. Failed. Entire session failed because model didn’t call a tool/ran out of reasoning tokens. Note that Sol was running with 64K max tokens and Deepseek with 128K - as our goal is not comparing models to each other, this is appropriate.

We also track number of tokens:

Deepseek

Quality clearly goes down, and model runs out of reasoning even at low effort. Number of tokens seems growing, but variance is high. If we were to tweak reasoning settings - logit bias, hard budget, etc - we should focus on n_games = 4-8 for evaluation.

Sol 6.1

Sol is consistently strong, but gets several failures at higher end of n_games. What is interesting here is the token count. Fitting linear model shows n_tokens = 2.8k + 2.87k * n_games - the fixed part of the cost is very small. That was a little surprising, I expected model to spend more tokens on figuring out overall shared algorithm, which doesn’t seem to be the case.

For benchmarking/tuning/improving Sol we need more complicated version of the test - more samples, larger reasoning window, maybe stronger solver to produce games.