Simul Endgame
Here we continue experimentation with the mastermind-like puzzles for model evaluation, except we focus on endgame.
Let’s use same terminology:
- Game/Puzzle: a single instance of Mastermind task with a single secret code by Codemaker.
- Session: a set of K >= 1 games running concurrently within one conversation.
We give model a past history of a simul session, formatted as
list of tuples (Codemaker, Guess, Reply), and expect
model to produce a final answer for all puzzles in a single turn.
The goal is to evaluate long single-turn reasoning, not multi-turn
conversation. This sort of inference is not nesessarily ‘what models
are good at’ - I’m sure models would have much better luck
implementing a solver in python. Yet, this reasoning is definitely
relevant to intelligence.
Dataset
Greedy
player solved 1000 games to completion. In each of those games
last turn is removed - so, greedy player would solve each game from
the dataset in a single guess. Then, for each session we pick
4 <= n_games <= 12 games from the dataset, select
history for them and tell model to solve them all. Corruption
probability was fixed at 0.2. The history is ‘stable-shuffle-merged’
- the order of moves within each game is preserved, but the merge
itself is randomized.
Rules
Instructions tell model that it has a single turn to complete as many games as possible. Model needs to call the single tool available with a list of (Codemaker, Guess) combinations. The fact that greedy player can finish it in a single turn doesn’t mean there’s only one valid option left - we might have large set of valid guesses with different probabilities. Corruption probability is shared with model in prompt. Model is allowed to produce one guess per game.
Results
We study how model behavior changes based on the simul size, and,
more importantly, if the test itself is reasonable and can be used
to understand and optimize reasoning settings. Deepseek
v4.1 flash and GPT-6.1-Sol
are evaluated. For each size 4 <= n_games <= 12
we run 10 sessions, and the main outcome we look for is the number
of correct solutions to puzzles.

There are three types of outcomes:
- Solved. Model submitted a guess and it was correct
- Incorrect. Model submitted a guess and it was incorrect
- Failed. Entire session failed because model didn’t call a tool/ran out of reasoning tokens. Note that Sol was running with 64K max tokens and Deepseek with 128K - as our goal is not comparing models to each other, this is appropriate.
We also track number of tokens:

Deepseek
Quality clearly goes down, and model runs out of reasoning even
at low effort. Number of tokens seems growing, but
variance is high. If we were to tweak reasoning settings - logit
bias, hard budget, etc - we should focus on
n_games = 4-8 for evaluation.
Sol 6.1
Sol is consistently strong, but gets several failures at higher
end of n_games. What is interesting here is the token count. Fitting
linear model shows n_tokens = 2.8k + 2.87k * n_games -
the fixed part of the cost is very small. That was a little
surprising, I expected model to spend more tokens on figuring out
overall shared algorithm, which doesn’t seem to be the case.
For benchmarking/tuning/improving Sol we need more complicated version of the test - more samples, larger reasoning window, maybe stronger solver to produce games.