Simul
Code.
Model needs to solve a variant of Mastermind (Bulls & Cows) puzzle: each opponent (‘Codemaker’) comes up with four-digit secret number: all digits distinct, number cannot start with 0. On each turn model (‘Codebreaker’) submits a query - another number satisfying the same constraints. The response consists of two numbers: how many digits were correct and in the right position, and how many were correct, but in wrong position. The game continues until the secret is guessed or maximum number of turns exceeded.
This well known version would be too easy, with many branches being directly exposed during training. Thus, the noisy modification of the puzzle is used: some of the digits might get replaced with ‘?’, and the response will be computed based on the corrupted guess. ‘?’ doesn’t match any digit. Corruption is stochastic, Codemakers are not strategically lying on purpose. Correctly guessed secret is a special case and is never corrupted - if player was supposed to get ‘4 0’, ‘4 0’ will be returned and the game will end. Player’s objective is unchanged: find the secret number based on this imperfect information.
Model will take a role of a grandmaster playing a simul session against multiple Codemakers. On each turn model will be given a specific Codemaker to query. This allows to partially decouple problem complexity from problem size and study long context quality degradation.
Let’s use following terminology:
- Game/Puzzle: a single instance of Mastermind task with a single secret code.
- Session: a set of K >= 1 games running concurrently within one conversation.
This task requires very little factual knowledge, has tiny input. Player needs to keep track of what was asked, deduce what it implies and decide on the query strategy. For each puzzle within a session we can track if it was solved or not and how many turns it took. For entire session we’ll know number of tokens. It will be misleading to try to attribute number of tokens to individual games, as reasoning about overall approach and algorithm will be reused.
As we can see, the task is a combination of long context + multi-turn reasoning.
Setup
The harness has just a single tool to make a guess. P(corruption) is fixed at 0.2, each digit corruption is independent. So, P(‘all 4 digits were corrupted’) is 0.2 ^ 4 = 1.6%. P(‘No digits were corrupted’) is 0.8 ^ 4 = ~40%.
Model gets 50 turns for each puzzle, the budget is shared across every puzzle in the session. Model is explicitly told the objective: optimize for average number of turns (not worst case), the rules and limits. After simul session is ended, new one is restarted from scratch - games within session share entire context, sessions share nothing.
Errors
If model submits invalid guess (for example, ‘0123’, ‘1’, ‘2222’), the turn is wasted, but game continues. Other soft failures include making more than one guess, or directing the guess to the wrong Codemaker.
Hard error scenario happens when model exceeded output token limit or total turn budget. In this case, entire session is stopped, and only games solved by that moment are counted.
Baseline results
Before looking at the models, let’s establish baseline for individual game as a sanity check.
Greedy player
Greedy player starts with uniform prior for every valid secret code and after every round computes posterior probability for each secret and picks the max. It is a suboptimal player. For example, typical opening like
1234 | ...
5678 | ...
will be not be played at all, if at least one of ‘1234’ is a match and it was not corrupted. P(secret = 5678) = 0, and greedy player won’t even consider it. Player which estimates information gain and combines it with greedy P(win-right-now) will do better, but for our purpose greedy baseline will work. With P(corruption) = 0.2 greedy baseline averages 7.7-7.8 steps.
Randomized player
This player is the exact same greedy one, but move sampling temperature is set to +inf instead of 0.0. Player picks the action uniformly from all candidates with p > 0. Such player can solve puzzle in 11.3 turns on average.
LLM Results
This is a pilot study to get a sense of what we can expect from modern models. Let’s establish the expectations first and see if they will be confirmed.
- single-game sessions will be reliably solved by modern models.
- single-game sessions turn count will be worse than greedy baseline, but better than randomized;
- larger simuls will see performance degradation, as context grows
- we’ll see some variance and might need to adjust the task accordingly.
Deepseek-v4.1-flash
The model has 1M context window, 128k output per turn. reasoning_effort=‘low’ was used.
| n_games | n_turns | n_tokens | pct_solved | n_sessions | n_warnings |
|---|---|---|---|---|---|
| 1 | 16.80 | 44223.55 | 100.00 | 20 | 0.00 |
| 2 | 17.65 | 31309.17 | 97.50 | 20 | 0.10 |
| 4 | 16.27 | 20828.00 | 100.00 | 10 | 0.30 |
| 6 | 25.13 | 23709.05 | 96.67 | 10 | 0.43 |
| 8 | 21.18 | 18980.41 | 96.25 | 10 | 0.22 |
| 12 | 31.61 | 14483.87 | 85.00 | 10 | 1.52 |
| 16 | 34.44 | 13019.00 | 90.62 | 10 | 1.21 |
| 20 | 39.36 | 12209.55 | 64.00 | 10 | 1.72 |
| 24 | 36.39 | 11295.10 | 77.50 | 10 | 1.31 |
| 28 | 47.39 | 13395.43 | 77.14 | 10 | 6.69 |
Note: n_turns, n_tokens and n_warnings are normalized per attempted game, not solved game.
As we can see from this table, for sessions with a single game, model reliably solves it at ~16-17 turns. As number of simultaneous games grows, number of turns increases and more puzzles are not getting solved. The main reasons for not being solved are turn limit and token limit.
This is very interesting, as the complexity of each individual puzzle has not changed at all - just mixing them up into longer sessions caused this.
We can also see a sharp increase in warnings. The most common warnings were: model was querying wrong Codemaker, so, instruction-following had issues.
Variance looks a little high, yet we can be confident in the trend.

Here we see a drop in number of games solved.
And here - increase in number of turns per game:

Note that here number of turns counts total number of turns, including failed. If we consider only succeeded turns, we’ll get a biased number - only ‘luckier’ puzzles will be counted.
GPT Terra 5.6
| n_games | n_turns | n_tokens | pct_solved | n_sessions | n_warnings |
|---|---|---|---|---|---|
| 1 | 12.15 | 15113.75 | 100.00 | 20 | 0.00 |
| 2 | 10.05 | 11651.85 | 100.00 | 20 | 0.00 |
| 4 | 10.28 | 9319.00 | 100.00 | 10 | 0.00 |
| 8 | 11.66 | 6809.83 | 100.00 | 10 | 0.01 |
| 12 | 18.16 | 5234.13 | 100.00 | 10 | 0.00 |
| 16 | 14.99 | 5729.38 | 100.00 | 10 | 0.01 |
| 20 | 22.33 | 5859.34 | 100.00 | 10 | 0.01 |
| 24 | 19.95 | 5126.12 | 100.00 | 10 | 0.01 |
Terra is less token-hungry even at max reasoning. It manages to solve all puzzles consistently, yet number of turns shows the same pattern.

GPT Sol 6.1
| n_games | n_turns | n_tokens | pct_solved | n_sessions | n_warnings |
|---|---|---|---|---|---|
| 1 | 9.10 | 12637.10 | 100.00 | 10 | 0.00 |
| 20 | 8.30 | 7318.11 | 100.00 | 10 | 0.00 |
| 32 | 8.56 | 5834.64 | 100.00 | 5 | 0.02 |
| 64 | 8.98 | 5195.62 | 100.00 | 3 | 0.01 |
This model is one of the best at the moment of writing. The numbers confirm it - even at more extreme 32/64 simul sessions the model can achieve much better results - under 9 turns average. Running with n_games=128, however, failed due to context window being exhausted. If context requirements were the same as for 64 games on average, we should get 5200 x 128 < 700k tokens, which is well within the 1M limit. This suggest that model gets less efficient at longer context.
Conclusion
The most interesting observation is that combining multiple tasks of certain complexity into same context window impacts performance and thus can be used to understand long-context performance.