# Solving Riemann: the whole hand of Poker 51 as one game

**Riemann v0.6 · working note · 1 October 2026** · Jeffery Lyn Huckstead / Cerebral Graphix · CC BY 4.0
Unpublished working note; no DOI. RH STATUS: OPEN. "Riemann" is the dealer's name, not a claim: nothing here uses zeta data or bears on the Riemann hypothesis.
Status labels as on the Rule Card: COMPUTED (exact or converged for a stated model), EVID (evidence, not a theorem), OPEN. **Nothing in this note is PROVED.**

## 1. Summary

Until v0.5 the Poker 51 dealer, Riemann, bet by a published read-and-threshold policy: estimate its equity from a belief table, narrow your range by the size of your bet, and compare with the price of a call. Since v0.6 Riemann bets from a **solved strategy**, one published table per game (standard, Wild Ⅰ, the bug). The rules you play by are unchanged.

| Game | Classes · your draw states | Value of the game | What the v0.6 table guarantees | A perfect player gains at most | The same against v0.5 | Four modelled types, fresh real-card deals: v0.5 → v0.6 |
|---|---|---|---|---|---|---|
| Standard | 19 · 46 | +0.015 | −0.099 | 0.115 | 3.055 | +0.554 → +1.281 |
| Wild Ⅰ | 28 · 67 | +0.005 | −0.108 | 0.114 | 3.009 | +0.612 → +1.169 |
| The bug | 36 · 81 | +0.031 | −0.087 | 0.118 | 3.030 | +0.362 → +0.954 |

Values are Riemann's average jbits per hand (ante 1 jbit each).

- **COMPUTED, for the abstraction.** The whole hand is nearly fair with perfect play on both sides (first column of values). The v0.6 table guarantees the value in the next column against *every* strategy the abstraction allows; a script on the site recomputes it from the published table (`guarantee-v0-6.mjs`, receipt `GUARANTEE_v0_6.json`). So a player who knows the table and plays perfectly against it gains about 0.1 jbits per hand. Against the v0.5 policy the same player gained about 3 jbits per hand.
- **EVID, against modelled players.** On fresh real-card deals against four player types played by a language model, v0.6 wins 2.3× (standard), 1.9× (Wild Ⅰ) and 2.6× (the bug) what v0.5 won, and more against each of the four types in every game.
- **Not claimed.** An equilibrium, optimality or any return for the real card game, or against real players.

## 2. The game that was solved

Order of play (Poker 51 v0.5): each side antes 1 jbit and is dealt five cards; you choose which cards to keep, which announces your draw count k, and you act first before the draw; Riemann answers knowing k; both draw, Riemann by the published House rule; after the draw you act first again, both sides knowing k and Riemann's draw count j; showdown. Both betting rounds use the same tree: you check, or bet 1, 2, 4 or 8 jbits; checked to, Riemann checks or bets 2; one raise and one re-raise by the same amount.

The abstraction:

- **Dealt hands** are grouped into classes: made hands by type, three of a kind (low, high), two pair by top pair, four-card flush and straight draws, pairs by rank, and no pair by top card. The standard game has 19. A variant's class also records whether the hand holds Ⅰ and the House rule's draw count for it, since Riemann always keeps Ⅰ: 28 classes for Wild Ⅰ and 36 for the bug. Class frequencies are exact (every five-card hand of the 51- or 52-card deck is counted).
- **Your draw** is part of the game. For each class you choose among the House-rule draw, standing pat, keeping a kicker with a pair or three of a kind, and keeping two high cards with no pair. That gives 46, 67 and 81 draw states. Riemann's draw is the House rule, taken as given.
- **Final hands** fall into 40 strength buckets, the percentiles of House-rule final hands under the game's own ranking. At showdown the higher bucket wins and equal buckets split. Card removal between the two hands is ignored.
- **Transitions** from a class and a draw to a final bucket are sampled (6,000 dealt hands per class in the standard game, 4,000 in the variants; seed 777).
- **Information.** Perfect recall: after the draw you know your draw state and your bucket, and Riemann knows its class and its bucket.
- **Stacks are deep.** There are no all-ins in the solved game.

## 3. The solver

CFR+ (counterfactual regret minimisation with non-negative regrets and linear averaging), alternating updates, vectorised over hand types. Convergence is reported as NashConv: the sum of what each side could gain by deviating.

| Game | Iterations | Value of the game for Riemann | NashConv |
|---|---|---|---|
| Standard | 3,000 | between +0.0147 and +0.0157 | 0.00096 |
| Wild Ⅰ | 3,000 | between +0.0054 and +0.0061 | 0.00071 |
| The bug | 3,000 | between +0.0307 and +0.0314 | 0.00070 |

A first version of the solver stalled at a NashConv of 0.13. It let each side forget its own dealt hand after the draw (imperfect recall), which CFR does not handle. It was diagnosed on a small game and replaced by the perfect-recall version above; nothing from it is used.

## 4. The player population

The solve needs a model of how people actually play, to know which mistakes are worth exploiting. Here that model is a language model: TypeSafe System One, `jev-1.13.0`, asked as a judge. Each spot was described in poker words (the hand, the pot, the action so far) and the model chose among the legal actions with probabilities. All arithmetic stayed in code. Four player types were prompted: **default** (best play as it sees it), **cautious**, **aggressive** and **calling station**.

- **Before the draw:** which draw to make, the opening action for each draw, and the answers to Riemann's bet or raise. Ten dealt hands per class per type: 760 requests for the standard game, 672 for Wild Ⅰ, 864 for the bug.
- **After the draw:** two representative hands for each (draw count × strength bucket) cell, against each of Riemann's draw counts, at pots of 2, 6, 18 and 50 jbits: 7,580, 3,800 and 3,820 requests.
- **The types behave as named.** In the standard game the aggressive type stands pat as a bluff on 49–60% of no-pair hands and checks a weak hand 11% of the time; the calling station folds a medium hand to an 8-jbit raise 8% of the time; the cautious type checks weak hands 87% of the time.

The population used in the solve is the average of the four types. **These are a model's answers, not people** (EVID). In all, the Poker study asked 45,113 questions (129.2 million input tokens).

## 5. Restricted Nash response: how much to trust the model

An equilibrium strategy cannot be beaten, but it does not punish mistakes. A best response to the model wins the most from it, and can be beaten badly by anyone who plays differently. A **restricted Nash response** sits between: with probability p you are the population model, and otherwise a free player who knows Riemann's strategy and plays perfectly against it. Riemann cannot tell which. p = 0 is the equilibrium; p = 1 is the best response to the model.

Each row below is one solve (2,000 CFR+ iterations). "A perfect player takes" is the value of the game minus what the strategy guarantees, in the abstraction (COMPUTED). The other columns are Riemann's expected jbits per hand on fresh real-card deals against each modelled type (EVID).

**Standard** (600 fresh deals per type)

| Riemann | a perfect player takes | vs default | vs cautious | vs aggressive | vs calling station | average |
|---|---|---|---|---|---|---|
| v0.5 policy (read and threshold) | 3.055 | +0.603 | +0.268 | +0.541 | +0.806 | +0.554 |
| equilibrium of the abstraction | 0.000 | +0.645 | +0.468 | +0.869 | +0.826 | +0.702 |
| restricted Nash response, p = 0.25 | 0.024 | +0.954 | +0.550 | +1.875 | +1.044 | +1.106 |
| **restricted Nash response, p = 0.5 (published as v0.6)** | **0.114** | **+1.026** | **+0.572** | **+2.426** | **+1.097** | **+1.281** |
| restricted Nash response, p = 0.75 | 0.699 | +1.184 | +0.654 | +3.396 | +1.208 | +1.611 |
| restricted Nash response, p = 0.9 | 2.031 | +1.328 | +0.728 | +4.391 | +1.193 | +1.910 |

**Wild Ⅰ** (300 fresh deals per type)

| Riemann | a perfect player takes | vs default | vs cautious | vs aggressive | vs calling station | average |
|---|---|---|---|---|---|---|
| v0.5 policy (read and threshold) | 3.009 | +0.660 | +0.235 | +0.824 | +0.727 | +0.612 |
| equilibrium of the abstraction | 0.000 | +0.518 | +0.425 | +0.782 | +0.548 | +0.568 |
| restricted Nash response, p = 0.25 | 0.017 | +0.864 | +0.528 | +1.729 | +0.778 | +0.975 |
| **restricted Nash response, p = 0.5 (published as v0.6)** | **0.113** | **+0.993** | **+0.576** | **+2.250** | **+0.859** | **+1.169** |
| restricted Nash response, p = 0.75 | 0.576 | +1.183 | +0.665 | +3.093 | +0.968 | +1.477 |

**The bug** (300 fresh deals per type)

| Riemann | a perfect player takes | vs default | vs cautious | vs aggressive | vs calling station | average |
|---|---|---|---|---|---|---|
| v0.5 policy (read and threshold) | 3.030 | +0.449 | +0.191 | +0.346 | +0.463 | +0.362 |
| equilibrium of the abstraction | 0.000 | +0.445 | +0.409 | +0.480 | +0.482 | +0.454 |
| restricted Nash response, p = 0.25 | 0.024 | +0.705 | +0.490 | +1.141 | +0.593 | +0.732 |
| **restricted Nash response, p = 0.5 (published as v0.6)** | **0.118** | **+0.845** | **+0.532** | **+1.684** | **+0.754** | **+0.954** |
| restricted Nash response, p = 0.75 | 0.771 | +1.130 | +0.736 | +2.793 | +0.835 | +1.373 |

**Why p = 0.5.** The frontier bends sharply, and in the same place in all three games. In the standard game, going from the equilibrium to p = 0.5 raises the average win by 0.58 jbits per hand while the most a perfect player can take rises by 0.11. Going on to p = 0.9 adds 0.63 to the win and 1.92 to what a perfect player takes. Because the table is published, that last figure is the price of publishing, so it is kept near 0.1. p = 0.25 is the near-equilibrium alternative (a perfect player takes about 0.02). The choice of p is a design choice.

The published tables round each probability to a byte. The figures of section 1 are recomputed from those bytes, which is why they read 0.115, 0.114 and 0.118 where the rows above read 0.114, 0.113 and 0.118.

## 6. Validation on real cards

The deals in section 5 were not used to build the model or the solves: 600 whole hands per player type for the standard game (seed 9191) and 300 per type for each variant (seed 9292), stratified by the House-rule draw count of the dealt hand and weighted back to exact frequencies. The model played each of them in full: its draw, its opening action, and its answers. Riemann's result is the exact expectation over its own mixed strategy and the model's stated probabilities, with the real cards and the real showdown.

- Scored inside the bucket game instead, the candidates come out in the same order, except that in the bug the v0.5 policy and the equilibrium trade places. The averages themselves differ from the real-card ones by up to about 0.2 jbits per hand: buckets hide differences between hands.
- The published, hash-mixed standard module scores +1.297 on the same deals, against +1.281 for the mixed strategy it encodes.
- The validation is out of sample in deals, not in player types: the same four prompts built the population. What protects Riemann if real players differ is the guarantee of section 1, not this table.

## 7. What is published

- `area-51/poker-51/poker51-riemann-v0-6.js`: the lookup. Before the draw it reads the class of Riemann's dealt hand, your announced draw count and the node of the betting tree; after the draw it adds Riemann's strength bucket, its own draw count and the betting before the draw.
- `area-51/poker-51/riemann-solved-GAME-v0-6-data.js`: one table per game. Each probability is a byte, with the bytes of a decision summing to 255 and the last action implied.

| Game | Table bytes | SHA-256 of the table bytes |
|---|---|---|
| Standard | 799,330 | `a0cf62457f7c62e5dfdecb88cc815b0bd4e072ed1b8955bf6d283b5970a6c9dc` |
| Wild Ⅰ | 1,177,960 | `78f4ef23c719d859f7b581bd5c7794d355e3128e44432d8726adc903a0078b69` |
| The bug | 1,514,520 | `d3b76fc0ee774de75da310f1ca7bad80a6116d8a05f1bd23ac4acfe5c4e3a0cf` |

- **The mixed strategy is made replayable by a public hash.** FNV-1a over Riemann's sorted cards and the spot gives a number u in [0, 1), which picks the action from the table's probabilities: the device of the v0.2 bluff rule. The same cards in the same spot always give the same action; across hands the frequencies follow the table (checked within 2 points at four kinds of spot in each game).
- **Where the table has no entry.** An all-in for an amount other than 1, 2, 4 or 8 jbits is read at the nearest of those sizes (in ratio). A five-card draw is read as a four-card draw. A solved raise against a player who is all in becomes a call, and Riemann bets only while you still have jbits to answer, as in v0.5.
- `area-51/poker-51/poker51-table-v0-6.js`: the v0.5 hand with these decisions. Dealing, the House draw rule, the bets you may make, settlement and receipts are unchanged; a hand saved under v0.2 to v0.5 finishes under its own module.
- `tests/v4_1/verify-poker51-v0-6.mjs`: each table is intact; over 9,000 seeded hands every decision is the table's, jbits are conserved and receipts are unique; hands replay byte for byte; and whenever Riemann makes the same choice, the v0.6 hand equals the v0.5 hand.

## 8. What the solve says about the game

- **The game is nearly fair.** With perfect play on both sides a hand is worth about +0.015 jbits to Riemann in the standard game (+0.005 in Wild Ⅰ, +0.031 in the bug): the value of acting second. Everything Riemann earns beyond that comes from the other side's mistakes.
- **Draw counts are hiding places.** At equilibrium in the standard game you keep two high cards with no pair and draw three, hiding among the pairs; with a small pair (2–5) you keep a kicker and draw two about two times in three, hiding among three of a kind; you stand pat as a bluff on 11.6% of low no-pair hands. A stand-pat is then a made hand only 39% of the time. The v0.5 belief tables assumed every draw count was honest, and that is the root of its 3-jbit leak.
- **A signal is only as believable as its honest version is common.** An honest stand-pat happens on 0.78% of deals (exact count over all 2,349,060 hands), so if just 1% of weak hands stand pat as a bluff, a stand-pat is a bluff 53% of the time. A common honest signal, such as drawing three, is barely moved by the same bluff rate.
- **Minimum defence frequency.** Facing a bet B into a pot P, a defender who folds more than B/(P + B) of the time hands any bluff a profit. The v0.5 policy continued against an 8-jbit bet into a 2-jbit pot with 44% of its hands; the solved last round continues with about 20%, which is P/(P + B).
- **A rules observation.** Checked to, Riemann can only bet 2 jbits, while you choose among four sizes. In the solved last round alone (honest draws on both sides) that costs Riemann at large pots: the round is worth +0.027, −0.022 and −0.279 jbits to it at pots of 2, 6 and 18. This is a property of the v0.5 rules, not of any policy, and is left as it is.

## 9. How the study got here

1. **A decision map.** 1,344 of Riemann's v0.5 decisions were put to the model as a second opinion, and 2,600 dealt hands were played by it as the player, including draw bluffs. Finding: a perfect player who stands pat on a weak hand gained 4.67 jbits per hand in the last round against v0.5.
2. **The draw.** 1,500 dealt hands: which cards a typical modelled player keeps. Announced draw counts meant something quite different from what v0.5 assumed (a two-card draw was one pair 94% of the time, not three of a kind).
3. **The last betting round as a game.** 225 solved games (25 draw-count pairs × 3 pots × 3 beliefs). A strategy solved for a mixed belief was robust under each belief; one solved for honest draws alone was not.
4. **The whole hand,** sections 2 to 6: the standard game, then the two variants.

Two hand-tuned policies from steps 1 and 2 (bluff-raise only a 1-jbit bet; doubt a stand-pat more) closed part of the leak but stayed exploitable by about 1 jbit per hand, and were set aside for the solve.

## 10. Status

| Statement | Status | Where |
|---|---|---|
| Class frequencies of dealt hands in each game | COMPUTED | every five-card hand counted; `fullgame.mjs`, `vabs.mjs` |
| Value of each abstract game, with NashConv below 0.001 jbits per hand | COMPUTED | CFR+; section 3 |
| What each published v0.6 table guarantees in its abstraction | COMPUTED | `guarantee-v0-6.mjs`, `GUARANTEE_v0_6.json`: the abstraction is rebuilt from its seed and the best response recomputed from the published bytes |
| What a perfect player takes from the v0.5 policy in the abstraction (about 3 jbits per hand) | COMPUTED | the v0.5 policy mapped into the abstraction by sampled hands of every class; study scripts |
| The v0.6 table module plays the published tables, replays byte for byte and keeps the v0.5 rules | COMPUTED | `tests/v4_1/verify-poker51-v0-6.mjs`; finite tests, not a formal proof |
| v0.6 wins more than v0.5 from each of four modelled player types on fresh real-card deals | EVID | the players are a language model's answers; section 6 |
| Returns against real players | OPEN | not measured |
| An equilibrium, optimality or a return for the real card game | not claimed | the solve is of an abstraction, with the House draw rule taken as given |

## 11. Limits

- **The players are a language model, not people.** The four types were also the ones validated against.
- **The abstraction:** classes and 40 buckets stand in for exact cards; no card removal between the two hands; deep stacks. A short-stacked player at the real table goes all in, which the solved game does not model.
- **The guarantee is a statement about the abstraction.** A player using exact cards, card removal or stack sizes may take more than 0.1 jbits per hand from the real table. How much more is not measured.
- **Sampling.** Transitions and bucket edges are sampled with fixed seeds; the validation deals are finite samples.
- **Riemann's draw is not solved.** It is the published House rule, as before.

## 12. Reproduce

From the site root, with Node 20 or later and no packages:

```sh
node tests/v4_1/verify-poker51-v0-6.mjs                                  # the table module: 9 checks, 9,000 hands
node research/project-51/poker-51/riemann-v0-6/guarantee-v0-6.mjs        # the three guarantees, about 15 minutes
node research/project-51/poker-51/riemann-v0-6/fullgame.mjs eq.json      # the standard equilibrium (CFR+, 3,000 iterations)
```

In `research/project-51/poker-51/riemann-v0-6/`: `fullgame.mjs` (the standard abstraction and solver), `fullgame2.mjs` (restricted Nash response, best responses), `vabs.mjs` and `vsolve.mjs` (the variants), `popmodel.mjs` (the population model, from the model's answers), `packsite.mjs` (a solved strategy to a published table). The restricted Nash responses need the model's answers; those, the questions, the solved strategies and every score table are a data bundle prepared for deposit with Project 51 (no DOI yet).
