I had lunch with my dad outside Reading Terminal Market in Philadelphia. He plays Tigris & Euphrates seriously. He is markstephens404 on Board Game Arena, with 1,749 games, 892 wins and a 74% win rate. He has been inside the world top 15 at the game.
So I asked him what the strongest players actually do. He talked. I listened. Then I asked a different question, and it is the reason this post exists.
Could a machine find a new idea about this game?
Not by playing itself ten million times. That works, and other people do it well. I meant something narrower: take the published rule, treat it as mathematics, and see whether anything falls out that strong players do not already say to each other.
This post is that attempt. I should tell you the result before you spend your time.
I did not find a new strategy. I found four exact statements about the scoring rule. Three of them I have not seen written down anywhere, and one of them corrects a mistake that is easy to make and that I made myself while working. They are small, they are certain, and they change how I read a position. That is less than I hoped for and more than I expected.
The idea this post is built on
Before you spend actions chasing points in one color, count how many of your four spheres are level at the bottom. That count is your floor multiplicity. Call it w. It is 1, 2, 3 or 4, and it changes turn by turn.
Two things follow from w alone, and both are exact:
Every guide says to keep your spheres level. That is the same advice in every position. w tells you what being unlevel is costing you right now, and what a war is really worth before you start it.
I also found the edge, quite fast. Past that edge the math stops and guessing starts. This post marks that edge clearly, because the interesting failure was not in the arithmetic. It was in the labeling.
This post is about one rule in the board game Tigris & Euphrates. You do not need to know the game. You need four facts.
Most guides give you one sentence: keep your spheres level. The sentence is true. It is not enough. This post shows what the rule says. It also shows what the rule does not say.
Each new word appears in bold the first time you meet it. A short definition comes with it. You do not need to remember them all at once.
Each claim below has one label. The labels mean different things.
Rule The rulebook gives this result. It is true of the game.
Model A toy model gives this result. It is true of the model. It is not yet true of the game.
Hypothesis A guess about real play. Nobody has tested it here.
Your result is not one number. It is four numbers.
Write your four sphere totals together as a portfolio. Sort the portfolio from low to high. To compare two players, compare the two sorted portfolios one position at a time. The first difference decides. This comparison is leximin.
Now give the four sorted positions names. Position 1 is criterion 1. Position 2 is criterion 2. The rule reads criterion 1 first, then criterion 2, and so on.
Criterion 1 is your lowest sphere. This post also calls it your floor, because it is the level everything else sits above.
Two or more spheres can hold that same lowest total. The count of spheres at your floor is the floor multiplicity. It is 1, 2, 3 or 4. It does more work in this post than any other number.
Two facts come from this. Both are easy to miss.
First, the comparison does not stop at the floor. One player holds 5, 5, 8, 8. A second player holds 5, 6, 6, 7. They are not equal. The first player loses. Criterion 1 is 5 for both players. Criterion 2 is 5 against 6.
Second, the rule does not see color. It sorts your portfolio first, and the sort removes the labels. Color controls what you can get while you play. Color does not change the comparison.
This post uses a few short symbols. Each one stands for a phrase about the game. Nothing here is more than counting. Here is the whole list.
Rule You hold t treasures. You want the highest floor you can get. There is a formula for it.
Proposition 1 (the level treasures reach). Your floor goes up to λ*.
λ* = max { λ : Σi max(0, λ − xi) ≤ t }
In words: pick a level λ. For each sphere, work out how many points it needs to reach that level, and count zero for a sphere that is already there. Add up those numbers. That total is the cost of the level. λ* is the highest level with a cost you can pay from your t treasures.
The sum Σi max(0, λ − xi) is the number of points you need to bring every sphere up to λ. This number goes up as λ goes up. So the levels you can pay for make one unbroken range. λ* is the top of that range. You can pay for λ*. You cannot pay for λ* + 1. ■
This method is called water-filling. You can do it at the table. Pour water into four columns. Stop when you have no water left. The surface shows your floor.
Here is an example. Your spheres hold 3, 5, 9 and 5. You hold six treasures. Work out the cost of each level.
| Level | Points needed, sphere by sphere | Cost | Can you pay? |
|---|---|---|---|
| 5 | 2 + 0 + 0 + 0 | 2 | yes |
| 6 | 3 + 1 + 0 + 1 | 5 | yes |
| 7 | 4 + 2 + 0 + 2 | 8 | no |
Level 7 costs 8 treasures, and you hold 6. Level 6 costs 5, so you can pay for it. So λ* is 6. One treasure is left over. It lifts one sphere to 7.
Six treasures poured into a portfolio of 3, 5, 9, 5
The solid columns show points you already hold. The lighter part shows treasure. The surface stops at λ* = 6.
Greedy filling is not only a good method. It is the best method for the whole comparison.
Proposition 2 (greedy filling is best). Put each treasure into a sphere that is at the floor. This gives the portfolio g. No other way to place the same t treasures beats g.
Why it works, in plain terms. Water-filling makes your four totals as level as they can be. The rule reads the low end of the sorted portfolio first. So the most level portfolio you can reach is also the best one. The formal version follows.
Lemma. Water-filling gives the smallest portfolio in the majorization order — a standard way of comparing how evenly a set of numbers is spread. A portfolio is smaller in that order exactly when each of its k-lowest sums is larger. So g gives the largest sum of the k lowest criteria, for every k. This is the Hardy–Littlewood–Pólya result. The source code also tests it by brute force.
Now assume that some portfolio y beats g. Let k be the first criterion where they differ. Then y↑k is above g↑k, and every criterion below k is equal. Add the first k criteria. The sum for y is larger than the sum for g. This contradicts the lemma. ■
Four points can arrive in one sphere at once. This is what a war pays. This post calls it a windfall. Look at two cases.
In the first case, one sphere is alone at the floor. The four points lift your lowest sphere. Your floor goes up.
In the second case, two or more spheres hold the floor value. You put four points into one of them. Your floor does not move. The other sphere at the floor is still there.
It is easy to say that the four points did nothing. That is wrong. Here is the correct statement.
Proposition 3 (the price of one point of score). To lift your lowest sphere by 1, you must pay exactly w points.
Put one point into each of the w spheres at the floor. All of them go to m + 1. Every other sphere is already above m. So the new floor is m + 1. Now pay less than w points. At least one sphere stays at m. So the floor does not move. ■
So floor multiplicity is a price. One sphere at the floor: one point of score costs one point. Four spheres at the floor: one point of score costs four points. This is the most useful number in this post.
Proposition 4 (which criterion moves). Add b ≥ 1 points to one sphere at the floor. This improves exactly criterion w. Criteria 1 to w−1 do not change. Criterion w goes up.
The floor value m appears w times. So x↑1 to x↑w are all m, and x↑w+1 is above m. Add b to one copy of m. Now w−1 copies of m are left. One entry becomes m + b, which is above m. Every other entry does not change, and each one is already above m. Sort again. Positions 1 to w−1 still hold m. Position w now holds min(m + b, x↑w+1), and that value is above m. So the first difference is at position w. ■
Read Proposition 4 again. A single-sphere windfall never does nothing. The value of w is 1, 2, 3 or 4. So one criterion always moves. Floor multiplicity decides which criterion moves. It does not decide if one moves.
A trap. It is easy to score a portfolio with min(). Do this, and a windfall into a level floor measures as zero. You then say the windfall is worth nothing. But min() is only criterion 1. In this case criterion 1 is the one criterion that cannot move. The zero comes from your measurement. It does not come from the game.
This also separates two tools. A treasure lifts the floor level, because you can put treasures into several spheres. A windfall cannot lift a level floor, because all the points go into one sphere. The windfall moves criterion w instead.
Does the difference matter? That question needs an opponent. So it leaves the rules and enters a model. The next section describes that model in full. For the chart below you need one word. A run is one instance of the model. A run is not a game, and this post never mixes the two.
Does a “worthless” windfall change who wins?
Each run makes one portfolio and one opponent portfolio. It compares them with the rulebook comparison. It does this with the windfall and without it. The groups show how many spheres hold the floor value.
Model At floor multiplicity 2 the lowest sphere cannot move. The windfall still lifts the win rate from 32.1% to 76.4%. It changes the result in 44.3% of matches. This is not a small effect.
Everything below comes from one model. Here is what the model does.
Points arrive one at a time. Each point has a color. On a steered step, the point goes into a sphere that is at your floor. On an unsteered step, the color comes at random from the tile bag: 57 red, 36 blue, 30 green, 30 black. Your discipline is how often you steer.
The bag assumption is the largest leap in this model. It is not a rule. Tile counts are not point counts. You need a leader or a king to get a point. A king collects for a sphere with no leader. A revolt pays red at any size. A war pays several points at once. A monument pays again and again, and it uses no tiles. A trader takes the treasures. Every leader needs a temple next to it, so red tiles do structural work.
The model has no board. It has no kingdoms, no leaders, no monuments, no opponents and no hidden information. So a result that leans on the bag assumption is a guess about the game, and no more.
The same 36 points, with more or less discipline
Your lowest sphere against how often you steer a point into a sphere at your floor.
Model Discipline is worth about 3.8 points of lowest sphere. The first quarter of the discipline gives you 56% of that gain.
A metric that is easy to misname. It is easy to call 57.4% the share of raw points that become score. It is not that. It is the lowest sphere as a share of the perfect value. Only 14.3% of your raw points sit in your lowest sphere. Perfect play reaches 25%, because a quarter is the limit with four spheres. Call the number balance efficiency and it is useful. Call it a conversion rate and it is four times too large.
What one more point does to your lowest sphere
Move the slider. Rank 1 is your weakest sphere now. This shows the effect on the lowest sphere only. It is not the value of a point, because that needs an opponent.
Model With 20 points still to come, the four ranks are almost equal: 0.31, 0.30, 0.28 and 0.16. The order can still change, so a strong sphere keeps value. As the end comes near, the effect moves onto the floor.
You can read this as a rule for play: go wide early, go exact late. Do not. This chart sees criterion 1 only. It cannot see the other three criteria. Read it as one idea about one criterion in one model.
The tile bag is not equal. Red is 37% of it. Green and black are 19.6% each. So you can guess that the scarce colors set your floor.
Share of the bag against share of runs where that color is alone at the floor
In the toy model, with no discipline. Both bars are shares of all runs, so you can compare them directly.
Model Green or black is the only lowest sphere in 66.8% of all runs. Red does this 1.1% of the time.
A denominator that is easy to lose. Green and black are 78.9% of the runs that have one sphere alone at the floor. It is easy to quote that as “almost four in five”. But only 84.7% of runs have one sphere alone at the floor. So the correct figure for all runs is 66.8%. The simulation reports the conditional number first. That is why it escapes.
Now the important part. I ran the same measurement again with a different arrival rule. The result is mostly an echo of my assumption.
Discipline removes the effect as fast. At 35% discipline the figure falls to 36.6%. At 50% it falls to 31.3%.
Hypothesis So do not say that the game is decided in green and black. Say this instead. If points in real play arrive in proportion to tile counts, then scarce colors set the floor more often for a player with no discipline. I have not shown the first part. The section above lists the reasons it is probably false. There is a clean test: count the colors of the points scored in a few hundred recorded games. Until somebody does that test, color advice has no support here.
Two analyses look good from here. Both fail against the rules. I name them so that nobody builds them.
Do not count the bag. It looks possible to work out what your opponent holds. Page 9 of the rulebook stops this: “The number of tiles in the bag is hidden information; players cannot deliberately count the tiles in the bag.” It also says that tiles removed from the game “are kept facedown in the box and cannot be viewed by any player.” Tiles leave the game face down through hand replacement, revolts and wars. So you cannot know what is left. The rules also forbid the attempt.
Do not trust one war calculation. It is easy to pick one position and find the best target. That number changes with your own strength, with the tiles you commit, and with what you believe about the unseen tiles. It also leaves out the cost of the action, the value of both kingdoms, and the priest rule on page 10. Under that rule a temple with a treasure, or a temple next to another leader, stays on the board and pays nothing. One position is not a strategy.
Steering is not a resource you can save. The model finds that late discipline is worth about 0.58 of a point. The number is real inside the model. But the model gives both policies the same points and the same steering chances. The game does not work that way. Your control over color comes from leaders, position, monuments, treasures, and from what your opponents allow. Early actions build the board that lets you score late.
The scoring rule holds more than “keep your spheres level”. Most of the extra content sits in the later criteria, and short summaries drop them. What I cannot tell you is how this works on a real board. That needs a program with the board, the opponents, and strong self-play.
One Python file gives every number above. It uses numpy. It gives the same result every time, and it runs in about three seconds. Each function carries the label of the claim it can support.
Two functions carry the whole file. Everything else calls them.
tigris_sim.py
def leximin_key(pts):
"""The comparison key: victory points sorted ascending.
Rulebook p. 14 compares lowest spheres, then second-lowest, and so on.
That is exactly lexicographic order on the ascending-sorted vector.
"""
return np.sort(pts, axis=-1)
def leximin_cmp(a, b):
"""-1 if portfolio `a` loses to `b`, +1 if it beats it, 0 if tied.
Elementwise on the sorted keys, first difference decides.
"""
ka, kb = np.sort(a), np.sort(b)
for x, y in zip(ka, kb):
if x < y:
return -1
if x > y:
return 1
return 0The scoring rule of the game is one sort. Sorting ascending is “compare the lowest sphere, then the second lowest”, so the rulebook sentence and the line of code are the same instruction.
That sort is also where color stops existing. Four totals go in carrying red, blue, green and black. They come out as positions 1 to 4. No line after it can ask which sphere a number came from, which is why color cannot change the comparison however much it decides what you are able to collect.
leximin_cmp then walks the two keys together and returns at the first difference. Notice what it never does: it never reduces a portfolio to a single number. Nothing in the file calls min() on one. That is the trap from Floor multiplicity sets the price, kept out by construction rather than by care, because min() is only criterion 1 and criterion 1 is the one criterion a windfall into a level floor cannot move.
All four propositions have short proofs. All four also have a brute-force test, because a short proof can hide a case. The tests are simple loops, and each one decides whether the claim held by calling the nine lines above. Run the file and they report like this.
stdout
[RULE] P1 floor after t treasures: 0 counterexamples in 90000 cases
[RULE] P2 greedy filling is optimal: 0 counterexamples in 12005 cases
[RULE] P3 cost to lift the floor = w: 0 counterexamples in 4096 cases
[RULE] P4 windfall improves criterion w: 0 counterexamples in 24576 cases
worked example: [3, 5, 9, 5] + 6 treasures -> level 6 (spent 5, leftover 1), final [7, 6, 9, 6]The worked example on the last line is the 3, 5, 9, 5 portfolio from Treasures fill your portfolio like water. The run that prints it is the run that draws the figure, so the two cannot drift apart.
The simulation is tigris_sim.py. Its output is sim_results.json. A script writes the chart numbers into this page from that file. A check script fails if the two do not agree.
I checked the rules against the Z-Man rulebook. See page 9 for hidden information, page 10 for the priest rule, and page 14 for the winner.