The Elo rating calculator above does the two things the Elo system consists of. It converts a rating difference into an expected score, and it converts the gap between the expected score and the actual result into a rating change, scaled by a K-factor. Enter two ratings, a result and a K-factor and you get both players' new ratings.
Arb Digital publishes this because Elo has escaped chess and turned up everywhere — football models, esports ladders, tennis rankings, matchmaking systems and internal league tables. The arithmetic is identical wherever it appears; what changes between implementations is the K-factor, the treatment of draws and the starting rating, and those choices matter more than most users realise.
What This Elo Rating Calculator Does
It calculates the expected score for each player from the rating difference alone, applies the result, and returns updated ratings for both. It also shows what each of the three possible results would have done to player A's rating, so you can see the whole outcome space rather than one branch of it. If you set a number of identical games, it recalculates the expected score after each one rather than multiplying a single change, which is the correct way to iterate a rating.
Elo is a self-correcting rating system rather than a scoring statistic, which makes it a different animal from most of our sports tools. The winning percentage calculator and the earned run average calculator summarise what happened; Elo estimates relative strength and updates it. For working between probabilities and betting odds, the odds probability converter and the implied probability calculator handle that conversion directly.
How to Use It
- Enter both ratings. They must come from the same rating pool; a figure from one federation or platform does not translate to another.
- Choose the result from player A's point of view. Win, draw or loss. Player B's result is the mirror image and the calculator handles it.
- Set the K-factor that the rating system in question uses. This is the single biggest choice you make here.
- Optionally repeat the game. Setting several identical games shows how a rating converges rather than moving in equal steps.
- Read the bars to compare the win, draw and loss outcomes from the same starting position.
The Formula
The expected score for player A is EA = 1 ÷ (1 + 10(RB − RA) ÷ 400). The rating update is R'A = RA + K × (SA − EA), where S is the actual score: 1 for a win, 0.5 for a draw, 0 for a loss. Player B's expected score is 1 minus player A's, and their update uses the same equation.
The 400 in the exponent is what defines the scale. It means a 400-point advantage corresponds to an expected score of about 0.909, and a 200-point advantage to about 0.760. Nothing physical fixes that number — it is a scaling choice that determines what a rating point means, which is why ratings from systems with different scales cannot be compared even when both are called Elo.
A worked example makes the mechanics concrete. Player A on 1600 against player B on 1800 has an expected score of 1 ÷ (1 + 100.5) = 0.240. With a K-factor of 20, a draw gives A a change of 20 × (0.5 − 0.240) = +5.2 points, a win gives +15.2, and a loss gives −4.8. The underdog gains more from a draw than they lose from a defeat, which is the system working exactly as designed.
Choosing a K-Factor
K sets the maximum a single game can move a rating and therefore the trade-off between responsiveness and stability. A high K tracks a changing player quickly but makes ratings noisy; a low K is stable but slow to catch up with someone improving fast.
Real systems vary K by circumstance rather than fixing it. FIDE's rating regulations set K = 40 for a player new to the rating list until they have completed events totalling at least 30 games, K = 20 while a player's rating remains under 2400, and K = 10 once a published rating has reached 2400. The logic is that new players need to find their level quickly and established players at the top should not swing on one result.
National systems take different approaches to the same problem. The US Chess rating system, documented in Glickman and Doan's published description of the US Chess rating system, derives its equivalent of K from an effective-games count rather than fixed rating bands, so a player's volatility depends on how much they have played. Different mathematics, same underlying idea: uncertainty about a player should govern how far a result moves them.
Why the Underdog Cannot Lose Much
The asymmetry in the worked example above is the most misunderstood feature of Elo, and it is a feature rather than a flaw. Because the rating change is K times the gap between actual and expected score, a heavily favoured player who wins gains almost nothing — they did what was expected. The same player losing drops a lot, because the result was far from expectation.
This produces a real strategic consequence in any pool where players can choose opponents. Playing far weaker opponents is close to pointless for rating purposes: winning gains a fraction of a point and losing is catastrophic. Playing stronger opponents is cheap in the other direction. Rating systems that let players pick their fields have to think hard about this, and it is the reason many sports ladders restrict how far apart matched opponents can be.
Expected Score Is Not Win Probability
Elo's expected score is a score, not a probability, and the difference matters wherever draws exist. An expected score of 0.75 could describe a player who wins three quarters of their games and loses the rest, or one who wins half and draws the other half. Both produce the same score and the same rating change, and Elo has no way to distinguish them.
In sports with no draws the two coincide and expected score can be read as a win probability directly. In chess, football or any draw-permitting sport it cannot, and models that need genuine win, draw and loss probabilities have to layer an extra assumption on top of Elo about how the score splits. Anyone converting an Elo expectation into betting odds should be alert to this; the odds probability converter will happily convert a number that was never a probability to begin with.
Zero-Sum Ratings and Pool Drift
In the classic formulation, what one player gains the other loses exactly, so the total rating in a pool stays constant. That property is why Elo is stable, and also why comparing ratings across eras or across pools is treacherous. If a pool takes in new players at a fixed starting rating and they lose points before leaving, rating slowly leaks into the pool and average ratings drift upward over time. If the opposite happens, they deflate.
The practical rule is simple: a rating means something relative to the pool that produced it and the period it was earned in. A rating from a small local league and one from a large international pool can be numerically identical and describe very different strength. Real federations spend considerable effort on this problem, which is why their published systems include machinery that a plain Elo update does not.
Arb Digital builds fast, dependency-free interactive tools that rank on their own merits and earn links. Browse the free library, or tell us what you want built.
Browse All Free Tools Talk to Arb DigitalCommon Mistakes to Avoid
- Comparing ratings across pools. A number from one federation, platform or league says nothing about a number from another, even when both are called Elo.
- Using one K-factor for everyone. Published systems vary K by experience or rating band precisely because uncertainty differs between players.
- Reading expected score as win probability. Where draws are possible, the two are different quantities and cannot be substituted.
- Multiplying a single rating change by the number of games. The expected score changes after every result, so ratings must be iterated, not scaled.
- Assuming a rating is a fixed measure of skill. It is an estimate that only has meaning relative to its pool and the moment it was calculated.
Related Free Tools From Arb Digital
For records and results, see the winning percentage calculator and the earned run average calculator. For probability work, the odds probability converter and the implied probability calculator. For statistical comparison generally, the z-score calculator and the average percentage calculator. Everything else is in the free online tools hub.
Frequently Asked Questions
The expected score for a player is 1 divided by 1 plus 10 raised to the power of the opponent's rating minus the player's rating, all over 400. The new rating is the old rating plus the K-factor multiplied by the difference between the actual score and the expected score.
It is the development coefficient that sets how far one game can move a rating. FIDE uses 40 for players new to the rating list until they have completed at least 30 games, 20 while a rating remains under 2400, and 10 once a published rating reaches 2400.
An expected score of about 0.76 for the stronger player. A 400-point gap corresponds to roughly 0.91. The 400 in the formula is a scaling choice that defines what a rating point is worth, not a physical constant.
Only in sports without draws. Where draws exist, an expected score of 0.75 could come from winning three games in four or from winning half and drawing half, and Elo cannot distinguish between those two.
Because the rating change is proportional to the gap between the actual and the expected result. A heavy favourite who wins has done what the ratings already predicted, so there is very little information in the result and very little movement.
No. Elo is a relative system, so a rating only means something within the pool that produced it. Two pools with different entry ratings, K-factors and player populations can produce identical numbers that describe very different strength.
They must be iterated. After each result the ratings move, which changes the expected score for the next game. Multiplying a single game's rating change by the number of games overstates the total, which is why this calculator recalculates each time.
This tool implements the published Elo arithmetic for general reference. Official ratings are produced by the relevant federation or platform under its own published regulations, which may include provisions this simplified calculation does not.