← Blog
Articles

Building Arena Stats That Tell the Whole Story

Arena Assistant30/7/20268 min de lectura

Most stat sites can tell you what ended up in a winning inventory. We want Arena Assistant to answer a harder question:

Did this choice help, or does it only look strong because of when and where players were able to choose it?

That distinction is the focus of our current statistics work. We have been rebuilding the foundation beneath item and augment recommendations so that the numbers describe decisions, not just final screenshots of successful games.

This is the first entry in Insignia, a series where we open the workshop and explain the systems, research, and tradeoffs behind Arena Assistant.

The invisible bias inside raw stats

Raw averages are useful, but they are not neutral.

Imagine that Void Staff appears mostly as a late purchase. A player has to survive enough rounds, earn enough gold, and still have room in the build before buying it. Those conditions already describe a relatively successful run.

If we only inspect final inventories, the item inherits some of that success. Its average placement may look excellent even when the item itself added little at that moment.

The same problem appears in several forms:

  • Survivorship bias: teams that live longer can buy more items, so late purchases naturally come from stronger runs.
  • Purchase timing: the same item can have a different job as the first legendary than as the fourth.
  • Augment timing: an augment picked early has more rounds to create value than one picked late.
  • Rarity context: a gold augment should be compared with the other gold choices available in that decision, not with every augment in the game.
  • Popularity pressure: a common choice can look safe because it has volume, while a situational choice can look extreme because it has a tiny sample.

Raw placement, first-place rate, top-half rate, and pick rate all remain valuable. The mistake is asking one of them to carry the entire explanation.

Reconstructing the decision, not the final inventory

The first major step is preserving order.

Arena Assistant now derives an ordered item-acquisition history from match timeline data. We identify the first distinct legendary purchase, then the second, third, and fourth. Boots, consumables, prismatics, and non-shop rewards are classified through the item catalog instead of being mixed into the legendary sequence.

That produces a much more useful comparison:

Question Coarse comparison Our direction
Is this item strong? Every final inventory containing the item Same champion, same legendary purchase number
Is this augment strong? Every game containing the augment Same champion, same pick slot, same rarity
Is the result stable? One displayed average Sample size, uncertainty, and choice share
Is a rare option secretly optimal? Sort by raw placement Require enough evidence before promoting it

For augments, Riot's match data preserves pick order across up to six slots. That lets us compare Phenomenal Evil with alternatives chosen by the same champion, at the same point in the run, and within the same rarity.

The goal is simple: compare like with like.

Placement Added

The working metric behind this research is Placement Added, or PA.

WORKING METRIC PA
Placement Added = Expected placementfor comparable situations Observed placementwith this choice

Because a lower Arena placement is better, positive PA means the choice performed better than its conditioned baseline. A PA of +0.20 means the observed result was about 0.20 places better than the other choices in that same comparison cell.

COMPARE LIKE WITH LIKE Fair comparison cells
Item
ChampionLegendary purchase number
Augment
ChampionAugment pick slotRarity

We use a leave-one-out baseline, so the choice being measured does not inflate the expectation it is compared against.

This removes several large distortions at once. Champion strength is held constant. Purchase depth is held constant. Augment timing and rarity are held constant.

It does not turn observational match data into a laboratory experiment. PA means adjusted for the situation we can observe, not proven causal. Unknown factors can still exist inside a comparison cell, and we intend to keep that limitation visible.

What the research showed

The research confirmed that purchase order changes the story.

Late-buy items often had excellent raw averages but became far more ordinary after comparing them with other purchases at the same depth. Void Staff, for example, still showed real value when bought early, while much of its late-purchase advantage disappeared after conditioning on survival.

Guardian Angel remained strong across several positions, but the detailed view showed that its best relative performance did not simply come from appearing in completed builds.

We also found the opposite pattern: choices that looked mediocre in raw placement could become reasonable once compared with the actual situation in which they were selected. That is especially important for narrow counters and adaptations. A situational anti-shield or defensive purchase should not be judged as though it were competing to be the universal first item.

On LeBlanc, for example, Zhonya's Hourglass became more impressive when evaluated in its actual purchase position. That context matters: the item was not merely riding the strength of completed builds.

On the augment side, timing and rarity changed rankings enough to expose both underrated options and popular choices whose raw numbers were doing too much of the storytelling.

These are exactly the details a serious stats reader wants to inspect.

Confidence is part of the metric

A fine-grained model creates smaller samples, so uncertainty cannot be an afterthought.

Every conditioned result carries:

  • the number of observed choices;
  • its share of the comparison cell;
  • a confidence interval around Placement Added;
  • whether it passes the reliability gate;
  • whether the choice is so universal that a useful counterfactual barely exists.

We tested a flat sample requirement against a confidence-based gate. The confidence approach kept substantially more useful item and augment cells while still rejecting noisy results.

That matters because a sample of 400 tightly clustered placements can be more informative than a larger but highly variable sample. The interface should communicate that difference instead of hiding everything behind a single tier letter.

When evidence is insufficient, the honest answer is not a dramatic ranking. It is not enough data yet. When it passes the full contract, the interface can identify it as reliable evidence.

What this foundation unlocks

This work is still in progress, but it opens the door to a much richer statistics experience:

  1. Order-aware build paths that explain why an item belongs in a specific purchase slot.
  2. Timing-aware augment analysis across pick number and rarity.
  3. Raw and adjusted metrics side by side, so advanced users can inspect the difference.
  4. Confidence-aware recommendations that separate a strong signal from an exciting small sample.
  5. Better situational-item discovery without turning rare counter purchases into universal advice.
  6. Patch, region, champion, and build-style slices built on the same statistical contract.
  7. Bundle and interaction analysis for item pairs, augment combinations, and complete build directions.

The same foundation can improve the champion pages, overlays, build paths, and the recommendations inside our Team Builder. More importantly, every surface can consume one shared evidence model instead of inventing its own interpretation of the data.

What remains unsolved

Good statistics should include their boundaries.

  • Match timeline coverage is high but not complete, and availability may not be perfectly random.
  • Riot does not expose the exact three-augment offer set, so rarity and pick slot are the strongest observable approximation.
  • Choice bundles can still confound one another. Two options frequently selected together need a deeper interaction model.
  • Very new patches may require a broader evidence window while enough current-patch games arrive.
  • Rare adaptations are often the most interesting findings and the least certain.

We are building provenance, drift checks, and reliability gates around these limitations. As the model improves, older recommendations can be recalculated from the same source data instead of being patched by hand.

For the people who enjoy looking closer

Some players only want a quick answer. Others want to know the sample, the purchase slot, the confidence interval, the alternative pool, and why a ranking moved.

We are building for both.

Arena Assistant will continue to provide clear recommendations, but the detailed evidence underneath them is becoming deeper, more transparent, and more honest. This is not a finished victory lap. It is a look at the statistical foundation we are laying now, and at the analysis it will make possible next.

Better stats are not just more numbers. They are fairer comparisons.

Recibe recomendaciones personalizadas en seleccion de campeones

  • Sugerencias de aumentos + objetos en vivo basadas en TUS ofertas de aumentos
  • Avisos de counter-pick para el equipo enemigo que enfrentas
  • Score de sinergia de duo mientras tu y tu companero hacen draft