Solo Queue Analytics

Next champion prediction

Lulu

Lulu

Predicted pick · 21% confidence
Logistic argmax, the most accurate of the six so far
Rift games
9
Win rate
78%
Avg KDA
6.72
AI playstyle read

How Jackson will probably play Lulu.

Model comparison

2

Logistic argmax

Shown above
Lulu
Lulu
21% confidence
Lulu
21%
Pyke
11%
Graves
11%
Jarvan IV
9%
Ornn
8%
1

Two-step sampler

Jarvan IV
Jarvan IV
Jungle · 13% confidence
Graves
15%
Jarvan IV
13%
Bel'Veth
11%
FiddleSticks
11%
Hecarim
11%
3

XGBoost

Lulu
Lulu
29% confidence
Lulu
29%
Jarvan IV
16%
Graves
14%
Pyke
14%
Karthus
5%
4

Decay frequency

Graves
Graves
28% confidence
Graves
28%
Jarvan IV
14%
Pyke
11%
Ornn
8%
Bel'Veth
8%
5

Uniform random

Yasuo
Yasuo
2% confidence

Draws at random, so there's no shortlist.

6

Markov transition

Lulu
Lulu
38% confidence
Lulu
38%
Bard
12%
Jarvan IV
12%
Karma
12%
Pyke
12%

Model accuracy

2

Logistic argmax

34% top-1 accuracy
95% CI 30–39%
Beats decay frequency, McNemar p = 0.001
Games scored
427
Hits
146
Last 20 games
20%

Logistic regression over the pool, auto-selected against a plain recency baseline.

3

XGBoost

34% top-1 accuracy
95% CI 30–39%
Beats decay frequency, McNemar p = 0.008
Games scored
427
Hits
145
Last 20 games
20%

Boosted trees on the same features. Data-hungry, so it usually trails on this little history.

4

Decay frequency

27% top-1 accuracy
95% CI 23–32%
Baseline: the bar the others must clear
Games scored
427
Hits
116
Last 20 games
0%

No learning: what's been played lately, older games fading out. The bar to beat.

6

Markov transition

26% top-1 accuracy
95% CI 23–31%
Not separable from decay frequency, McNemar p = 0.810
Games scored
427
Hits
113
Last 20 games
5%

Scores by what usually follows the champion played last. The only model reading game order.

1

Two-step sampler

15% top-1 accuracy
95% CI 12–19%
Loses to decay frequency, McNemar p = 0.000
Games scored
427
Hits
64
Last 20 games
0%

Samples a role, then a champion in it. The only model with randomness, trading accuracy for variety.

5

Uniform random

4% top-1 accuracy
95% CI 2–6%
Loses to decay frequency, McNemar p = 0.000
Games scored
427
Hits
15
Last 20 games
0%

A seeded random pick. The floor: beat this or you aren't predicting anything.

Recent scored predictions

Actually played Two-step samplerLogistic argmaxXGBoostDecay frequencyUniform randomMarkov transition
Pyke Graves Lulu Lulu Graves Karthus Lulu
Jarvan IV FiddleSticks Graves Lulu Graves Fizz Graves
Lulu Karthus Lulu (correct) Lulu (correct) Graves Talon Lulu (correct)
Seraphine Braum Graves Lulu Graves Ekko Bard
Taric FiddleSticks Lulu Lulu Graves Caitlyn Pyke
Bard FiddleSticks Graves Graves Graves Karthus Graves
Lulu FiddleSticks Graves Graves Graves Shaco Graves
Lulu Lillia Lulu (correct) Lulu (correct) Graves Seraphine Bard
Lulu Lillia Lulu (correct) Lulu (correct) Graves Blitzcrank Jarvan IV
Lulu Jarvan IV Lulu (correct) Lulu (correct) Graves Lillia Graves
Lulu Hecarim Graves Jarvan IV Graves Jayce Graves
Nautilus Graves Lulu Lulu Graves Hecarim Graves
Blitzcrank Lillia Graves Jarvan IV Graves Yasuo Jarvan IV
Lulu FiddleSticks Graves Graves Graves Graves Graves
Jarvan IV Garen Lulu Lulu Graves Bard Lulu
How this is tested

For every ranked game, each model is trained only on games before it, then its pick is written beside the champion actually played. No model can peek at the answer, and correctness is inherent in every row, so there's no separate verification to trust.

A few hundred games is not many, so each accuracy carries a 95% confidence interval, and the gap to the decay frequency baseline is tested with McNemar's test rather than eyeballed. McNemar's is the right test because every model predicts the same games: the samples are paired, so only the games where two models disagreed carry information. A four-point gap on this much data can easily be noise.

Why these six models

Everything has to beat the decay baseline. "What he's played lately" is free and has no parameters to overfit, which makes it hard to improve on with this little data, so it stays in the line-up as model 4.

Capacity is matched to data size. A regularized logistic regression learns a handful of coefficients a hundred-odd games can support; XGBoost carries far more and tends to memorize noise here, so it usually trails. It's scored anyway, so the gap is a number rather than a claim.

Model parameters

Two-step sampler

ParameterValueWhat it does
role halflife (days)120How fast old games stop counting toward role odds: a game this many days old carries half the weight of one played today.
role temperature0.6Sharpens the role odds before sampling. Below 1.0 favours the main role more; 1.0 would sample the raw frequencies.
champion recency halflife (days)100Same idea for champions: how quickly a champ fades from the pool when it stops being played.
sampling temperature0.7The variety knob. The displayed pick is sampled from the model's odds reshaped by this: lower is safer and more repetitive, higher is more chaotic.
min games to be a candidate3A champion only enters the candidate pool after this many games in the sampled role, so one-off picks don't pollute the odds.
regularization grid (C)[0.03, 0.1, 0.3, 1.0, 3.0]Candidate regularization strengths; grouped cross-validation on the training games picks whichever generalizes best.
featureschamp_freq_score, days_since_last_on_champ, in_role_score, is_last_pickDeliberately tiny feature set: a few hundred training games can't support more.
decision rulesample role, then sample championThe only model allowed randomness. Both draws are seeded from the previous game's id, so a re-run reproduces the identical pick.

Logistic argmax

ParameterValueWhat it does
recent-games window10Each champ's share of the last N games -- a strong 'what am I on right now' signal the two-step lacks.
min games to be a candidate3Same pool rule as the two-step model, but across ALL roles at once (no role step).
holdout games20The most recent games are hidden from training and used only to measure the model, so its choices are earned on unseen data.
featureschamp_freq_score, days_since_last_on_champ, is_last_pick, recent_shareRole-agnostic features over the whole champion pool.
decision ruleargmax + holdout auto-selectAlways shows its single most-likely champion. If the plain decay baseline beats the fitted model on held-out games, it deploys the baseline instead -- simpler wins ties.

XGBoost

ParameterValueWhat it does
n_estimators150How many small trees are stacked in sequence, each one correcting the errors of the trees before it. More trees means more capacity to fit -- and on tiny data, to memorize noise.
max_depth2How many yes/no splits each tree may make. Depth 2 keeps every tree nearly trivial, the main defence against overfitting a small dataset.
learning_rate0.1How much each new tree is allowed to change the running prediction. Lower is more cautious: many small corrections instead of a few big ones.
min_child_weight5A leaf must cover at least this much data to exist. Stops the model carving out rules that only ever applied to one or two games.
subsample / colsample_bytree0.8 / 0.8Each tree only sees a random 80% of the games and 80% of the features, so no single quirk of the data gets baked into every tree.
reg_lambda2.0An L2 penalty that shrinks leaf values toward zero, the boosted-tree cousin of the C knob on the logistic models.
featureschamp_freq_score, days_since_last_on_champ, is_last_pick, recent_shareIdentical inputs to the argmax model, imported from the same module. Any difference in the scores is the model, never the plumbing.
decision ruleargmax, always deploys itselfNever falls back to the decay baseline even when the baseline evaluates better. Deploying itself is the experiment: its hit rate is tracked head-to-head on the same table.

Decay frequency

ParameterValueWhat it does
recency halflife (days)100A champion's score is its pick count with old games fading out exponentially; a game this many days old counts half.
decision ruleargmax of decayed frequencyNo learning at all: 'whatever he's been playing lately'. The yardstick every fitted model must beat to justify existing.

Uniform random

ParameterValueWhat it does
poolevery champion ever playedAll champions in the history get an equal share, no matter how long ago or how rarely played.
decision ruleseeded uniform drawThe deliberately bad floor: long-run accuracy should sit near 1 / pool size. A model that can't clearly beat this isn't predicting anything.
Glossary
TermWhat it means here
logistic regressionThe simplest useful classifier: a weighted sum of the features pushed through a squashing function to get a probability. Few parameters, hard to overfit, easy to read (each feature gets one coefficient).
gradient boosting (XGBoost)Builds many small decision trees in sequence, each tree correcting the mistakes of the ones before. The industry default for tabular data, but it is data hungry: with only a few hundred games it tends to memorize noise instead of learning patterns.
decay / recency baselineNot a learned model at all: just 'what has he played lately', with older games fading out exponentially. Every real model must beat this to justify existing. It is embarrassingly hard to beat.
uniform floorA pick drawn at random from every champion ever played. Its accuracy is what zero skill looks like on this data -- the number every other model's accuracy should be read against.
argmax vs samplingArgmax always outputs the single most likely champion (accurate, repetitive). Sampling draws from the whole probability distribution (less accurate per pick, but varied and honest about uncertainty). The two-step samples; the comparison models argmax.
walk-forward backtestEvery row in the predictions table was computed as if standing just before that game: trained only on earlier games, seeded from the previous game's id. A backfilled prediction is informationally identical to one made live at the time.
temporal holdoutWithin each training pass, the most recent 20 games are hidden and used only for scoring, so model choices (like argmax's auto-select) are earned on games the model never saw.
grouped cross-validationDuring tuning, all candidate rows belonging to one game stay in the same fold, so the model can never peek at part of a game it is being tested on.
regularization (C / reg_lambda)A penalty on model complexity. Stronger regularization forces the model to stay simple, trading a little training-set fit for better behavior on new games.
accuracy (top-1)The share of games where the champion actually played was the model's pick. With ~275 scored games, differences of a few percentage points are statistical noise (roughly a ±6-point margin) -- read the scoreboard accordingly.
overfittingWhen a model learns the training data's coincidences instead of its patterns: great scores on games it has seen, poor scores on games it has not. The expected story here: the flexible XGBoost trails the simpler models, and that gap is the concrete reason simple models are trusted on small data.