ohfootball.io

Methodology

How The Ratings Work

The ratings are calculated using a modified elo system. This page is meant to serve as an overview of the model, its parameters, and its known limits.

The Update Rule

Before a game, each team has a rating. The expected score for team A against team B is a logistic function of the gap between them:

Ea=11+10(Rb−Ra)/400

After the game, the rating moves by the difference between what happened and what was expected, scaled by the K factor and by a per game multiplier:

Ra′=Ra+K·m·(Sa−Ea)

Sa is 1 for a win and 0 for a loss. Team B receives the exact opposite change.

Production Parameters

These values are used for the published snapshots. Each parameter was tuned independently using the 2000–2023 seasons as the training/validation set, with final performance evaluated on the 2024–2025 seasons as the test set.

Initial rating
1500
The rating of a team with no history and no division.
K factor
148
The maximum rating a single game can move. It is large because a season is only about ten games.
Rating scale
400
A 400 point gap means the stronger team is expected to win about 91% of the time.
Home advantage
30
Added to the home team's rating before the probability is calculated. It never changes the stored rating.
Season carryover
0.85
The share of last season's ending rating that a returning program keeps.
Division rating step
140
Points per division of separation in the preseason prior.
Provisional games
3
How long a team's early-season rating moves faster than normal.
Provisional K multiplier
1.6
The size of that early-season boost. It decays linearly to 1.0.
Margin weight
0
Margin of victory is available in the model but is switched off in production.

Where A Season Starts

A team's preseason rating is built from its division, then pulled toward what the program finished with in the most recent season it played:

prior=1500+140·(4−division)start=prior+0.85·(last played rating−prior)

Division I sits 420 points above the baseline and Division VII sits 420 below it, with Division IV at the baseline. Independent teams get no division adjustment. A program that has never played stays at its prior.

A program that stops for a season or more keeps the rating it last earned rather than starting again at its prior, because a team that comes back is not a new team. The pull toward the prior is applied once, however long the program was away. A program that returns in a different division is pulled toward the prior of the division it returns in.

Early Season Behavior

For a team's first three games, the K factor is multiplied by a boost that starts at 1.6 and decays linearly to 1.0. Both teams in a game share one multiplier, taken from whichever team is further from settled. This is an attempt to reduce the amount of variance in early season matchups since the model doesn't take into account things like roster changes, injuries, coaching changes, etc.

What Counts

A win, a loss, and a tie all move ratings, with a tie counted as half a win for both teams. A forfeit never moves a rating, because no team played the game. Cancellations are excluded for the same reason. Upcoming games get a probability but never change a rating.

How It Is Measured

Every configuration is scored by replaying history one game at a time and grading the prediction that was made before the result was known. Three numbers are tracked: Brier score, log loss, and straight accuracy on games with a decided result. Brier score and log loss both reward calibration, so a model that says 90% needs to be right about 90% of the time, not merely on the correct side.

Where The Data Comes From

The data is sourced from a combination of Joe Eitel and the Ohio Highschool Football Database.

Are The Predictions Any Good?

Idk. Historical accuracy sits around 80%, so it's better than a coin flip. I think a definition of "good" would be when it's able to consistently outpredict humans. An example might be checking if the model's predictions are better than WFMJ's predictions.

Known Limits

Margin of victory is ignored
A one point win and a forty point win move a rating by exactly the same amount. This keeps the model resistant to running up the score.
Ratings are zero sum inside a season
Every point one team gains, its opponent loses. Ohio as a whole cannot get stronger or weaker, so ratings compare teams to each other and not to some fixed standard.
Schedules are regional
Most teams play a narrow local schedule. A team that dominates a weak area can carry a higher rating than it deserves until it plays outside that area, which often means the playoffs.
Out-of-state opponents are invisible
Only games between rated Ohio teams update ratings. A game against an out-of-state program is skipped, so it neither helps nor hurts.
The preseason prior is a guess about school size
Division is mostly an enrollment bracket, not a strength measure. Using it as a prior helps on average and is wrong for any specific program that is unusually strong or weak for its size.
Program continuity can break
Carryover follows a program identifier across seasons, so a program that misses a season keeps the rating it last earned. Co-ops, mergers, and renames can still split one program into two histories, which resets a team to its prior. This will cause really awful predictions for brand new schools.
The model knows nothing about football
There is no roster, no injury report, no weather, no travel distance, and no notion of matchup style. It simply looks at who played and what the result was.
Source data can be wrong
Results are scraped from a public sites. While these sites are awesome, any errors will flow straight to this model.