MLB Bullpen Analytics · Decision Support

The Arm Tax

A decision tool that turns a reliever's recent workload into a run-value penalty — measured on the exact scale a coaching staff already uses to compare matchups. So the dugout can finally see the moment a tired top-tier arm stops being the better call than a fresh second option.

Built for an MLB coaching staff 11,000+ pitches analyzed Output: a nightly bullpen briefing

Developed hands-on with MLB teams — working directly with a club's GM and pitching coaches to fit the tool to how real in-game and workload decisions actually get made.

The Problem

The matchup grid is blind to fatigue

Staffs already rank every reliever-vs-hitter matchup on a run-value scale. But that grid treats a pitcher the same whether he's fully rested or three appearances into a hard week — and the people tracking arm workload sit a department away from the people making the 7th-inning call.

The matchup grid alone

  • ✕ Ranks arms as if every night is fresh
  • ✕ Ignores pitches, back-to-backs, and bullpen warm-ups
  • ✕ Workload data lives outside the in-game decision
  • ✕ No way to compare a tired ace to a fresh middle arm

With the Arm Tax applied

  • ✓ Every arm's value is discounted for recent use
  • ✓ Game pitches, warm-ups, and rest days all count
  • ✓ Fatigue speaks the same unit as the matchup grid
  • ✓ Tonight's true pecking order, ranked and ready

The Idea

One number, in the unit they already trust

The whole design rests on a single move: express fatigue in the same run-value units as the staff's projections. If the penalty and the projection share a scale, the dugout can simply subtract one from the other — no new mental model, no competing dashboard.

Their number

Projected value

What each arm is worth in a clean matchup, on a run-value scale.

Our number

The Arm Tax

A fatigue penalty in the very same units — earned from this week's usage.

The call

Adjusted value

Projection minus tax. The honest ranking of who's actually best tonight.

How It Works

Four stages, public and private data in, a ranking out

The pipeline fuses public pitch-tracking and game logs with the club's private feeds — bullpen warm-up counts and their proprietary talent numbers — ingested automatically off their existing internal workflows. Each stage is calibrated to the individual pitcher rather than a league-wide rule of thumb.

01

Ingest

Public pitch tracking and game logs, plus the club's warm-up counts and talent numbers — auto-pulled from their internal workflows.

→

02

Workload spine

An acute-vs-chronic ratio flags whoever has spiked off their own baseline.

→

03

Arm Tax

A decay model converts recent load into a run-value penalty per pitcher.

→

04

Briefing

Adjusted values ranked into a clean go / caution / hard-stop card.

The Model · the shape, not the recipe

What goes into the tax

Fatigue isn't one thing, so the penalty is built from a handful of load sources, then decayed by how long ago they happened and escalated when a pitcher is used on back-to-back days.

# recent load, decayed by recency, scaled to runs
load = game pitches + warm-up cost + over-threshold penalty
decayed = Σ ( loadd × recoverydays ago )
arm tax = decayed × streak multiplier × per-pitcher scale
warm-up cost

Bullpen pitches thrown while getting hot count too — weighted toward game intensity, not ignored.

over-threshold penalty

Pitches beyond a pitcher's workload threshold cost more, mirroring how recovery time climbs.

recovery decay

Yesterday hurts more than three days ago; load fades on a per-pitcher recovery curve.

streak multiplier

Back-to-back days escalate the penalty, with a hard stop at three straight.

Workload Spine

Aligned to the sheet they already read

The sports-performance group monitors an acute-to-chronic workload ratio — recent volume against each pitcher's own baseline. The tool computes the same ratio so the model lives inside the staff's existing language instead of competing with it. A spike off baseline is the trigger.

< 0.8

Under-worked

0.8 – 1.3

In the sweet spot

1.3 – 1.5

Elevated — watch

> 1.5

Spike — at risk

Calibration

Tuned per pitcher, anchored to their projections

Two calibration steps keep the model honest: one learns each arm's personal fatigue sensitivity from the tracking data, the other reverse-engineers the blend of public projection systems that best matches the club's proprietary number.

Per-pitcher sensitivity

Learned from the tape

Each arm's recovery rate and fatigue scale are estimated from rested-vs-tired performance splits — durable arms get taxed lightly, fragile ones heavily. Defaults give way to data as the season fills in.

Projection blend · NNLS

Matching the trusted number

A non-negative least-squares fit finds the mix of public systems that tracks the staff's internal projection. The solution is deliberately sparse and legible — "these two systems carry it" — so the staff can trust what they can read.

A Signature Finding

Some arms warn you with spin, not speed

Staffs widely watch velocity for fatigue. Testing whether spin or velocity drops first on short rest — controlling for within-game fade and pitch mix — surfaced something the existing process misses.

🌀 Per-pitcher fatigue fingerprint

For one reliever, spin fell sharply on short rest while velocity held steady

The drop was statistically significant; the velocity reading was flat. For that pitcher, spin is the earlier warning light — and a velocity-only watch would miss it entirely. Another arm showed the textbook opposite: velocity led, spin was noise.

The takeaway isn't a league-wide rule. It's that fatigue has a personal signature, and the tool flags which metric to watch for each arm. Just as telling: where public data can't read a pitcher cleanly — usually because he's only sent out on short rest when he feels sharp — that selection bias is exactly what the club's private warm-up and workload feeds, ingested automatically, are there to cut through.

The Output

Tonight's bullpen briefing

Everything resolves to one card the dugout can read in seconds: each arm's projection, its fatigue tax, the adjusted value, and a plain verdict. Illustrative names below — the structure is the point.

RelieverAdjusted valueTonight
Arm A Tier 1 · 1 day rest +6.1 − 4.0 tax Caution
Arm B Tier 2 · 3 days rest +4.4 − 0.3 tax Full go
Arm C Tier 1 · 3rd straight day +1.2 − 7.6 tax Hard stop

The decision the card forces

On paper, Arm A is the better matchup. After fatigue, the fresh Tier-2 arm grades out ahead — exactly the trade-off the staff couldn't see before. And Arm C is off the board tonight on a biomechanical hard stop, not a hunch.

Python Streamlit pandas · numpy · scipy Statcast / pitch-tracking MLB Stats API NNLS calibration

The Arm Tax — fatigue-adjusted bullpen decision support.  ·  Client and player details anonymized.  ·  Portfolio piece.