Most lead-scoring pitches ask you to take the number on faith. We would rather show you the check. Every time Follow Up Ace meaningfully changes a contact's score, it logs that prediction — and an hourly job comes back later to stamp what actually happened at 24 hours and 7 days.
On the most recent measured week, the response rates lined up in the order the model predicted. Contacts scored Hot responded within 24 hours 47.5% of the time. Warm came in at 15.2%, Cool at 9.4%, Cold at 1.0%, and Dormant at 0.6%. That is a clean monotone ladder measured against outcomes, not a claim about outcomes.
Two details make it trustworthy rather than flattering. Model output is calibrated with isotonic regression before anyone sees a percentage, because a raw class-weighted classifier score is not a probability. And because the prediction log only fires on contacts whose scores moved, it also writes a small unbiased random control arm every run — so the selection bias in the main stream can be measured directly instead of assumed away.
24-hour response rate by score tier, measured from logged predictions against recorded outcomes on the most recent full week of production data. Rates shift week to week with lead mix and seasonality; the ordering is the durable result. Full engineering detail in the case study.