- Collecting outcome measures is table stakes. Packaging them for a payer is the part almost nobody does.
- A payer cares about three things: cost per episode, readmission or return to a higher level of care, and completion rates.
- Clinical significance beats statistical significance in this conversation. Say how many people got measurably better, not what your mean change was.
- Bring the denominator. A percentage without the count underneath it reads as marketing.
Most treatment programs walk into a rate conversation with a story about the quality of their care and walk out with the same rate they had before. The programs that move a rate walk in with a document.
The difference is not that one program delivers better care. It is that one of them can prove something the payer's actuary can use.
What the payer is actually optimising
A network manager is not evaluating your clinical philosophy. They are managing total cost of care for a population, and behavioral health matters to them mostly through its effect on other spend: emergency department use, inpatient admissions, and the cost of people cycling through repeated episodes of treatment.
That means the three questions behind any rate conversation are:
- What does an episode with you cost, end to end? Not your per-diem — the whole episode including length of stay.
- Do people come back? Return to a higher level of care within thirty, sixty, or ninety days is the number that shapes their view of you.
- Do people finish? A program with a high completion rate is a program whose episodes are predictable to price.
Your PHQ-9 trend is relevant, but only because it supports an answer to one of those three. Lead with theirs, not yours.
Twenty-four standardized assessments is table stakes. Packaging them as evidence is not.
What the document should contain
A defined population
“Our IOP patients in 2026” is not a population. “All patients admitted to IOP between January and June 2026 with this payer, excluding those who transferred in above week four” is. Define it once, at the top, and use the same definition for every number that follows.
A baseline and a follow-up, on the same people
The most common self-inflicted wound is reporting intake scores across everyone and discharge scores only across completers. It inflates your result and it is immediately visible to anyone who reads carefully. Report matched pairs, and report how many people you could not match.
Clinically significant change, not just averages
“Mean PHQ-9 change of −9.4” is a statistic. “Sixty-eight percent of episodes showed clinically significant improvement, using a five-point threshold” is a claim about people. Use the second and show the first underneath it.
The counts
Every percentage needs its denominator adjacent to it. 68% of 131 matched episodes is credible. 68% on its own reads like a brochure.
What did not work
Include a number that is not flattering — a cohort that did not improve, a level of care where your results are ordinary. It costs you nothing with a payer who was going to find it anyway, and it changes how they read everything else on the page.
A structure that works
| Section | What goes in it |
|---|---|
| Population | Inclusion rules, date range, payer, level of care, N |
| Volume | Episodes, bed-days or visits, by month |
| Outcomes | Matched baseline and discharge, % clinically improved, N matched |
| Utilization | Average length of stay, completion rate, step-down pattern |
| Return to care | Readmission or higher level of care within 30 / 90 days |
| Limitations | What you could not measure and why |
Timing
Bring this six months before renewal, not six weeks. A network manager who sees your data for the first time during a negotiation has no way to validate it and every incentive to discount it. A network manager who has seen the same report twice already is being asked to act on something familiar.
Do not send patient-level data. Aggregate everything, state your minimum cell size, and suppress anything below it. A rate conversation that becomes a privacy incident has gone badly wrong in a way no rate increase compensates for.
Why most programs skip this
Not because it is hard analytically. Because assembling it means pulling assessment scores from one system, admissions and discharges from another, and claims from a third, then reconciling three different definitions of what an episode is.
Programs where those three live in one record produce this in an afternoon. Programs where they do not produce it never, which is the real reason the conversation does not happen.
