A star rating is widely used as a forward-looking recommendation. It is built entirely from realised performance, and the distance between those two things is measurable on the rating agency’s own published data.

What follows uses Morningstar’s own work on how its ratings have performed - not an outside critique - and concerns one narrow question: what happened to funds after they reached the top of the scale.

Three Consecutive Cohorts, One Direction

Average star rating of three vintages of Morningstar five-star domestic funds, at issue versus five years later.

The 2004 class of five-star domestic funds averaged 3.2 stars five years on. The 2005 class averaged 3.1. The 2006 class averaged 2.9. Three consecutive vintages, three declines, landing at or below the middle of a five-point scale. Across these cohorts, five-star funds went on to underperform their own benchmark by roughly two percentage points.

The consistency is what makes this more than an anecdote. A single bad vintage is variance. Three in a row, moving the same way by similar magnitudes, is a property of the measurement rather than of the period.

Only One in Seven Held the Rating for a Decade

The longer window is starker. Of the funds carrying the highest five-star rating in July 2004, just 58 of 403 - 14% - still held that rating in July 2014. Six in seven had fallen off the top of the scale within ten years.

Put differently: an investor selecting on a five-star rating in 2004 had roughly a one-in-seven chance that the attribute they selected on would still describe the fund a decade later. For a holding period that long, the rating was closer to noise than to signal about the terminal state.

This Is Regression, Not Mismeasurement

None of this shows the rating is calculated badly. It shows regression to the mean operating on any ranking built from a noisy process.

A fund reaches five stars by combining genuine skill with a run of favourable variance. The skill component persists to some degree. The variance component - which did real work in lifting the fund to the top of the distribution, precisely because reaching an extreme requires the luck to have run the right way - does not repeat on demand. Selecting on realised outcomes therefore selects partly on the non-repeating part, and the stronger the selection, the larger the share of it that will not repeat.

This is why the effect is strongest at the top of the scale rather than in the middle. Five stars is an extreme, and extremes are where the luck component is largest.

What the Rating Still Does Well

The rating compresses genuinely useful information about risk-adjusted history into a single number, and it does that honestly. A consistently low-rated fund remains a meaningful warning - persistent poor risk-adjusted returns are more likely to reflect structural problems, cost drag or mandate drift than a run of bad variance.

The asymmetry is worth stating plainly, because it is easy to over-read this in the opposite direction. The evidence undermines the rating as a forecast of future outperformance. It does not undermine it as a summary of what already happened, and it does not make a one-star fund a contrarian buy.

What This Means for Allocators

Use the rating as a screen for what to examine, not as a conclusion. It tells you where to start reading a factsheet, not which fund to hold.

Weight persistence over level. A fund that has held a mid-to-high rating across several distinct market regimes is making a different claim from one that reached five stars inside a single favourable one.

And read the rating as tense. It is a description of where a fund has been. The published record on what happens next says the gap is wide enough to matter.

Back to Writing