You are reading the scores wrong
We all read an 85 as high. Nearly every score lands between 80 and 92, which puts 85 near the middle of the pile.

You see an 85 on a monitor and file it away as good. Not the best, but clear of the middle, with room above and plenty below. That habit comes from school, and reading a review score like a school grade is the most common mistake a shopper makes.
The mistake is not that 85 is a low number. It is that the scale you are picturing, the one running from 0 to 100 with failure at the bottom, does not exist in practice. SetupScore's set holds 104 products, each with at least 2 verified expert ratings. Nothing in it scored below 50 or above 97.
Why do most review scores land in the same narrow band?
Half of that scale is empty, and the half that remains is narrower than it looks. Most of the field crowds into a single run of scores, from 80 to 92, which means almost four products in five sit in a twelve point band. SetupScore's analysis of 104 product scores found 78.8% of them packed into a twelve point stretch of the scale. Only 16 products, 15.4% of the set, sit below 80.
There is not much room left to separate 104 products. The median score is 85.4, already above the number most shoppers treat as respectable. Above an 85 sit 54 products, and below it 46. So 51.9% of the field outranks an 85, which is the opposite of what an 85 feels like.
Almost the whole axis is empty
Tight packing changes what a small difference is worth, and near the middle 2 points can move you past a sixth of the field. Across the 104 products SetupScore has scored, an 85 beats 44.2% of the field. An 87 beats 61.5% of them. That gap is worth 17.3 points of rank. The same 2 points near the top would barely move a product.
The top of the range has almost nothing left to move through. Only 6 products in the whole set score above 92, and a score of 92 already beats 92.3% of everything we cover. The bottom edge is just as thin, where a score of 75 beats 6.7% of the field. Almost everything else sits in the middle.
A score is a position, not a grade
View the numbers as a table
| score | share of the field this score beats |
|---|---|
| 75 | 6.7% |
| 80 | 15.4% |
| 82 | 27.9% |
| 85 | 44.2% |
| 87 | 61.5% |
| 90 | 79.8% |
| 92 | 92.3% |
| 95 | 98.1% |
Two forces produce a field this tight, and the effect has a name: score compression. The first is selection, or range restriction: the products that get reviewed are already survivors, picked because readers are already shopping for them, which strips out the broken and the obscure before scoring starts. A field of finalists does not spread out.
The second force works on the numbers themselves. Reviewers publish on different scales, and a five point scale and a hundred point scale do not mean the same thing. One offers a handful of steps, the other a hundred.
A scale with 5 steps has only 5 available answers, and converting it to a 100 point axis does not create more. A rating of 4 out of 5 becomes an 80. A rating of 4.5 out of 5 becomes a 90. Coarse scales can only land on a handful of fixed values, and those values sit near the middle of the axis.
There is an obvious objection here: we may have created the clustering ourselves. We publish an average of several ratings per product, and averaging always pulls numbers closer together. Two things answer that objection.
The first is that the clustering was there before we averaged anything. According to SetupScore's analysis of 599 expert ratings, 63.9% of them land inside a ten point range. The individual ratings behave the same way on their own, and 383 of them fall between 80 and 90. The middle half spans the same 10 points.
The second is a test of that worry. We dealt the 599 real ratings out at random into imaginary products of the same sizes, and repeated that 4,000 times. If averaging were doing the work, the imaginary products would spread as widely as the real ones. They did not. The random sets came out tighter, 5.6 points of spread against 7.3 for the real ones. In SetupScore's shuffle test of 599 expert ratings, only 27 of the 4,000 random reassignments reached the spread of the real products. Averaging does squeeze the numbers, but it is not what put these products so close together.
So the number is not a grade. It is a position in a short, tightly packed queue, and it behaves like one. An 85 does not mean 85 percent of anything, and it does not mean the product got 85 percent of the way to some ideal. It means several people who tested it landed near the middle of a narrow band, and none of them found a reason to push it higher.
Once you see the queue, the number changes job. A score is a rank, not a measurement. It does not tell you how good a product is in some absolute sense. It tells you where that product stands in a line of 104.
Read the distance, not the number
A position in that line is only useful if you know what it is worth. Figure 2 turns each score into a percentile rank, the share of the field it beats, so you can read across and see where any score really sits. The shape matters more than any single row: the steps are crowded through the middle of the scale and stretched almost flat at the top.
The practical move is to compare two scores against each other rather than one score against an imaginary 100. A gap of 3 points between two products near the middle is a large difference in rank. The same 3 points above 92 covers almost nothing, because almost nothing is there.
Money follows the top of the scale, but the link is weaker than it looks. Across the 103 priced products in SetupScore's set the correlation is r = 0.089, so price barely predicts where a product ranks. Median prices do still climb, from $160 in the lowest band to $640 in the highest.
What the top of the scale costs
Shopping only at 90 and above means choosing to spend roughly 4 times more. We have a stake in that decision. SetupScore earns affiliate commission as a percentage of the sale price, so advice that pushes you up the scale increases what we are paid.
How should you read a review score?
None of this makes scores useless. It makes them a narrower instrument than they look, and there are better questions to ask of one. Start with position. Out of everything that reviewer has rated, where does this product sit? Any outlet with a back catalogue has built a line of its own, and the score in front of you has a place in it. That place is the information. The distance to 100 is not.
How the number was reached changes what it is worth, so it pays to work out where the score came from. Some are the output of a system, fed by measurements and repeatable across products. Others are a verdict, reached in one move by somebody who lived with the thing for a week. Both are worth reading. They are not the same claim.
A decimal is a cheap tell. Seeing 8.7 rather than 9 probably means arithmetic happened somewhere, though it settles nothing on its own, because a finer opinion is still an opinion. What it does suggest is that the number was built out of parts.
An overall score is a blend, and the weights inside it belong to the reviewer. A pair of headphones can lose ground on portability because the cups do not fold. That is a real fault for somebody carrying them onto a train every morning, and no fault at all for somebody who listens at a desk. The total hides the difference. Aspect scores show it, which is why the parts are worth more to you than the total.
How much evidence sits behind a score matters as much as the score itself. A number resting on 2 opinions is not the same object as one resting on 16, even when both land on an 85. Our own set thins out fast at that end. Of the 104 products SetupScore covers, 18 rest on the minimum of 2 ratings, where one more review can move the score a long way. Ask that of any score you meet, including ours.
How we handle it here
This is the problem SetupScore was built around. Scores from different outlets do not share an axis, so they are normalized onto one common scale before anything is combined, which is why a 4 out of 5 and an 80 can land in the same column. Most of what feeds a page carries no score at all. Only 792 of the 3,993 sources behind SetupScore pages arrive with a number attached, and the rest, largely YouTube and Reddit, are read as text and tone. Each product is then split into named aspects with their own scores and quotes, so the weighting can be yours instead of ours. None of that removes the need to read. It only makes the number easier to argue with.
