Research · How we draw the boundaries

One method, very different worlds: why these numbers never share one ranking

We run one identical method across every app we track: pull the public review feed, classify each written complaint, and report the 1–2 star share against Apple's own star distribution. Below is what that same method returns for 3 tracks serving completely different people.

7.4%
median complaint rate · Field service
21 apps
6.5%
median complaint rate · AI companion & wellbeing
14 apps
3.8%
median complaint rate · Shift work & scheduling
11 apps
46
apps across 3 tracks
Different collection dates. Field service read on 17 Sep 2026;AI companion & wellbeing read on 19 Sep 2026;Shift work & scheduling read on 19 Sep 2026. They are not a same-morning snapshot, so read the gap between tracks as structural, not as "this week vs last week". We date every page by when the reviews were actually pulled, never by when the page was last rebuilt.
Uneven group sizes. We track 21 apps in Field service but only 11 in Shift work & scheduling. A median over a larger group moves less when one app changes; treat the smaller group's median as the looser of the two.

The headline gap

across 21 Field service apps it is 7.4%;across 14 AI companion & wellbeing apps it is 6.5%;across 11 Shift work & scheduling apps it is 3.8%. That is not a rounding difference, and it is not one bad app dragging a line down — these are the middles of 3 separate distributions.

1.9×
between the loudest track (Field service, 7.4%) and the quietest (Shift work & scheduling, 3.8%). If we had pooled all 46 apps into one list and ranked them against a single median, every app in the quieter tracks would have looked like a top performer and every app in the louder one would have looked like a disaster — and neither would have been true.

Where the complaints differ most

These are the categories with the widest spread across the tracks. Read them as structure, not verdict: the apps are used by different people in different situations, so they fail in different places.

CategoryHighest shareTrack SpreadRatio
crash18.0%Shift work & scheduling11.7 pp2.9×
login11.5%Shift work & scheduling9.0 pp4.6×
price7.2%AI companion & wellbeing6.0 pp6.0×
subscription6.0%Field service5.2 pp7.5×
billing5.6%Field service5.0 pp9.3×

The full comparison

Every category, every track, computed the same way. Highlighted = the highest in the row.

CategoryField serviceAI companion & wellbeingShift work & schedulingSpread
crash11.9%6.3%18.0%11.7 pp
login2.6%2.5%11.5%9.0 pp
price5.4%7.2%1.2%6.0 pp
regression3.0%2.6%6.9%4.3 pp
missing5.7%3.9%6.4%2.5 pp
subscription6.0%2.9%0.8%5.2 pp
billing5.6%0.8%0.6%5.0 pp
support5.0%1.6%1.5%3.5 pp
ads0.4%2.6%1.2%2.2 pp
sync1.1%1.1%2.5%1.4 pp
perf1.7%0.5%2.4%1.9 pp
competitor1.4%1.4%1.2%0.2 pp
ux1.4%0.5%1.2%0.9 pp
accuracy0.9%0.8%1.3%0.5 pp

Shares are the median across the apps in each track of that category's percentage of the app's written complaints. Categories are assigned by rule-based classification and then manually reviewed; unclassified complaints are left out rather than forced into a bucket.

Why we keep the tracks apart

Because a complaint rate is only a comparison if the things being compared actually compete for the same user. A field-service app and a shift-scheduling app might both sit in the same vendor's product suite, but the person tapping the screen is different, what they were trying to do is different, and what counts as a bad evening is different.

So each track gets its own median, its own ranking, its own coverage figures, and its own collection date. A track with fewer than 3 apps gets no rank and no multiple on its pages — it still gets published, just without pretending it has a distribution behind it. Two numbers have no median worth quoting; three is where we start.

What this costs us to maintain

3 tracks means 3 baselines to recompute every time anything changes, and every new app has to answer "which group does it compete in?" before it gets a page. That answer is a judgement call and we write it down per app. It also means the reply to "can you rank us against everyone you track?" is sometimes no, and here is why.

Want this run on your own app?

Free, no signup. Send us your App Store link and we'll classify your reviews and send back the breakdown — ranked against your own track, not against every app we happen to track.

How it works →