How we test
A score is only worth something if you can reconstruct it. So here's the whole protocol: what we buy, how long we wear it, what we measure it against, and exactly how six axis scores become one number.
Applies to: every review published after 1 September 2026
Changelog: published at the bottom of this page
The protocol
Buy it ourselves
Retail purchase, personal cards, standard shipping. We never accept review units, press access or comped memberships. If we can't buy it, we don't review it.
31 days minimum
Two testers minimum, four when the marketing claims are loud. Programme-based products get the full programme length — up to 90 days.
Benchmark it
Sleep against home PSG. Heart rate against an ECG chest strap. Glucose against paired fingersticks. Blood markers against an accredited lab draw.
Tear it down
Network capture on first launch, SDK inventory, permission audit, export test, and a real account-deletion request that we follow up 30 days later.
The six axes
Each is scored 0–10 by the lead reviewer, challenged by a second reviewer, and defended in the Monday meeting. Weights are equal — we've never found a defensible reason to make one axis worth more than another.
Composite = (sum of six axes ÷ 6) × 10, rounded to the nearest whole number.
What each axis actually asks
- Measurement validity — does the number match a reference instrument, and does the app admit its error bars?
- Evidence & claims — we read every study cited. Company-funded, unpublished or n<20 gets named in the review.
- Data rights — trackers, SDKs, export formats, deletion that deletes, and what happens when you stop paying.
- Daily UX — 31 days of real use, plus a dark-pattern count against the taxonomy below.
- Price & value — total 24-month cost at list price, including hardware and consumables.
- Staying power — funding, shipping cadence, support response times, and shutdown risk.
85–100 Best in category · 70–84 Good, with named trade-offs · 55–69 Works, something meaningful is broken · 40–54 Narrow use case only · <40 We'd ask for a refund.
The dark-pattern taxonomy
We count, we don't vibe. Every instance is screenshotted, dated and logged against the app's build number. The category median is five per 31-day window.
| Pattern | What it looks like | Penalty |
|---|---|---|
| Roach motel | Cancellation takes more steps than signup | −1.0 UX |
| Manufactured urgency | Countdown timers on evergreen offers | −0.5 UX |
| Guilt streaks | Punishing language for a rest day | −0.5 UX |
| Pre-checked consent | Marketing or data-sharing opt-in enabled by default | −1.0 Data rights |
| Data hostage | Historical data locked behind an active subscription | −1.5 Data rights |
| Silent price increase | List price raised without notifying subscribers | −1.0 Price |
| Fake personalisation | "Your plan" is identical across test accounts | −1.0 Evidence |
| Pathologised normal | Normal physiology framed as damage, unsourced | −1.0 Evidence |
The rules we don't bend
- We buy everything. No review units, no press accounts, no comped memberships.
- No affiliate revenue. Not now, not "just for hardware," not ever.
- No sponsored content. We don't publish it, and we don't accept ads from companies we score.
- No pre-publication review. App makers see the piece when you do. Factual queries only, after publication.
- The firewall. Intelligence clients get no influence over scores, no advance copies, and no say in the schedule.
- Disclosed conflicts. Any team member holding equity in a scored company recuses themselves, and we say so in the review.
- Corrections in public. Dated, signed, kept at the bottom of the piece forever.
- Re-testing. Scores expire. Every app is re-tested every 12 months or every two major releases.
Methodology changelog
Composite moved from five axes to six after three shutdowns in twelve months.
Per-pattern deductions published rather than applied at reviewer discretion.
Two-week tests were missing month-two retention behaviour and renewal nags.