The latest reviews look worse. Your competitor shipped an update last week. It is tempting to connect the two events with a thick red arrow.

Keep the arrow in pencil until you inspect the timing and the sample.

Establish the release boundary

Record the version, release date, store and country. Reviews may arrive after the experience they describe, and not every reviewer identifies their version. Some people update immediately; others continue using older releases.

Group reviews by observed version when available. Otherwise, label the grouping as based on review date and retain that limitation.

Compare similar windows

Use comparable lengths before and after the release. Keep the same language, storefront and collection method. A shift from recent reviews to “most helpful” reviews can manufacture a change that belongs to the interface filter.

Count the relevant complaint theme as a share of your collected sample, while acknowledging that the sample itself may not represent all users. A higher count can simply reflect more total reviews.

Look for a precise symptom

“App got worse” is difficult to test. “Export now produces an empty file on this device” is more actionable. Repeated descriptions of the same new symptom, aligned with version information, deserve closer investigation.

A later patch or developer response can provide additional context. Preserve those observations separately instead of rewriting your original notes as though you knew the outcome all along.

For a founder studying opportunities, the lesson may be about release discipline rather than a new niche. Add a regression check for the critical workflow in your own app.

Publishing an accusation that an update “destroyed” a product requires much stronger evidence than a handful of unhappy reviews. A careful before-and-after notebook can still uncover a useful engineering lesson without turning correlation into a public verdict.