Microsoft measured what happens when smart people guess. Across years of controlled experiments, about a third of ideas improved the metric they targeted. A third did nothing. A third made things worse. Ron Kohavi, who ran Bing's experimentation platform, published the split with Stefan Thomke in Harvard Business Review in 2017.
Sit with that bottom third for a moment. These were funded ideas, argued for in meetings by experienced people, and shipping them untested would have quietly damaged the product. Without the experiments, nobody would have known which third they were standing in. The changes ship, the metrics drift, and the story becomes seasonality.
The same paper carries the other half of the lesson. An engineer proposed a small change to how Bing displayed ad headlines. It sat in the backlog for months because it looked minor. When it finally ran, it produced a 12 percent revenue lift, worth over $100 million a year at the time. The best idea of the year had looked like the least interesting ticket in the queue.
- 1/3of ideas improved the metric they were designed to improve
- 1/3had no measurable effect
- 1/3made the metric worse
- +12%revenue from the ad-headline change nobody prioritized
Kohavi and Thomke, Harvard Business Review, 2017.
Both stories argue for the same posture. Intuition is a bad ranker. Cheap, frequent testing beats expensive, occasional testing, because the wins hide where nobody would bet and the losses wear convincing costumes.
For a small site the practical floor is lower than the folklore says, but it is real. You need enough conversions for the arithmetic to close. A two-variant test on your highest-traffic step will finish. A five-variant test on a trickle will not, and the checkout, where the money already flows, is a better first subject than the tagline.