Loyalty program reporting is close to perfectly designed to mislead, and almost nobody involved is being dishonest.
Members spend more than non-members. They buy more often. They stay longer. Every one of those statements is true, and every one of them was true before anyone enrolled. Programs recruit your best customers, and then the program is credited with the behavior that made them join in the first place.
Self-selection, in one sentence
The customers who opt into a loyalty program are, by construction, the customers who expected to buy from you enough for it to be worth the effort.
That is the entire mechanism. It requires no bad faith and no bad math. It only requires the standard comparison, members versus non-members, which measures who enrolled rather than what enrollment did.
The gap between those two things is usually large. It is not unusual for a program reporting a substantial return on the member-versus-non-member comparison to return a small fraction of that against a proper control, and occasionally nothing at all. Finding out which is true of your program is, for most loyalty teams, the single highest-value analysis available.
The comparison that measures the program
Withhold the program, or a specific mechanic inside it, from a randomly assigned share of eligible customers. Read the result as the difference between the two groups.
Random assignment is what makes it work. It is the same instrument that separates incremental return from self-reported credit in paid media, and there is no principled reason retention should be exempt from a standard we apply to media. If a win-back campaign cannot show lift against a control, it is a report of customers who came back, not a program that brought them.
The common objection is that the program already launched and everyone is already in it. That is a real constraint on measuring the program as a whole, and it is not a reason to measure nothing. You can randomize new mechanics: a bonus event, a new benefit, a tier promotion, a points multiplier. Programs that start doing this usually discover that a small number of mechanics are doing most of whatever incremental work is being done, and the rest is margin moving quietly out the door.
What the cost actually is
The loyalty budget line that gets scrutinized is the platform and the team. That is the smaller number.
The real cost is margin: the rewards, the points liability, the discounts that go to customers whose behavior did not change. That number rarely appears in the same report as the benefit, and until it does the conversation cannot resolve, because there is nothing to weigh the return against.
Put both in one view and the discussion improves immediately. It stops being about whether members like the program, which they reliably do, since it gives them things, and becomes a question about whether the behavior created is worth what the program spends. That is a question a CFO can engage with. A satisfaction score is not.
Mechanics that change a decision, and mechanics that reduce a price
Once the measurement is honest, design gets more interesting, because the question becomes what would actually alter behavior rather than what would be popular.
Accumulated progress. Status, tiers, and stored value create a genuine cost to leaving. This works when the status is worth something and becomes an expensive fiction when the benefits are cosmetic.
Personalization that improves with tenure. A relationship that genuinely gets better the longer it runs, better fit, better recommendations, faster service, is a switching cost no competitor can match on day one.
Access rather than discount. Early access, exclusive assortment, and service-level benefits change the choice without moving the price, which protects the margin the program is spending.
Earn-and-burn discounting. The default, the easiest to launch, and the least likely to be incremental. It carries a specific long-term risk beyond the margin cost: it teaches a base that was not price sensitive to become price sensitive, and that is very difficult to reverse.
Which of these is right is not a matter of taste. It follows from what the churn decomposition says. A base leaking through never-activated churn does not need a tier structure; it needs an onboarding fix, and a loyalty program layered on top will spend margin on customers who were already fine.
The uncomfortable quarter
A team that introduces holdouts to a program that never had them should expect the reported number to fall, because the previous number included everyone who was coming back anyway.
That is not a failure of the program and it should not be handled as one. It is the first quarter in which the number means something. The programs that survive this are the ones where leadership decided in advance that they wanted the real figure, and the ones that do not survive it were spending margin on a story.
More on how this fits the rest of the work: loyalty and win-back covers program design against the churn mix, and subscription economics sets the ceiling on what recovering any customer is worth.
So: if your loyalty program were switched off tomorrow for a randomly chosen tenth of members, do you know what would happen? And if not, what exactly is the reported return describing?