What the say-do gap is
The say-do gap is the difference between what people tell a researcher they will do and what they actually do afterwards. Market-research practice uses exactly that name[1]; the academic literature calls the same phenomenon the intention-behavior gap, or the attitude-behavior gap, and has been measuring it for decades[2]. Economists describe the money version of it as the difference between stated preference (what you say a thing is worth) and revealed preference (what you hand over at the till).
It is not lying. People answer honestly and then behave differently, because the answer and the behavior are produced by different things: one by deliberate reasoning in a research setting, the other by habit, context, mood, price, and whatever else is in the fridge.
- “I’d definitely buy this.”
- “The pack is easy to open.”
- “I’d have it for breakfast.”
- “The price feels fair.”
- Bought it once, then went back to their usual brand.
- Fought the lid open over the sink, then decanted it into a jar.
- Ate it at 21:40 in front of the television.
- Waited for the promotion.
How far stated intent overshoots
The size of the gap is not a matter of opinion. Four findings are worth knowing, because they are the ones that survive a challenge from a food scientist in the room.
Intentions convert about half the time
Sheeran and Webb’s synthesis of the intention-behavior literature concludes plainly: “the intention-behavior gap is large — current evidence suggests that intentions get translated into action approximately one-half of the time.”[2] Intention is not useless — across 10 meta-analyses covering 422 studies it correlates with later behavior at r = .53, so it explains something like a quarter of the variance[4]. The other three quarters is your risk. And the people who create most of it are not the undecided; they are the ones who genuinely intended to act and then did not — what the literature calls inclined abstainers[5]. Those are precisely the respondents who gave you a top-box score.
Moving intent barely moves behavior
Correlation is not the number that matters to you anyway; you want to know what happens when you change intent. Across 47 experiments that deliberately did so, a medium-to-large change in intention (d = 0.66) produced only a small-to-medium change in actual behavior (d = 0.36)[3]. If your reformulation lifts declared purchase intent, that lift is real. It is just a great deal smaller by the time it reaches a shelf.
Hypothetical money is worth more than real money
Across 28 valuation studies that asked people what they would pay and then made them actually pay, the median ratio of hypothetical to real value was 1.35 — overstatement of roughly a third, on a severely right-skewed distribution, so plenty of individual studies were far worse[6]. An earlier meta-analysis of 29 studies, differently selected, found subjects overstated by a factor of about three on average[7]. In food specifically, methods that involve real money and real product predicted actual retail sales better than hypothetical choice experiments[15].
Intent is weakest exactly where NPD needs it most
Purchase intentions predict sales better for existing products than new ones, for durables than for non-durables, over short horizons than long, and at brand level than category level[8]. Read that list again from an NPD seat: new, non-durable, long horizon. That is a new food product, and it is the worst case on all three counts.
There is no credible peer-reviewed “concept tests overstate trial by X%” figure for food. Every version of that number we could trace led to a vendor blog with no method attached, so it is not on this page — ask anyone who quotes one for the paper. The defensible statement is directional and very well evidenced: stated intent systematically overpredicts behavior, by an amount that varies with category, method and question wording. That alone is enough to change how you make a launch decision.
Why it happens
Five mechanisms do most of the work. They are additive, and a survey triggers all of them at once.
- Social desirability. People shade answers towards what looks good. This is measurable, not theoretical: in a study of 484 adults where true energy expenditure was known from doubly-labelled water, social desirability was among the best predictors of who under-reported what they ate — for both men and women[10]. If people misreport their own dinner against a biological reference, a 7-point purchase-intent scale is not a hard target.
- Hypothetical bias. Nothing is at stake in an answer. Saying yes to a concept costs nothing; buying it costs €4.79 and a slot in the weekly shop[6].
- Context. Food is eaten somewhere, by someone, at a time, alongside something else. Strip that away and you are measuring a different act. The same food scored differently in a restaurant, a laboratory and a cafeteria — identical product, different room[13]. Consumers also rate products higher at home than in a test facility, and for some products the conclusion of the test itself flips with the setting[14].
- Habit and unawareness. Most eating is automatic and cue-driven, so it bypasses intention altogether[2]; habit is one of the strongest predictors of eating behavior, and when behavior is habitual, intentions predict it poorly[17]. Köster’s conclusion after a career in this field is blunt: past behavior, habit and hedonic appreciation predict actual food choice better than attitudes and intentions do[12]. Consumers are not withholding the reason they chose something. They frequently do not have access to it, so they supply a plausible one instead — Köster’s example is asking Germans why they like a coffee and being told, by most of them, that it is “mild”, a word they got from advertising[11].
- The question changes the answerer. Asking someone about their intent measurably strengthens the link between what they then say and what they then do — a self-fulfilling effect of the measurement itself[9]. The act of asking is not neutral.
An identical snack was sold for three weeks in four Dutch works canteens: in two it was labelled new, in two healthy. It sold better as “new” (5.2% of people bought it) than as “healthy” (3.8%). But on the questionnaire, the “healthy” group rated it higher, said they would eat much more of it in future, and between them claimed to have bought about twice as many snacks as had actually been sold. The “new” group overestimated their own consumption by less than 10%. Same product, same weeks — the say-do gap opened purely on framing.[11]
Where traditional methods break down
Every method is good at something. The question is what each one lets you conclude, and where the gap opens.
| Method | Good at | Where the say-do gap opens |
|---|---|---|
| Survey / concept test | Reach, cheap comparison, screening many concepts | Nothing is at stake, so intent is inflated; you learn the score, not the reason |
| Focus group | Language, associations, early positioning | People perform for the room; the loudest opinion becomes the group’s opinion |
| Central location test | Isolating a real sensory difference under control | No kitchen, no household, no second serving; scores shift versus at home, and can change the verdict[14] |
| Home use test | Real context, real preparation, repeat occasions | Traditionally slow, and if it only collects a questionnaire at the end, it inherits recall error |
Notice the last row. Sending product home is necessary but not sufficient. If the only thing that comes back is a form, you have moved the location of the survey without closing the gap. What closes it is capturing the behavior while it happens.
To be fair to the central location test: it is more stable than a home test, and it is the right tool for proving that two formulations genuinely differ. The literature shows that the setting changes the result — it does not on its own prove that home testing predicts the market better. The study that does test against real retail sales is the one comparing hypothetical to non-hypothetical methods, and it comes out the same way: the closer the task is to real behavior, the better it predicts[15].
What behavioral evidence looks like instead
The fix is not better questions. It is to stop treating the answer as the primary evidence.
Behavioral evidence means watching the product get used: unpacked, stored, prepared, served, finished or quietly abandoned. In the consumer’s own kitchen, with their own pans, their own timing and their own family objecting to it. Unscripted, with no moderator steering and no other participants in the room. And across several days, not one sitting — because the second and third use is where most food products actually die.
That is what Eatpol runs. Consumers we recruit to match your target audience use your product at home and record it on their own phone; you get the footage plus the analysis, typically one week from kickoff. You still ask them questions — you just stop relying on the answer alone, because now you can check it against what they did thirty seconds earlier. See how the platform works.
What comes back is not a score. It is a list of specific, fixable things: the step where four of ten people got the preparation wrong, the moment the texture changed their face, the point in the pack where they gave up and reached for scissors.
Using this before a gate review
You are not buying insight. You are buying a decision you can defend.
Peer-reviewed estimates put new-product failure at 50–75% of launches removed from the market well short of their financial targets (industry panels put it as high as 85%)[18]. The same paper argues the cause is institutional: companies systematically under-use behavioral science when they develop and evaluate products. That is the say-do gap, seen from the top of a P&L. In fairness, the failure rate is contested — a later empirical study of actual food launches estimated success rates between 58% and 88% depending on category[19]. Either way, the share of launches that miss their targets is large enough that how you make the go/no-go call is worth arguing about.
At a gate, the question is rarely “did people like it?” It is “what happens when this is on shelf, and how do you know?” A top-two-box number does not survive the follow-up question. Twelve households on video does. So does a ranked list of the friction that cost you repeat purchase, with the clip attached.
- Run it before the gate, not after. The point of finding the problem is to fix it while fixing it is still cheap. After the gate, the same finding is a delay.
- The cost of being wrong is asymmetric. Killing a weak concept costs you a few weeks. Launching one costs the listing, the slot, the trade spend and the year.
- Retailers ask behavioral questions. A category manager wants to know what the shopper does with it and whether they come back — not what a panel said in a hall test.
- Bring evidence a sceptic can watch. Numbers get argued with. Footage of your own target consumer struggling with your own pack ends the argument in about four seconds.
Still deciding what to launch rather than whether to launch it? Eatpol Nova works the other end of the same problem.