For innovation & NPD teams

The say-do gap

Short answer

The say-do gap is the distance between what consumers say they will buy and what they actually do. Stated intent is measured in a survey; behavior happens in a kitchen. In food and drink the gap is wide, well documented, and it is why products that test well still fail after launch.

Updated 2 September 2026 · 9 min read · Eatpol, Wageningen

Your gate review runs on what people said. Your P&L runs on what they did.

What the say-do gap is

The say-do gap is the difference between what people tell a researcher they will do and what they actually do afterwards. Market-research practice uses exactly that name[1]; the academic literature calls the same phenomenon the intention-behavior gap, or the attitude-behavior gap, and has been measuring it for decades[2]. Economists describe the money version of it as the difference between stated preference (what you say a thing is worth) and revealed preference (what you hand over at the till).

It is not lying. People answer honestly and then behave differently, because the answer and the behavior are produced by different things: one by deliberate reasoning in a research setting, the other by habit, context, mood, price, and whatever else is in the fridge.

What they said
  • “I’d definitely buy this.”
  • “The pack is easy to open.”
  • “I’d have it for breakfast.”
  • “The price feels fair.”
What they did
  • Bought it once, then went back to their usual brand.
  • Fought the lid open over the sink, then decanted it into a jar.
  • Ate it at 21:40 in front of the television.
  • Waited for the promotion.

How far stated intent overshoots

The size of the gap is not a matter of opinion. Four findings are worth knowing, because they are the ones that survive a challenge from a food scientist in the room.

Intentions convert about half the time

Sheeran and Webb’s synthesis of the intention-behavior literature concludes plainly: “the intention-behavior gap is large — current evidence suggests that intentions get translated into action approximately one-half of the time.”[2] Intention is not useless — across 10 meta-analyses covering 422 studies it correlates with later behavior at r = .53, so it explains something like a quarter of the variance[4]. The other three quarters is your risk. And the people who create most of it are not the undecided; they are the ones who genuinely intended to act and then did not — what the literature calls inclined abstainers[5]. Those are precisely the respondents who gave you a top-box score.

Moving intent barely moves behavior

Correlation is not the number that matters to you anyway; you want to know what happens when you change intent. Across 47 experiments that deliberately did so, a medium-to-large change in intention (d = 0.66) produced only a small-to-medium change in actual behavior (d = 0.36)[3]. If your reformulation lifts declared purchase intent, that lift is real. It is just a great deal smaller by the time it reaches a shelf.

Hypothetical money is worth more than real money

Across 28 valuation studies that asked people what they would pay and then made them actually pay, the median ratio of hypothetical to real value was 1.35 — overstatement of roughly a third, on a severely right-skewed distribution, so plenty of individual studies were far worse[6]. An earlier meta-analysis of 29 studies, differently selected, found subjects overstated by a factor of about three on average[7]. In food specifically, methods that involve real money and real product predicted actual retail sales better than hypothetical choice experiments[15].

Intent is weakest exactly where NPD needs it most

Purchase intentions predict sales better for existing products than new ones, for durables than for non-durables, over short horizons than long, and at brand level than category level[8]. Read that list again from an NPD seat: new, non-durable, long horizon. That is a new food product, and it is the worst case on all three counts.

What we will not claim

There is no credible peer-reviewed “concept tests overstate trial by X%” figure for food. Every version of that number we could trace led to a vendor blog with no method attached, so it is not on this page — ask anyone who quotes one for the paper. The defensible statement is directional and very well evidenced: stated intent systematically overpredicts behavior, by an amount that varies with category, method and question wording. That alone is enough to change how you make a launch decision.

Why it happens

Five mechanisms do most of the work. They are additive, and a survey triggers all of them at once.

  • Social desirability. People shade answers towards what looks good. This is measurable, not theoretical: in a study of 484 adults where true energy expenditure was known from doubly-labelled water, social desirability was among the best predictors of who under-reported what they ate — for both men and women[10]. If people misreport their own dinner against a biological reference, a 7-point purchase-intent scale is not a hard target.
  • Hypothetical bias. Nothing is at stake in an answer. Saying yes to a concept costs nothing; buying it costs €4.79 and a slot in the weekly shop[6].
  • Context. Food is eaten somewhere, by someone, at a time, alongside something else. Strip that away and you are measuring a different act. The same food scored differently in a restaurant, a laboratory and a cafeteria — identical product, different room[13]. Consumers also rate products higher at home than in a test facility, and for some products the conclusion of the test itself flips with the setting[14].
  • Habit and unawareness. Most eating is automatic and cue-driven, so it bypasses intention altogether[2]; habit is one of the strongest predictors of eating behavior, and when behavior is habitual, intentions predict it poorly[17]. Köster’s conclusion after a career in this field is blunt: past behavior, habit and hedonic appreciation predict actual food choice better than attitudes and intentions do[12]. Consumers are not withholding the reason they chose something. They frequently do not have access to it, so they supply a plausible one instead — Köster’s example is asking Germans why they like a coffee and being told, by most of them, that it is “mild”, a word they got from advertising[11].
  • The question changes the answerer. Asking someone about their intent measurably strengthens the link between what they then say and what they then do — a self-fulfilling effect of the measurement itself[9]. The act of asking is not neutral.
The cleanest demonstration we know of

An identical snack was sold for three weeks in four Dutch works canteens: in two it was labelled new, in two healthy. It sold better as “new” (5.2% of people bought it) than as “healthy” (3.8%). But on the questionnaire, the “healthy” group rated it higher, said they would eat much more of it in future, and between them claimed to have bought about twice as many snacks as had actually been sold. The “new” group overestimated their own consumption by less than 10%. Same product, same weeks — the say-do gap opened purely on framing.[11]

Where traditional methods break down

Every method is good at something. The question is what each one lets you conclude, and where the gap opens.

MethodGood atWhere the say-do gap opens
Survey / concept testReach, cheap comparison, screening many conceptsNothing is at stake, so intent is inflated; you learn the score, not the reason
Focus groupLanguage, associations, early positioningPeople perform for the room; the loudest opinion becomes the group’s opinion
Central location testIsolating a real sensory difference under controlNo kitchen, no household, no second serving; scores shift versus at home, and can change the verdict[14]
Home use testReal context, real preparation, repeat occasionsTraditionally slow, and if it only collects a questionnaire at the end, it inherits recall error

Notice the last row. Sending product home is necessary but not sufficient. If the only thing that comes back is a form, you have moved the location of the survey without closing the gap. What closes it is capturing the behavior while it happens.

To be fair to the central location test: it is more stable than a home test, and it is the right tool for proving that two formulations genuinely differ. The literature shows that the setting changes the result — it does not on its own prove that home testing predicts the market better. The study that does test against real retail sales is the one comparing hypothetical to non-hypothetical methods, and it comes out the same way: the closer the task is to real behavior, the better it predicts[15].

What behavioral evidence looks like instead

The fix is not better questions. It is to stop treating the answer as the primary evidence.

Behavioral evidence means watching the product get used: unpacked, stored, prepared, served, finished or quietly abandoned. In the consumer’s own kitchen, with their own pans, their own timing and their own family objecting to it. Unscripted, with no moderator steering and no other participants in the room. And across several days, not one sitting — because the second and third use is where most food products actually die.

That is what Eatpol runs. Consumers we recruit to match your target audience use your product at home and record it on their own phone; you get the footage plus the analysis, typically one week from kickoff. You still ask them questions — you just stop relying on the answer alone, because now you can check it against what they did thirty seconds earlier. See how the platform works.

What comes back is not a score. It is a list of specific, fixable things: the step where four of ten people got the preparation wrong, the moment the texture changed their face, the point in the pack where they gave up and reached for scissors.

Using this before a gate review

You are not buying insight. You are buying a decision you can defend.

Peer-reviewed estimates put new-product failure at 50–75% of launches removed from the market well short of their financial targets (industry panels put it as high as 85%)[18]. The same paper argues the cause is institutional: companies systematically under-use behavioral science when they develop and evaluate products. That is the say-do gap, seen from the top of a P&L. In fairness, the failure rate is contested — a later empirical study of actual food launches estimated success rates between 58% and 88% depending on category[19]. Either way, the share of launches that miss their targets is large enough that how you make the go/no-go call is worth arguing about.

At a gate, the question is rarely “did people like it?” It is “what happens when this is on shelf, and how do you know?” A top-two-box number does not survive the follow-up question. Twelve households on video does. So does a ranked list of the friction that cost you repeat purchase, with the clip attached.

  • Run it before the gate, not after. The point of finding the problem is to fix it while fixing it is still cheap. After the gate, the same finding is a delay.
  • The cost of being wrong is asymmetric. Killing a weak concept costs you a few weeks. Launching one costs the listing, the slot, the trade spend and the year.
  • Retailers ask behavioral questions. A category manager wants to know what the shopper does with it and whether they come back — not what a panel said in a hall test.
  • Bring evidence a sceptic can watch. Numbers get argued with. Footage of your own target consumer struggling with your own pack ends the argument in about four seconds.

Still deciding what to launch rather than whether to launch it? Eatpol Nova works the other end of the same problem.

The part nobody measures
A product usually fails on one small thing — and nobody ever tells you which one.

Most food products that fail are not bad. They fail on something small and specific. The texture after two days in the fridge. A pack that will not reseal. A cook time nobody actually follows.

The consumer notices. They do not complain, they do not fill anything in, and they are not angry about it. They simply do not buy it again — and they never say why. Nobody completes a form to explain why they stopped buying a yoghurt.

So the brand watches a number fall and never learns the reason. Trial was bought with marketing; repeat had to be earned by the product, and the product lost it somewhere between the second and third use, in a kitchen nobody was watching.

This is also the part a single-exposure score cannot see. When people ate the same product at home once a week for ten weeks, boredom rose and acceptance fell over the run — and it fell least when they had variety and choice[16]. A first-taste liking score is a measurement of week one. Your forecast is a claim about week eight.

That lost reason is the real waste. It was observable, it was fixable, and it was gone before anyone thought to look. Closing the say-do gap is, in practice, mostly about being in the room for the second use.

References

Sources

  1. Quirk’s Media. Say-Do Gap, Glossary of Marketing Research Terms. quirks.com/glossary/say-do-gap
  2. Sheeran, P. & Webb, T. L. (2016). The Intention–Behavior Gap. Social and Personality Psychology Compass, 10(9), 503–518. doi.org/10.1111/spc3.12265 — open-access accepted version: eprints.whiterose.ac.uk/107519
  3. Webb, T. L. & Sheeran, P. (2006). Does changing behavioral intentions engender behavior change? A meta-analysis of the experimental evidence. Psychological Bulletin, 132(2), 249–268. doi.org/10.1037/0033-2909.132.2.249
  4. Sheeran, P. (2002). Intention—Behavior Relations: A Conceptual and Empirical Review. European Review of Social Psychology, 12(1), 1–36. doi.org/10.1080/14792772143000003 — 10 meta-analyses, 422 studies, r+ = .53 (as reported in [2]).
  5. Orbell, S. & Sheeran, P. (1998). ‘Inclined abstainers’: A problem for predicting health-related behaviour. British Journal of Social Psychology, 37(2), 151–165. doi.org/10.1111/j.2044-8309.1998.tb01162.x
  6. Murphy, J. J., Allen, P. G., Stevens, T. H. & Weatherhead, D. (2005). A Meta-analysis of Hypothetical Bias in Stated Preference Valuation. Environmental and Resource Economics, 30(3), 313–325. doi.org/10.1007/s10640-004-3332-z
  7. List, J. A. & Gallet, C. A. (2001). What Experimental Protocol Influence Disparities Between Actual and Hypothetical Stated Values? Environmental and Resource Economics, 20(3), 241–254. doi.org/10.1023/A:1012791822804
  8. Morwitz, V. G., Steckel, J. H. & Gupta, A. (2007). When do purchase intentions predict sales? International Journal of Forecasting, 23(3), 347–364. doi.org/10.1016/j.ijforecast.2007.05.015
  9. Chandon, P., Morwitz, V. G. & Reinartz, W. J. (2005). Do Intentions Really Predict Behavior? Self-Generated Validity Effects in Survey Research. Journal of Marketing, 69(2), 1–14. doi.org/10.1509/jmkg.69.2.1.60755
  10. Tooze, J. A., Subar, A. F., Thompson, F. E., Troiano, R., Schatzkin, A. & Kipnis, V. (2004). Psychosocial predictors of energy underreporting in a large doubly labeled water study. The American Journal of Clinical Nutrition, 79(5), 795–804. doi.org/10.1093/ajcn/79.5.795
  11. Köster, E. P. (2003). The psychology of food choice: some often encountered fallacies. Food Quality and Preference, 14(5–6), 359–373. doi.org/10.1016/S0950-3293(03)00017-X — the canteen study is reported here, citing Köster, Beckers & Houben (1987).
  12. Köster, E. P. (2009). Diversity in the determinants of food choice: A psychological perspective. Food Quality and Preference, 20(2), 70–82. doi.org/10.1016/j.foodqual.2007.11.002
  13. Meiselman, H. L., Johnson, J. L., Reeve, W. & Crouch, J. E. (2000). Demonstrations of the influence of the eating environment on food acceptance. Appetite, 35(3), 231–237. doi.org/10.1006/appe.2000.0360
  14. Boutrolle, I., Delarue, J., Arranz, D., Rogeaux, M. & Köster, E. P. (2007). Central location test vs. home use test: Contrasting results depending on product type. Food Quality and Preference, 18(3), 490–499. doi.org/10.1016/j.foodqual.2006.06.003
  15. Chang, J. B., Lusk, J. L. & Norwood, F. B. (2009). How Closely Do Hypothetical Surveys and Laboratory Experiments Predict Field Behavior? American Journal of Agricultural Economics, 91(2), 518–534. doi.org/10.1111/j.1467-8276.2008.01242.x
  16. Zandstra, E. H., de Graaf, C. & van Trijp, H. C. M. (2000). Effects of variety and repeated in-home consumption on product acceptance. Appetite, 35(2), 113–119. doi.org/10.1006/appe.2000.0342
  17. van ’t Riet, J., Sijtsema, S. J., Dagevos, H. & De Bruijn, G.-J. (2011). The importance of habits in eating behaviour. An overview and recommendations for future research. Appetite, 57(3), 585–596. doi.org/10.1016/j.appet.2011.07.010
  18. Dijksterhuis, G. (2016). New product failure: Five potential sources discussed. Trends in Food Science & Technology, 50, 243–248. doi.org/10.1016/j.tifs.2016.01.016
  19. Salnikova, E., Baglione, S. L. & Stanton, J. L. (2019). To Launch or Not to Launch: An Empirical Estimate of New Food Product Success Rate. Journal of Food Products Marketing, 25(7), 771–784. doi.org/10.1080/10454446.2019.1661930
Good to know

Frequently asked.

What is the say-do gap?

The say-do gap is the distance between what consumers say they will buy and what they actually do. Stated intent is measured in a survey; behavior happens in a kitchen. It matters most in food and drink, where choices are habitual, contextual and largely automatic, so what people report in a research setting is a poor guide to what they will put in a basket week after week.

Is the say-do gap the same as the intention-behavior gap?

They describe the same thing in different vocabularies. Say-do gap is the market-research term. Intention-behavior gap is the academic term used in psychology, alongside attitude-behavior gap. Economists frame it as stated preference versus revealed preference. If you are searching for evidence rather than opinion, the academic terms will find far more of it.

How much do consumers overstate what they will buy?

There is no single number that holds across categories, and you should distrust anyone who gives you one. What the evidence supports: intentions translate into action roughly half the time, and across 28 valuation studies the median ratio of hypothetical to real willingness-to-pay was 1.35, with a severely skewed distribution. The direction is consistent and the magnitude is material, but it varies by product, method and question wording.

Why do surveys and focus groups overstate purchase intent?

Four mechanisms stack up. Social desirability pushes answers towards what sounds good. Hypothetical bias means nothing is at stake in saying yes. Context is stripped away, so a test setting measures a different act from a real eating occasion. And most food choice is habitual and automatic, so people often cannot report the reason for it even when they want to. A focus group adds a fifth: people perform for the room.

Can you close the say-do gap by writing better survey questions?

You can narrow it. Non-hypothetical elicitation, where participants commit real money to real product, predicts retail behavior better than hypothetical choice questions. But wording changes cannot reach the part of the problem that is unconscious, habitual or contextual. To get at that you have to observe the behavior instead of asking about it.

What is the difference between stated preference and revealed preference?

Stated preference is what someone says a product is worth or how likely they are to buy it, collected by asking. Revealed preference is what their actual choices show, inferred from what they did. Stated preference is cheap, fast and scalable; revealed preference is harder to collect but is the thing your forecast actually depends on.

How does an in-home test close the gap?

It puts the product where it will really be used, so the context that drives food choice is present rather than removed. Consumers score the same products differently at home than in a test facility, and for some products the conclusion of the test changes with the setting. An in-home test only closes the gap, though, if it captures behavior rather than just collecting a questionnaire at the end. Eatpol records the preparation and the eating on video and analyses what people did, alongside what they said about it.

Why does the say-do gap hit repeat purchase hardest?

Trial can be bought with marketing and distribution. Repeat has to be earned by the product, on the second and third use, when nobody is observing. That is where a small, specific failure lands: a texture that changes after two days, a pack that will not reseal, a cook time nobody follows. Consumers notice, stop buying, and never report the reason, so the brand sees a falling number without a cause attached to it.

How quickly can we get behavioral evidence before a gate review?

Eatpol delivers results about one week from kickoff: brief, recruit consumers matching your target audience, test in their own homes on video, then analysis in your dashboard. Quick consumer video interviews can come back in around two days. That is fast enough to sit inside a stage-gate cycle rather than pushing it, which is the point — the evidence has to arrive while changing the product is still cheap.

Kickoff to results in one week

Stop guessing at
the second purchase.

We put your product in your target consumers’ kitchens, record what they actually do with it, and hand you the reasons behind the number — before your gate review, not after it.