In this article
Why test at all — and why structure matters
Most new food products don't make it: peer-reviewed estimates put failure at 50–75% within two years of launch. The pattern behind those failures is remarkably consistent — decisions made on enthusiasm, internal tastings and stated intent rather than on how real consumers behave with the product at home. That distance between what people say and what they do is the say-do gap, and a good test plan is designed to close it.
Structure beats volume here. One well-designed test with the right 80 consumers tells you more than five improvised rounds with whoever was available.
The six steps
Step 1 — Define the question and the success criteria
Before anything ships, write down what the test must prove: Which of the three recipes wins? Is liking at or above our action standard? Would category buyers purchase it at the target price? Set the go/no-go threshold now — deciding what "good enough" means after seeing the results is how weak products slip through.
Step 2 — Choose the method that fits your stage
Concept feedback, recipe comparison and launch validation are different questions with different tools; the table below maps them. If you are still choosing what to build at all, an ideation tool like Eatpol Nova comes before any product test.
Step 3 — Recruit real target consumers
The single most common shortcut — testing with colleagues, friends or family — is also the most expensive one. They aren't screened for category usage, they know you, and they want to be kind. Recruit consumers who actually buy in your category, screened on consumption habits, with no relationship to the team.
Step 4 — Run the test in a realistic context
Food is chosen, cooked and eaten at home — so validation belongs there. Ship the product to consumers and let them prepare it as they normally would: their pans, their kitchen, their family at the table. Standardize the instructions, then let real life happen. For the method comparison (facility vs. home), see consumer research for food products.
Step 5 — Measure behavior, not just opinions
Collect liking and purchase-intent scores — they are comparable and necessary. But weigh observed behavior more heavily: what people actually cook, how they take the first bite, whether the plate gets finished, what they say spontaneously on camera. Behavior predicts the repeat purchase; scores predict the survey.
Step 6 — Decide against the criteria, and iterate
Compare the results with the thresholds from step 1 and make the call: kill, fix or launch. After every meaningful recipe or packaging change, re-test — a product that won six months ago has not necessarily survived its own cost optimization.
Which test at which stage?
| Stage | Question | Fitting method | Typical scale |
|---|---|---|---|
| Ideation | Which concepts resonate? | Video interviews, concept screens | 15–50 consumers |
| Formulation | Which recipe wins? | Sensory analysis (discrimination, descriptive) | 8–12 trained / 60+ consumers |
| Validation | Will they buy and re-use it? | In-home use test (IHUT) | 60–100+ consumers |
| Pre-launch | Does the full proposition land? | Branded IHUT with price exposure | 60–100+ consumers |
The five most common mistakes
- Testing too late. If the first consumer test happens after production is committed, the test can only confirm or embarrass — it can no longer steer.
- Convenience panels. Colleagues and friends predict office enthusiasm, not shelf performance.
- Only asking, never observing. Stated purchase intent is systematically optimistic; watch what people do with the product.
- No benchmark. A liking score means little in isolation — test against the category leader in the same round.
- Moving the goalposts. Success criteria defined after the results arrive always fit the results.
What does it cost, and how long does it take?
Costs scale with participants and logistics: qualitative rounds are cheapest, facility studies (CLTs) carry venue and staffing costs, and in-home tests price mostly per participant including shipping. Methods that use the consumer's own smartphone keep fixed costs low.
On timing: facility-based rounds typically take 4 to 8 weeks each. Eatpol — a spin-off of Wageningen University & Research — recruits from its own tester community and analyzes in-home video with AI, delivering a full round in about one week. At that pace, testing twice before launch stops being a schedule problem.