HelloFresh Q4 2024 Senior Product Manager

Personalized Meal Defaults for Factor US

A third of Factor customers never chose their own meals. They got whatever the system assigned — and the system didn't know they couldn't eat pork.

  • A third of Factor US customers received preselected meals each week without making any choices. 30% were actively dissatisfied — the defaults included meals they couldn't eat or didn't want, and ratings for preselected meals had been declining for two years.
  • Designed and ran a three-variant experiment in eight weeks: collected two dietary restrictions (no pork, non-spicy) during signup and customized preselected meals accordingly. Scoped to mobile, two restrictions, seven sub-preferences — narrow enough to validate the mechanism, broad enough to be meaningful.
  • Among customers who selected an exclusion: retention +2.2%, conversion +2%, contributing to $30M+ in estimated customer lifetime value (USD). The broader experiment audience saw retention improve +4.5%, driven primarily by reduced pause rates.
  • Key decision: chose pre-built sub-preference meal strings over a real-time business-rules engine — less flexible, but shippable in weeks. The business-rules approach was the right long-term answer, but the infrastructure didn't exist and the hypothesis needed validating before the investment was justified.

Context

Factor, HelloFresh's ready-to-eat brand in the US, served millions of active subscribers. Customers chose a meal preference during signup — Keto, Protein Plus, Calorie Smart, and so on — and each week received a set of preselected meals based on that preference. If they wanted different meals, they had to go into the app and manually select from over 100 options.

33.5% never did. They received whatever the system assigned. And 30% of those customers were actively dissatisfied with what showed up. Meal ratings for preselected meals had declined over 6% from Q1 2022 to Q2 2024. Cancellation surveys and support tickets cited unwanted preselections as a driver. Customers called out meals they couldn't eat — pork when they didn't eat pork, spicy dishes when they avoided spice — and some assumed the system was already supposed to know this.

This was a separate initiative from the active-experience experimentation program I ran in 2023–2024. That work optimized how subscribers interacted with the product week-to-week. This one targeted what happened before they interacted at all — the default meals assigned to customers who never made a selection.

The longer-term vision was an ML-based personalization engine that would learn from behavior and ratings over time. That was over a year away. The question: could a minimal intervention — two dietary restrictions, collected during signup, applied to default meals — move the retention needle now?

The Problem

The existing preference system was too coarse. A “Protein Plus” customer got the same preset meal string as every other Protein Plus customer. No accounting for whether they ate pork. No accounting for spice tolerance. The system treated all customers within a preference as identical.

This mattered because of how the product worked. Each week, Factor loaded a set of preselected meals into a customer's order. Customers who did meal selection could swap these out. But for the third who didn't, these defaults were the product. And the defaults were wrong often enough that customers were pausing, canceling, or throwing meals away.

Multiple UXR studies showed customers expected their preselections to already reflect their preferences and past orders. They were surprised to learn the system didn't consider any of this. The gap between what customers assumed and what the product actually did was generating friction that showed up in pause rates, cancellation drivers, and declining meal ratings.

Insights

Two restrictions covered the highest-frequency pain points

Roughly 20% of users looked for non-spicy meals and 4% looked for pork-free options. These were the most common unmet restrictions — table-stakes dietary needs that the existing preference system didn't capture. Most Factor customers avoided at least one ingredient when selecting meals manually. Starting with no-pork and non-spicy kept the test scoped while addressing the restrictions that showed up most often in dissatisfaction data. Two restrictions was an explicit scope choice: validate the mechanism first, expand later.

The menu constrained what was possible

Not every preference-restriction combination could produce a credible experience. Keto plus pork-free returned only two unique recipes on some menu weeks — a consequence of shelf-life testing that had forced the Keto menu to rely on chicken recipes containing gelatin (derived from pork). Rather than ship an experience where customers would get the same two meals repeatedly, keto-porkfree was dropped from the test. Seven sub-preferences shipped instead of the originally planned eight. The menu assortment, not the product design, was the binding constraint on scope.

Sub-preferences traded sophistication for speed

Two approaches were on the table. A business-rules engine would evaluate each meal against a customer's restrictions in real time, rank alternatives, and replace ineligible meals dynamically. Pre-built sub-preference meal strings — curated default lists for each preference-restriction combination, created by the Factor menu team — were less flexible but could ship in weeks. The business-rules approach was the right long-term solution but required infrastructure that didn't exist. We chose sub-preferences to validate the hypothesis within eight weeks, knowing the approach wouldn't scale past a handful of restrictions.

Hypotheses

  1. Offering customers the ability to specify dietary exclusions during signup, and reflecting those exclusions in their preselected meals, would improve retention — because customers would stop receiving meals they couldn't eat or didn't want.
  2. The improvement would be driven primarily by reduced pause and cancellation — not by increased meal choice rate or AOV — because the problem was in the defaults experience, not in the selection experience.
  3. Moving the preference question earlier in the signup funnel (grouping it with goals rather than placing it after box size) would improve retention independently of any preselection changes — worth isolating because it could be rolled out to markets where custom preselections weren't feasible.

What We Built

Signup funnel changes

Added two questions to the Factor US mobile signup funnel, placed after the existing “What are your goals?” step. The first captured dietary preferences. The second captured restrictions (No Pork, Non-spicy). The existing preference selection step was removed — the new questions captured equivalent information and mapped to the same backend preferences, with the restriction layered on top. All inputs were saved to the customer's profile for future use.

Custom preselected meals

For each supported preference-restriction combination, the Factor team created a sub-preference with its own default meal string. When the weekly menu loaded for a customer who hadn't yet made selections, the system served their sub-preference string instead of the base preference string. Meals that violated the restriction were replaced with eligible alternatives matching the base preference.

Three-variant experiment

The experiment split the audience into three groups:

  • Control: existing funnel and preselections
  • movePreferenceQ: reordered funnel with preference question moved to follow goals — no change to preselections
  • customPreselections: reordered funnel, restriction questions added, and custom preselected meals served based on inputs

Targeting: users landing on the Goals page of the mobile signup funnel. Allocation ran for approximately five weeks with roughly 630,000 users per variant. Retention data was collected for four weeks post-allocation. Desktop was excluded because the goals/question flow didn't yet exist there.

Results

Retention and revenue

Of the customers who converted in the experiment variant, 38.8% selected a restriction (no pork or non-spicy). This was the cohort actually receiving personalized defaults — and the results were clear:

  • Retention (AOR): +2.2% (2.5 delivered orders vs 2.41 in control)
  • Conversion: +2%
  • Avg net revenue per customer: +1.11%

The full customPreselections variant — including customers who saw the reordered funnel but didn't select a restriction — showed even stronger aggregate numbers: delivered orders +4.5%, gross revenue +2.9%, gross margin +4.1%. AOV and meal choice rate showed no significant impact, which was consistent with the hypothesis: the problem was in the defaults experience, not the selection experience.

Conservative CLV estimate: $30M+ USD, deliberately constrained from higher raw numbers to account for potential confounders.

The funnel reordering drove its own impact

We hypothesized that moving the preference question earlier in the funnel — grouping it with goals rather than placing it after box size — would independently improve retention by giving customers more confidence in the offering. The movePreferenceQ variant was designed to isolate this effect: it reordered the funnel questions but made no changes to preselections. It showed significantly positive results on delivered orders and margin, confirming the hypothesis. The AOR improvement in both variants was driven by reduced pause rates.

Separating the two effects mattered for the rollout decision. For Factor US, the customPreselections variant was rolled out to all mobile users. For other RTE markets where custom preselections weren't yet feasible, movePreferenceQ could be rolled out on its own — it had shown positive impact independently and set those markets up for future personalization.

Margin improvement came from revenue, not cost savings

Initial analysis attributed the margin improvement to cheaper replacement meals — chicken and vegetarian replacing pork. Statsig cost data showed otherwise: costs didn't decrease. The margin improvement came entirely from higher gross revenue driven by more delivered orders. Customers with personalized defaults weren't generating cheaper orders — they were just ordering more often.

What Didn't Work

The menu assortment was the hardest constraint. Keto plus pork-free had to be dropped because some weeks produced only two eligible recipes — not enough for a credible customer experience. The sub-preferences approach hit a ceiling: every new restriction multiplied the number of preset strings that needed manual curation, and the menu wasn't deep enough to support all combinations.

The experiment's target duration was adjusted from 63 to 50 days mid-run. Several metrics — including delivered orders — moved from non-significant to significant at the new threshold. The data team confirmed the change shouldn't have affected results, but it meant extra scrutiny on whether the improvement was real versus an artifact of the shortened window. This is part of why the impact was reported conservatively.

The initiative was scoped to new users only — active customers couldn't update their exclusion preferences through the existing account settings. This kept the MVP contained, but meant the full value of personalized defaults wouldn't be captured until the active experience was updated.

What's Next

The sub-preferences approach validated the hypothesis but was not designed to scale. Every new restriction required manually curated meal strings for each preference, and the menu wasn't deep enough to support all combinations. The next step was the business-rules engine — dynamic replacement logic that could handle arbitrary restrictions without manual curation. The experiment results justified that investment. The profile data collected during signup also established the foundation for broader personalization: the customer preferences captured here became inputs for future recommendation work across RTE brands.

My Role

Shipped the initiative in eight weeks — from problem definition through experiment design, launch, analysis, and rollout decision. Coordinated with the Growth squad for funnel implementation, aligning with their backlog to include the signup flow changes.

  • Scoped the MVP to two restrictions — an explicit choice against pressure to include more attributes. A clean result on two restrictions would justify the investment in the full personalization infrastructure. A noisy result on five restrictions would justify nothing.
  • Chose the sub-preferences approach over a business-rules engine. Sub-preferences could be built and validated within the eight-week timeline. The trade-off was explicit: this wouldn't scale past a few restrictions, but it would answer the question of whether personalized defaults mattered at all.
  • Designed the three-variant experiment structure to separate the impact of funnel reordering from the impact of personalized defaults. This meant the rollout decision wasn't all-or-nothing — markets that couldn't support custom preselections could still ship the funnel change on its own.
  • Dropped keto-porkfree from the test when menu analysis showed it could only return two unique recipes on some weeks. Accepted a reduced test scope to avoid shipping a sub-par experience that would have muddied the data.
  • Navigated the “too good to be true” results. The raw experiment numbers exceeded what the model predicted. Rather than report the upper bound, worked with the analysis team to produce a conservative estimate that accounted for potential confounders.
Back to all work