HelloFresh 2023 – 2024 Senior Product Manager

Rapid Experimentation in Ready-to-Eat

The active experience for millions of RTE subscribers was underperforming, and nobody could agree on what to fix first. Nine experiments, each targeting a specific customer behaviour, settled the debate with data.

  • Six product surfaces, three competing theories about what to fix, and a debate between a large redesign and incremental work. Nobody had the data to prioritise confidently.
  • Ran nine targeted experiments with a survey-paired methodology (a first for the tribe) that produced both metric movement and the qualitative insight to interpret it.
  • $16M+ in cumulative customer value added. The methodology was adopted across other HelloFresh brands.
  • Key decision: chose a bigger Cart MVP than the first experiment required — slower to ship initially, but subsequent experiments built on it instead of starting from scratch.

Context

I owned the active experience for HelloFresh's Ready-to-Eat brands — Factor and Youfoodz — across seven geographies. The active experience is everything a subscriber interacts with between sign-up and churn: My Deliveries, Edit Meals, Add-ons, Settings, Cart, and pricing. It was the primary surface for retention, average order value, and meal choice rate. By early 2023, these metrics had flattened. Retention improvements had stalled. AOV growth was marginal. Meal choice rate was stagnant. The product served millions of active subscribers, and the organisation was debating whether to invest in a large-scale redesign or continue shipping incremental features. Both options had advocates. Neither had evidence.

The Problem

The RTE active experience was broad — six major product surfaces, two brands, seven geographies. Every stakeholder group had a theory about what would move the numbers:

Stakeholder Theory
Commercial team Pricing communication was the bottleneck
Design team Navigation overhaul needed
Brand team Richer recipe content would drive engagement

Each argument had supporting data — correlations, customer feedback, comparisons across markets. None was conclusive. And the risk of a large bundled release was real. A redesign that improved one segment could regress another on a product serving millions of subscribers. I argued against the redesign. The data needed to prioritise well did not exist yet — we were debating which wall to repaint when we hadn't checked which walls were load-bearing. Nine cheap tests would generate that data in months. A single redesign would take just as long and produce one data point. If we were going to redesign, we should know which surfaces actually mattered first. The risk was that stakeholders would lose patience before cumulative signal emerged — no single experiment would be transformative. I committed to transparent read-outs after every cycle so the data could build the case incrementally.

Insights

The metrics were lying by omission

Early experiments produced statistically significant metric movements, but the numbers alone did not explain why. The quick filters experiment showed a -2.3% cancellation rate reduction — but was that because customers found meals faster, or because the filter UI made the menu feel less overwhelming? The first interpretation pointed toward filter specificity. The second pointed toward a different navigation model entirely. I introduced experiment-paired surveys (Sprig integrated with Optimizely) — a first for the RTE tribe. The integration delayed experiment launches by a few days, but the qualitative data changed what we built next. When recipe card enrichment showed improved meal choice rates, the survey revealed that confidence in meal selection — not speed — was what mattered. That distinction redirected our next round of investment. The methodology was adopted across the broader TAM tribe. The Sprig implementation also opened the platform for Pets Table and GoodChop — valuable given limited UXR bandwidth across brands.

Features designed for extension pay off faster

The Cart was the clearest example. The immediate hypothesis was about add-on uptake. But the Cart was built to support future use cases: an Order Confirmation Modal, add-on promotion surfaces, pricing communication. The second and third experiments built on the first instead of starting from scratch. The same principle applied to Favourites. RTE lacked tooling to link recipes across menu weeks, which meant favourite information would not persist week-over-week. Rather than building a narrow workaround, I worked with assortment teams to introduce a new ProductCode field in CCM that persisted across multiple tools and services. Other teams built on this solution, and it became an input for future personalisation.

Hypotheses

  1. A cart experience would increase add-on uptake and AOV by making the purchase flow feel transactional and familiar — reducing the friction of the existing inline add-on interaction.
  2. An order confirmation modal with bulk meal choice would improve active order rate — the hypothesis was that prompting meal selection at the moment of order commitment (when attention is highest) would outperform the default flow where meal selection happened separately.
  3. Quick filters and goal-focused collections would reduce cancellation by addressing the cognitive load of browsing 100+ meals — the bet was that customers who could not find meals matching their dietary goals were churning, not that they disliked the food.
  4. A favourites feature would reduce pause rates by making it easier for returning customers to resurface previously enjoyed meals — the insight was that pausing correlated with the effort of re-discovering meals after a break.
  5. Enriched recipe cards would improve meal choice rate and re-engagement — the hypothesis was that confidence in meal selection was as important as convenience.

What We Built

Nine experiments shipped across 18 months, each targeting a specific customer behaviour with a clear success metric. The programme ran on two-week experiment cycles: ship, measure, survey, decide.

The Cart

The Cart was the foundation. RTE needed its own cart experience — HelloFresh's Editable Order Summary did not meet RTE use cases. I designed it as an extensible surface, not a standalone feature, and incorporated a 'Time Saved' element showing customers how much time they saved using Factor — consolidating messaging that had previously lived in separate banners (banner overload was a top consumer pain point identified through UXR). The Order Confirmation Modal followed, adding bulk meal choice at the point of order commitment. This was an extension point we had scoped into the Cart from the beginning — the Cart was always meant to support surfaces like this. The Cart and its extensions became the highest-value initiative in the programme.

Discovery and navigation

Quick filters and goal-focused meal collections tackled the cognitive load of browsing 100+ meals per week with no structured way to find what matched a dietary goal. Favourites addressed a different problem: customers returning after a pause had to rediscover the menu from scratch. The ProductCode fix made favourites persist across weeks — a systemic solution rather than a patch. Sorting meals initially failed on Youfoodz in 2023. The menu was too small and uniform for sort parameters to surface meaningfully different results. A retest six months later with a larger, more varied menu worked. Learnings from these experiments were shared with the HelloFresh Browse squad, enabling their own iterations on meal discovery — one squad's experiment data informing another's roadmap.

Content and engagement

Enhanced recipe cards with nutritional detail, preparation context, and ingredient transparency. The paired survey revealed that confidence — not speed — drove the meal choice improvement. Customers who felt uncertain about what they were ordering skipped meal selection entirely, which meant lower engagement and higher pause rates. One-click deselect removed friction from changing previous selections. After launch on RTE, GoodChop adopted the feature for their customers. Banner consolidation combined the convenience banner with the seamless prompt, reducing banner overload while preserving the messaging that mattered.

Results

The big numbers

UXUM score across the active experience: 8.1. CVA was not calculated for every experiment — some were measured on engagement and retention metrics rather than modelled for lifetime value. The $16M is concentrated in the Cart and discovery themes.

Theme Key result
Cart ~$11M CVA — AOV up, add-on uptake +3.2%, meal choice rate up across both brands
Discovery and navigation ~$5M CVA — cancellation -2.3%, pause rate down
Content and engagement Meal choice rate up, unpause +1.98%, net revenue per customer up
Total $16M+ in cumulative customer value added

What the data settled

The programme resolved the original prioritisation debate empirically. Navigation and discovery drove the largest retention impact — cancellation and pause rate. The Cart drove the largest revenue impact — AOV, add-on uptake, and pricing communication (which was built into the Cart experience). Recipe content improved engagement and re-activation. None of the three original theories was right on its own. Multiple levers contributed, with discovery and the Cart producing the highest returns in different dimensions.

What Didn't Work

Sorting failed on smaller menus. The first attempt on Youfoodz did not produce meaningful results because the menu was not large or varied enough for sort parameters to matter. A retest six months later with a larger menu succeeded — but the initial failure cost a full experiment cycle and forced us to revisit the hypothesis.

Not every experiment produced a clean win. Several variants improved one metric while regressing another. A filter change that improved meal choice rate also increased time-to-complete. The survey-paired methodology helped diagnose these trade-offs but added lead time: each read-out took longer because we waited for qualitative data before making ship/kill decisions. In a programme designed for speed, the methodology sometimes worked against the cadence.

One-click deselect shipped with a bug. After launch, I identified a spike in complaints about meal selections being reset — traced to a bug in the experiment. The team fixed it, relaunched, and put better testing and monitoring in place. We documented improved testing steps as part of ways of working, but the incident consumed time and eroded initial confidence in the feature.

Cross-pollination created confounding effects. When the Browse squad used our filter experiment data to inform their own roadmap, changes they shipped affected the baseline our subsequent experiments were measured against. We did not anticipate this feedback loop. Some later read-outs required re-analysis to account for confounding changes on shared surfaces.

What's Next

The experiment programme established both a results foundation and a methodology that outlasted the individual experiments. The Sprig + Optimizely survey-paired approach was adopted across the TAM tribe and paved the way for other brands to use it. The Cart supported subsequent work on promotional surfaces and pricing communication. The ProductCode field introduced for Favourites became infrastructure for personalisation. GoodChop adopted the one-click deselect feature after RTE validated it. Most importantly, the programme produced a prioritised map of what actually drove customer behaviour on the active experience — data that directly informed the Factor Weight-Loss Programme's active experience decisions the following year.

My Role

Sole product owner for one squad (5-8 developers plus designer). This was foundational work that built the judgment and methodology I later applied at a broader scope.

  • Chose experimentation over the redesign proposal. The information needed to redesign well did not exist yet. I argued that nine cheap tests would generate it faster than one expensive bet, and that if we were going to redesign, we should know which surfaces actually mattered first. The risk was that stakeholders would lose patience before the cumulative signal emerged. I managed this by sharing results transparently after each cycle and letting the data build the case.
  • Chose a bigger Cart MVP than the first experiment required. Slower start — but subsequent experiments built on it instead of starting from scratch.
  • Introduced the survey-paired methodology for the tribe. Spearheaded the Sprig + Optimizely integration and pushed for adoption across brands. The speed trade-off was real — experiments launched days later. But the qualitative data changed what we built, not just whether we shipped. The recipe card enrichment survey, for instance, showed that confidence was the driver — not speed — which redirected subsequent investment.
  • Solved systemic tooling gaps instead of working around them. The Favourites ProductCode fix took longer than a workaround but created infrastructure other teams used. Identifying and prioritising improved sold-out error messaging (spotted through Usabilla monitoring during capacity-constrained weeks) fixed a live pain point before it scaled.
  • Managed alignment across physical product, assortment, brand, and commercial teams. Each group had opinions about which experiment should run first and whether results justified scaling. The two-week cadence and transparent read-outs provided structure, but alignment still required active management — particularly when experiments contradicted a stakeholder's preferred theory.
Back to all work