Why Your Marketing Mix Model Results Might Be Misleading You

A model running is not the same as a model being right. Marketing Mix Modelling is a powerful analytical tool, but like all statistical methods, it can mislead you if the underlying data is flawed, the configuration is off, or the outputs are read without appropriate scepticism.

This guide explains the three most common ways MMM results mislead practitioners, how to spot a red flag in your outputs before presenting them to a client, and what a well-configured model actually looks like.

In this guide, you’ll learn how impact analysis introduces specific biases, why branded paid search is chronically over-attributed in MMM, and how lift tests remain the gold standard for validating model outputs.


The Difference Between a Model Running and a Model Being Right

The first thing to understand about MMM outputs is that a completed model run does not mean a correct model run. Robyn and other MMM frameworks will always produce outputs — the question is whether those outputs reflect the true structure of your market.

This is not unique to MMM. Any statistical model fitted to real-world data makes assumptions, and when those assumptions are wrong or the data is insufficient, the model’s estimates will be biased. The model won’t tell you it’s misconfigured. It will simply give you numbers.

“The most dangerous MMM outputs are the ones that look plausible but are wrong in ways that are hard to detect without domain knowledge. A model that shows TV contributing 5% of sales for a brand that spent 40% of its budget on TV for three years should raise immediate questions — but it’s a common output when Carry-Over Effect parameters are set too narrowly.”

Understanding how models mislead requires knowing where the structural failure points are. There are three that appear most frequently in practitioner-level MMM work.


Three Ways MMM Results Commonly Mislead

1. Regularisation Bias Toward Sparse Channels

What it is: Robyn uses ridge regression (a form of regularised regression) as its core estimation method. Ridge regression adds a penalty for large coefficients — it “shrinks” estimates toward zero to avoid overfitting. This is mathematically sensible, but it has a systematic consequence: channels with sparse, intermittent spend patterns are penalised more aggressively than channels with consistent, high-volume spend.

What it looks like in your output: A channel where spend has been low, variable, or intermittent (a new channel, a test-and-learn campaign, seasonal TV) will show a lower contribution estimate than a channel with consistent, high spend. This is partly real — low-spend channels have less statistical leverage — but it is also partly an artefact of the regularisation.

What to do about it: Interpret results for any channel accounting for less than 5–8% of total media spend with caution. Treat their contribution estimates as directional, not precise.

2. Branded Paid Search Over-Attribution

What it is: Branded paid search — ads triggered by searches for your own brand name — is one of the most consistently over-attributed channels in MMM. The reason is structural: branded search captures demand that was already created by television, social, outdoor, or organic brand activity. Someone who saw your TV ad and then searched for your brand is being credited entirely to branded search in a click-based model — and to a significant degree in MMM, if the model doesn’t properly account for the upstream channels that created demand.

What it looks like in your output: Branded paid search shows strong, consistent contribution across all time periods, including periods when other media was running at high levels.

What to do about it: The most rigorous mitigation is to run a branded search pause test — temporarily switching off branded paid search in a controlled geography or period to measure how much demand truly disappears. If your model hasn’t been calibrated against any lift test data, treat branded search contribution estimates as upper bounds.

3. Insufficient Data Volume for Key Channels

What it is: MMM is a data-intensive method. For a channel to be modelled reliably, the model needs enough variation in spend over time to observe how sales respond when spend goes up and when it goes down. Channels that launched recently, that run at a constant flat budget, or that have very low total spend simply don’t provide enough information for the model to estimate their contribution accurately.

What it looks like in your output: Channels with insufficient data show very wide uncertainty bands, or produce contribution estimates that are sensitive to small changes in configuration.

What to do about it: Before running your model, audit your channel data:

  • Minimum 20 weeks of active spend for any channel you want modelled reliably
  • Meaningful variation in spend levels — a channel that ran at exactly £10,000 per week for two years provides no variation to model
  • At least 10% of total media budget concentrated in the channel if you want precise estimates rather than directional ones

Channels that fail these tests should be included as context variables if possible, or excluded and noted in the model documentation.


How to Spot a Red Flag in Your Outputs

Before presenting MMM results to a client, run these sanity checks against your outputs.

Sanity Check 1: Does the Organic Base make sense?

The Organic Base is the share of your sales the model attributes to no marketing activity — the demand that would exist even with zero media investment. For most businesses, this should be somewhere between 30–70% of total sales.

If your Organic Base is below 20%, the model is likely over-attributing sales to media channels. If it’s above 80%, the model is likely under-attributing them.

Sanity Check 2: Does Carry-Over duration match channel logic?

Each channel’s Carry-Over Effect estimate should be broadly consistent with what you know about how that channel works:

  • Paid search (direct response): Short carry-over, typically 1–3 weeks
  • Social and video (mid-funnel): Moderate carry-over, typically 3–8 weeks
  • TV and out-of-home (brand building): Longer carry-over, typically 6–12+ weeks

Sanity Check 3: Do the outputs reflect known business events?

Identify three or four specific events in your data history that you know had a material effect on sales — a major promotion, a product launch, a competitor’s entry. Does the model’s output curve show a corresponding spike or trough at the right time?

Sanity Check 4: Are return-on-investment estimates plausible?

Calculate the implied return on media investment (revenue contribution divided by spend) for each channel. These figures should be in a plausible range for your industry and channel type. Cross-reference implied MMM ROI against your historical platform reporting. Large discrepancies in either direction are worth investigating before presenting.


The Role of Calibration: Why Lift Tests Are the Gold Standard

The most rigorous method for validating an MMM output is calibration against a lift test — also called a geo-experiment, matched-panel test, or incrementality test.

In a lift test, you run a controlled experiment: one geography or audience continues to see the media in question, while another is held out as a control. After the test period, you measure the difference in sales between the two groups. That difference represents the true incremental contribution of the media — with actual causal evidence.

The relationship between lift tests and MMM is complementary, not competing:

  • Lift tests tell you the truth about one channel, at one moment in time. They are expensive, slow (typically 4–8 weeks), and limited to testing one variable at a time.
  • MMM gives you a comprehensive view across all channels, historically. But it is correlational, not causal.

When you have lift test results for one or more channels, you can use them to calibrate your MMM — constraining the model’s estimates for those channels to be consistent with your experimental evidence. This dramatically improves confidence in the full model output.

If you haven’t run any lift tests and cannot run them currently, treat your MMM outputs as directional intelligence, not ground truth. They are still highly valuable for identifying relative channel performance — but present them with appropriate confidence levels.


What a Well-Configured Model Looks Like

A checklist for evaluating whether your model output is trustworthy before presenting it.

  • R-squared / model fit above 0.80 — the model explains at least 80% of weekly sales variation
  • Organic Base between 30–70% for a mature, multi-channel advertiser
  • Carry-Over durations consistent with channel type (search: short; TV: long)
  • All major promotions and seasonal events included as context variables
  • No channel contribution estimates below zero for spend-supported channels
  • Implied ROI estimates within a plausible range for your industry
  • Sanity-checked against known business events — model curve matches reality
  • Calibrated against at least one lift test result (preferred; disclose if not)

Frequently Asked Questions

How do I know if my MMM results are accurate?

There is no single test for MMM accuracy, but there are several reliability indicators. Check that the model’s R-squared or equivalent fit statistic is above 0.80, that the Organic Base percentage falls within a plausible range (typically 30–70%), that Carry-Over estimates are consistent with the nature of each channel, and that the output correctly reflects known business events. The most rigorous validation is calibrating the model against a geo lift test.

Why does paid search always show high contribution in MMM?

Paid search — especially branded search — often shows high contribution in MMM because it is the channel that captures purchase intent at the moment of conversion. Some of that contribution is genuine; some is attributable to demand originally created by upstream channels (TV, social, or brand awareness). The best mitigation is a branded search pause test and calibrating your model against the results.

What is an MMM lift test?

A lift test (also called a geo-experiment or incrementality test) is a controlled experiment designed to measure the true causal contribution of a specific media channel. One geography or audience continues to be exposed to the channel; another is held out as a control. The difference in sales outcomes between the two groups represents the genuine incremental effect of the media. Lift test results can be used to calibrate an MMM model, constraining its estimates to be consistent with experimental evidence.

Can MMM results be wrong even if the model runs successfully?

Yes. A completed model run does not guarantee accurate outputs. MMM results can be misleading if: key variables (promotions, seasonality) are missing from the data; channel spend data is inconsistent or incomplete; hyperparameter configurations are poorly calibrated; or a channel has insufficient spend history to be modelled reliably. Always apply sanity checks to your outputs before presenting them.

What is the Organic Base in MMM?

The Organic Base represents the share of your sales that the model attributes to factors other than any media investment — brand equity, word of mouth, distribution, and underlying demand. For most established businesses, this falls between 30–70% of total sales. An Organic Base outside this range is often an indicator that the model is misconfigured or that important variables are missing.

Is MMM more accurate than last-click attribution?

For strategic budget planning across multiple channels, yes — significantly. Last-click attribution assigns 100% of conversion credit to the final touchpoint regardless of what happened earlier in the customer journey. It systematically under-credits upper-funnel channels (TV, video, brand campaigns) and over-credits direct-response channels (branded search, retargeting). MMM accounts for Carry-Over Effects, cross-channel interactions, and external factors that click-based models cannot see.


Understand Your Model Before Trusting It

MMM is one of the most effective tools available for media budget decision-making. Used well, it gives you a perspective on your media mix that no platform can provide. But “used well” means critically evaluating the outputs before acting on them.

If you’re running MMM for the first time and want a structured workflow that validates your configuration and flags potential issues before you present, MMM Pilot includes built-in model quality checks and clear guidance on interpreting each output.

Run a model with built-in quality checks →

Discover more from MMM Pilot™

Subscribe now to keep reading and get access to the full archive.

Continue reading