How to Choose a CPG Marketing Agency: A Buyer’s Evidence Framework

A buyer’s framework for choosing a CPG marketing agency that can connect media, retail availability, creative claims, measurement, and commercial decisions.

The short answer: choose the agency that can solve your commercial constraint and prove it

A good CPG marketing agency is chosen by its ability to solve your brand's specific commercial constraint and to produce evidence that finance, sales, retail, and marketing can reconcile. It is not chosen by the longest channel list, the largest claimed ROAS, or a case-study deck that shows outcomes without showing method.

That distinction matters more in consumer packaged goods than in most categories. A food or beverage brand does not convert demand in one place. A shopper may see a paid social video, search the brand, buy on a retailer app, or walk into a store where the product may or may not be on the shelf. Media can be executed well and still fail commercially because distribution slipped, a promotion overlapped, inventory ran out, or the claim on the creative could not survive review.

So the buying decision is not "who runs ads best." It is: which partner can connect media decisions to product availability, retail and DTC economics, defensible creative claims, and measurement that survives scrutiny?

This is a buyer-side framework. It covers how to define the job before the RFP, six evaluation gates with evidence to request, how to rank the evidence you receive, how to compare fee models by incentive rather than price, a scorecard structure you control, and a diligence workflow that exposes contradictions early.

What a CPG marketing agency actually does

A CPG marketing agency plans and runs commercial demand work for a physical product sold through retail, DTC, or both. Depending on scope, that includes paid media planning and buying, retail media management, shopper and in-store activation, creative and content production, brand and packaging communication support, influencer and creator programs, and measurement.

The important structural point: no single one of those activities is the product. The product is a decision system — what to spend, where, on which SKUs, in which markets, with what creative, judged against which outcome.

Partner types differ substantially:

Partner type Core job Typically strong at Typically weak or out of scope ------------ General performance agency Efficient paid media across channels Platform execution, DTC demand capture, creative iteration cadence Retail data, shelf and distribution reality, retailer-specific mechanics CPG paid-media agency Paid media for a physical product with retail context Connecting media to velocity and availability, SKU-level planning Deep in-store execution, heavy production, advanced causal modeling Shopper / activation agency Convert in and around the store Retailer programs, displays, sampling, local activation Always-on digital performance, cohort-level DTC economics Retail-media specialist Manage on-platform retailer ad networks Retailer ad mechanics, share-of-shelf digital, promo alignment Off-platform demand creation, total-market causal reads Creative partner Produce the assets and messaging Concept quality, production volume, format coverage Buying, measurement, commercial reconciliation Measurement partner Establish evidence and causality Experiment design, modeling, source reconciliation Execution, day-to-day media operations

Most emerging brands do not need all six. Most established brands cannot get all six at equal quality from one firm.

When one lead agency plus specialists beats a single full-service claim

A single partner is usually the right answer when the scope is narrow and the coordination cost of multiple vendors exceeds the capability gain — for example, a DTC-led brand with limited retail distribution, one primary retailer, and a single-market focus.

A lead agency plus specialists becomes more credible when any of the following is true:

retail media, in-store activation, and off-platform demand all carry real budget;

measurement needs to be independent of the party spending the money;

production volume is high enough that creative becomes a supply problem rather than a planning problem;

multiple retailers have materially different mechanics and reporting.

Be direct in the evaluation: ask a full-service firm which of the six roles it performs itself, which it subcontracts, and which it would advise you not to buy from it. An agency that names its own boundary is usually more reliable than one that claims every capability.

Write a selection contract before you write the RFP

Most bad agency selections start with an unclear brief. Every agency then answers a slightly different question, and the comparison becomes a comparison of writing quality.

Before any outreach, write a one-page selection contract and send the identical version to every candidate. It should state:

Product and SKU scope. Which products are in scope, which are strategic, which are being deprioritized.

Retailer and DTC mix. Which accounts matter, what share of revenue each represents in broad terms, and where growth is expected.

Distribution and inventory constraints. Store counts, planned expansion, known supply limits, seasonal capacity.

Margin and contribution definition. What "profitable" means to you — gross margin, contribution after trade and fulfillment, or another internal standard.

Promotion calendar. Known price events, retailer programs, and periods where measurement will be confounded.

The buying decision. What you are actually buying: strategy, execution, measurement, production, or a combination.

Data available. What exists today — retailer reports, sales exports, media exports, first-party event data — and what does not.

Markets. Geographies in scope, and whether performance expectations differ by market.

Legal and claim-review owner. Who approves product claims, and how long review takes.

The 90-day decision. What you intend to decide at the end of the first quarter of the engagement.

That last item is the one buyers most often omit and the one that most improves the answers. If the agency knows the decision it must help you make, its proposal becomes testable.

A written selection contract also protects you commercially. It defines the scope you will compare fees against, so change-control conversations later are about documented deviation rather than memory.

Six evaluation gates

Score each candidate against the same six gates. For every gate, you are looking for four things: the evidence to request, a strong-answer pattern, a warning sign, and a disqualifying condition.

Gate 1 — Commercial-model fit

Does the agency understand how your brand makes money?

Evidence to request: a written interpretation of your margin structure and how it would change media decisions; an explanation of which SKUs they would prioritize and why; the efficiency metric they would manage to, and at which level of the business.

Strong-answer pattern: the agency distinguishes revenue from contribution, asks about trade spend, returns, fulfillment, and retailer terms, and proposes a business-level efficiency control rather than optimizing every campaign to its own ROAS. If you use a blended control, a partner should be comfortable working inside a framework like MER versus ROAS budget control rather than defending platform-reported numbers in isolation.

Warning sign: the proposal quotes a target ROAS before asking what a unit contributes.

Disqualifying condition: the agency cannot explain how its recommendations would differ between a high-margin and low-margin SKU.

Gate 2 — Distribution and retail-readiness fit

Can the agency reason about availability, not just demand?

Evidence to request: how they would handle a market where demand rises but distribution is thin; how they check product-page or store availability before scaling spend; how they treat out-of-stock periods in reporting.

Strong-answer pattern: the agency treats availability as a gating condition on spend, proposes market- or retailer-level pacing tied to distribution, and asks for store counts and velocity data before proposing a budget shape. Market-level thinking of this kind is the mechanic covered in the hyper-local activation playbook.

Warning sign: national budget recommendations with no reference to where the product can actually be bought.

Disqualifying condition: the agency proposes scaling into markets it has not checked for availability, and treats out-of-stock as a reporting footnote rather than a spend decision.

Gate 3 — Measurement and source-of-truth discipline

Can the agency tell you what it does not know?

Evidence to request: a proposed measurement map naming the source of truth for each decision; how it reconciles retailer reporting, platform reporting, and internal sales; its position on attributed versus incremental sales; a redacted example of a reporting artifact.

Strong-answer pattern: the agency separates operating metrics from proof metrics, discloses attribution windows and product scope explicitly, and proposes experiments for causal questions. Industry standards exist for exactly this: the IAB/MRC Retail Media Measurement Guidelines call for transparent definitions and methodologies, disclosed attribution windows, defined product scope, and explicit reporting of measurement limitations. A partner who already works that way will not be surprised by the request. The deeper architecture is covered in the CPG paid media measurement guide.

Warning sign: "closed-loop measurement" used as a synonym for proof. Closed-loop attribution connects an exposure to a recorded sale inside one system; it does not establish that the sale would not have happened anyway. The IAB and IAB Europe Guidelines for Incremental Measurement in Commerce Media frame incrementality around credible counterfactuals, bias control, and separating signal from noise — which is a different exercise from matching exposures to transactions.

Disqualifying condition: the agency cannot name a single condition under which its own reporting would be unreliable.

Gate 4 — Channel and activation fit

Does the channel plan follow the business job, or the agency's inventory of skills?

Evidence to request: the reasoning behind the proposed channel mix; how retail media and off-platform media would be coordinated; how in-store activity would be defined and measured if it is in scope.

Strong-answer pattern: channels are proposed as answers to named jobs — trial, repeat, retailer-specific conversion, new-market entry — with an explicit statement of what each channel cannot do. If in-store media is included, the agency defines formats, zones, and outcomes explicitly; the IAB and IAB Europe In-Store Retail Media definitions and measurement standards exist because those definitions are not self-evident, and a partner that skips them cannot report in-store work consistently.

Warning sign: an identical channel mix appearing across multiple client examples.

Disqualifying condition: the agency proposes a channel it cannot describe the measurement approach for.

Gate 5 — Creative, claims, and compliance workflow

Can the agency produce claims your legal owner can approve, at the volume the media plan requires?

Evidence to request: its claim-substantiation workflow; how it briefs creators and handles disclosures; a redacted example of a rejected claim and what replaced it; the production cadence it can sustain.

Strong-answer pattern: the agency has a documented review path, separates substantiated product claims from lifestyle messaging, and builds legal review time into the production calendar. For food-adjacent brands, objective health or safety claims in advertising must be truthful, non-misleading, and adequately substantiated — the FTC's Health Products Compliance Guidance sets out the expectation. For creator work, material connections between brand and endorser require clear, hard-to-miss disclosure, as the FTC's Disclosures 101 for Social Media Influencers explains.

Warning sign: claim strength treated as a creative-performance lever rather than a compliance question.

Disqualifying condition: the agency proposes health or efficacy claims without asking what substantiation exists, or treats creator disclosure as optional or aesthetic.

Gate 6 — Operating model, decision rights, and economics

Who does the work, who decides, and how does the commercial relationship behave when things change?

Evidence to request: named team members and their actual allocation; escalation path; meeting and reporting cadence; change-control terms; fee structure and what triggers a fee change; offboarding and data-ownership terms.

Strong-answer pattern: the people in the pitch are the people on the account, decision rights are written down, and data, accounts, and creative assets remain yours.

Warning sign: senior presence in the pitch that disappears from the staffing plan.

Disqualifying condition: the agency will not confirm that you own the ad accounts, the first-party data, and the creative assets produced under the engagement.

Rank the evidence you receive

Not all evidence carries the same weight. Use an explicit hierarchy so that a confident assertion never outranks a verifiable export.

Tier Evidence type How much weight Why ------------ 1 Verified first-party exports and retailer data Highest Reflects actual sales and availability, independent of the agency's narrative 2 Documented platform reports with stated windows and scope High, but bounded Real data, but attributed and self-reported by the selling platform 3 Controlled experiments with documented design High for causal questions Tests a counterfactual rather than a correlation 4 Modeled estimates with disclosed assumptions Moderate Useful for portfolio questions; sensitive to inputs and structure 5 Directional proxies and benchmarks Low Context only; rarely transfers across brands and categories 6 Unsupported pitch claims None Cannot be checked

Two practical implications.

First, "closed-loop" is not incrementality. A platform can correctly report that exposed shoppers purchased, and still not tell you whether the purchase was caused. Platform documentation is often explicit about its own scope: Google Ads store-sales measurement is subject to eligibility, geography, volume, and account conditions, and reported store-sales conversions can be delayed and depend on the configured reporting setup. Those are legitimate signals, not causal proof.

Second, experiments are the cleanest answer available to most brands, and some retail platforms document their own randomized approaches. Instacart, for example, has published a description of randomized holdout methodology for estimating incremental sales within its own platform. That is a documented first-party method with a defined scope — not evidence that every retail-media network measures the same way. If a candidate proposes experiments, ask what design and what geographies; the mechanics are covered in the geo-holdout guide.

For modeled evidence, ask what data the model needs. Google's Meridian documentation on collecting and organizing modeling data describes aligned media, spend, control variables, KPI, time, and preferably geographic granularity. If the agency proposes marketing mix modeling while your data cannot meet those input requirements, the proposal is aspirational.

Match partner type to the job

No partner type wins universally. The right answer depends on the job you are buying.

Business job General performance CPG paid media Shopper / activation Retail-media specialist Creative partner Measurement partner --------------------- DTC demand capture Strong fit Strong fit Weak fit Weak fit Support role Support role Retailer-specific activation Weak fit Moderate fit Strong fit Strong fit Support role Support role New-market launch Moderate fit Strong fit Strong fit Moderate fit Support role Support role Retail sell-through measurement Weak fit Moderate fit Weak fit Moderate fit Not applicable Strong fit Creator and UGC production Moderate fit Moderate fit Moderate fit Weak fit Strong fit Not applicable Incrementality evidence Weak fit Moderate fit Weak fit Weak fit Not applicable Strong fit

Read the matrix as a scoping tool, not a ranking. A brand whose constraint is retail measurement should not hire the best creative shop and hope measurement follows. A brand whose constraint is creative supply should not buy a measurement-heavy engagement and wonder why output did not increase.

Compare fee models by incentive, not by price

Fee structure changes behavior. Evaluate each model on four dimensions: incentive alignment, scope clarity, change-control behavior, and measurement risk.

Fee model Incentive it creates Scope behavior Change-control risk Measurement risk --------------- Fixed retainer Stable service, no push to inflate spend Needs explicit scope, or drift accumulates Moderate — disputes arise over "extra" work Low Percentage of spend Rewards larger budgets Scope often loosely defined Low on paper, but budget growth becomes an implicit goal Moderate — efficiency may be argued in favor of scale Project fee Delivers a defined artifact Very clear Low, if deliverables are specific Low, but no ongoing accountability Performance component Rewards a named outcome Requires precise definitions High — disputes over qualifying outcomes High — creates pressure on the measurement source Hybrid Balances base service with outcome interest Depends entirely on drafting quality Moderate Moderate to high, depending on the outcome metric

Two cautions. Any performance component needs a measurement source neither party can unilaterally interpret, agreed before signature; otherwise you have converted a commercial relationship into a reporting argument. And a lower fee is not automatically better value — the correct comparison is total cost of the scope you actually need, normalized across candidates.

This guide publishes no fee ranges or benchmarks. Market rates vary by scope, category, market, and team seniority, and any number quoted as universal should be treated as marketing rather than information.

Build a scorecard you control

Use a weighted scorecard with three components: criteria weights, an evidence-confidence modifier, and a disqualifier override.

Criteria. Start with the six gates. Add any criterion specific to your situation.

Weights. You choose them. A brand whose constraint is measurement weights Gate 3 heavily; a brand entering a new retailer weights Gate 2. There is no universal weighting, and anyone who supplies one is guessing at your business.

Score. Rate each gate 1–5 on the evidence provided.

Evidence-confidence modifier. Multiply each gate score by a confidence factor based on the evidence tier: verified artifacts 1.0, documented platform reports 0.9, described-but-unshown process 0.7, assertion only 0.4.

Disqualifier override. If any gate hits a disqualifying condition, the candidate is out regardless of total score. Do not average away a structural failure.

An illustrative arithmetic example

The numbers below are fictional and exist only to show the mechanics. They are not benchmarks, not market data, and not Sharply Labs results.

A buyer sets weights: commercial model 25, retail readiness 20, measurement 25, channel fit 10, creative/claims 10, operating model 10.

Gate Weight Agency A raw / conf Agency B raw / conf Agency C raw / conf --------------- Commercial model 25 4 / 1.0 5 / 0.4 3 / 0.9 Retail readiness 20 4 / 0.9 3 / 0.7 4 / 1.0 Measurement 25 5 / 1.0 4 / 0.4 3 / 0.7 Channel fit 10 3 / 0.9 5 / 0.7 4 / 0.9 Creative / claims 10 3 / 0.7 4 / 0.9 5 / 1.0 Operating model 10 4 / 1.0 3 / 0.7 4 / 0.9

Weighted confidence-adjusted totals: Agency A = 25(4.0) + 20(3.6) + 25(5.0) + 10(2.7) + 10(2.1) + 10(4.0) = 100 + 72 + 125 + 27 + 21 + 40 = 385. Agency B = 50 + 42 + 40 + 35 + 36 + 21 = 224. Agency C = 67.5 + 80 + 52.5 + 36 + 50 + 36 = 322.

Agency B presented the most confident narrative and scored lowest, because almost nothing was evidenced. That is the entire point of the confidence modifier.

A diligence workflow that surfaces contradictions

Send the identical scoped brief — the selection contract — to every candidate, with the same deadline and the same question list.

Run structured calls with a fixed agenda, same questions, same order, so answers are comparable rather than charismatic.

Request redacted artifacts: a reporting template, a measurement map, a creative brief, a test design, a QA checklist. Process evidence is more informative than outcome slides.

Test how each partner handles unavailable data. Tell them one dataset they expect does not exist. A strong partner narrows the claim or proposes a way to build the data. A weak one proceeds unchanged.

Run a claims and compliance scenario. Give a realistic product claim and ask what substantiation they would require and how they would rewrite it if substantiation is thin.

Normalize fees and scope into one comparison table: what is included, what is billed separately, what triggers a change order.

Check references with exact questions: what did the agency get wrong; how was it handled; what did they decline to do; how did reporting change when results were poor; who actually did the work.

Document contradictions between the pitch, the artifacts, and the references. Contradictions, not weaknesses, are the strongest negative signal.

If uncertainty remains, buy a bounded diagnostic sprint — a paid, scoped, time-limited piece of work with a defined deliverable — rather than signing a long agreement to resolve doubt.

What the agency should ask you for before launch

A serious partner's data request is itself evidence of competence. Expect them to ask for:

SKU-level retailer availability and distribution data;

sales, velocity, and inventory data at the grain you can provide;

the promotion and price-change calendar;

historical media exports with spend, impressions, and platform-reported outcomes;

creative history, including what was tested and what was retired;

site and retailer product-page performance;

first-party event definitions and how they are currently implemented;

contribution constraints and the internal definition of profitable;

privacy-safe access to the above, under documented permissions.

Two rules apply on your side. Never share raw personal data with a prospective partner, and never share it at all where hashed, aggregated, or permissioned access is sufficient. Conversion-signal work should be implemented through documented server-side paths with proper consent, which is the subject of the server-side tracking guide.

A 30-day operating-model test

If you want to reduce risk before a long commitment, structure the first 30 days as a test of collaboration and evidence quality — explicitly not a test of performance.

Days 1–5: definitions and access. Agree outcome definitions, the product and geography scope, and permissioned access. Failure mode: definitions still unsettled at day five.

Days 6–12: measurement map. The agency documents which source answers which question, with windows, scope, and known limitations. Failure mode: a tool list instead of a decision map.

Days 13–18: creative and claims workflow. One brief moves through production and legal review end to end. Failure mode: the review path is discovered rather than planned.

Days 19–24: first hypothesis and test. One documented hypothesis, one design, one pre-agreed read. Failure mode: a test whose result cannot change a decision.

Days 25–30: reporting cadence and decision rule. A recurring review with a written rule for what triggers a spend, creative, or scope change. Failure mode: reporting that describes activity without recommending an action.

Thirty days is enough to see how a partner behaves, reasons, and reports. It is not enough to judge commercial outcomes, and no honest partner will claim otherwise.

Direct answers to common buyer questions

What does a CPG marketing agency do? It plans and executes demand work for a physical product across retail, DTC, or both — media, retail media, activation, creative, and measurement — and translates those activities into commercial decisions about spend, SKUs, and markets.

When should a food brand use a CPG specialist rather than a general performance agency? When the constraint involves retail: distribution-dependent demand, retailer-specific mechanics, velocity and inventory interaction, or reconciliation between retail and DTC. When the constraint is purely DTC demand capture and creative iteration, a strong general performance partner may be the better fit. The category context for food-specific work sits on the food and CPG growth page.

What should be in a CPG agency RFP? The selection contract above: product scope, retailer and DTC mix, distribution constraints, margin definition, promotion calendar, available data, markets, claim-review ownership, and the 90-day decision — plus the artifacts you want returned and the criteria you will score.

How can a buyer compare case studies fairly? Ask the same four questions of every case: what was the baseline, what changed besides media, what measurement source produced the number, and what would have happened without the work. Cases that cannot answer question four are testimonials, not evidence.

Should one agency manage both retail media and paid social? It can, when coordination value is high and both capabilities are genuinely present. Ask which team runs each, whether reporting is unified at the business level, and how conflicts between retailer and off-platform budgets are resolved. Consolidation is only an advantage if it produces one coherent decision.

Which data should the agency request? The list in the previous section. A partner that asks only for ad-account access is planning to optimize platform metrics.

How should fees be compared? Normalize scope first, then compare total cost of that scope, then evaluate the incentive each structure creates. Price alone is the least informative dimension.

What are the clearest red flags? Guaranteed outcomes; ROAS promises before seeing your data; "closed-loop" presented as proof of causality; a pitch team that is not the delivery team; refusal to confirm your ownership of accounts and data; unwillingness to name a condition under which their reporting would be wrong.

Can an agency guarantee retail lift, ROAS, or distribution growth? No. A partner can improve execution quality, evidence quality, and decision speed. Commercial outcomes depend on product, price, availability, retailer decisions, competition, and demand — none of which an agency controls. Treat any guarantee as a disqualifying signal.

When this framework does not apply

This evaluation approach assumes a brand with something to measure and someone to decide. It is the wrong framework when:

the product is pre-distribution, and the real job is retailer acquisition rather than demand generation;

inventory is unstable, so scaling demand creates a service failure rather than growth;

no substantiated claims exist yet, and the required work is substantiation, not advertising;

no internal decision owner exists, so recommendations will have nowhere to land;

available data cannot support the analysis being promised, in which case the first engagement should build the data foundation.

In several of those cases the correct next purchase is a bounded diagnostic or data-foundation project, not a full-service retainer.

Minimum viable evaluation pack

Assemble this before the first call. None of it requires exposing personal data.

One-page selection contract, as defined above.

SKU list with strategic priority and margin tier (relative tiers are enough).

Retailer and DTC revenue mix, expressed in broad shares.

Distribution snapshot: store counts or weighted distribution, plus planned changes.

Twelve months of promotion and price-change dates.

Media export summary: spend by channel and period, with platform-reported outcomes.

Creative inventory: what exists, what has run, what was retired.

Measurement inventory: which systems exist, who owns them, what each reports.

Claim-substantiation status and the name of the review owner.

The 90-day decision statement.

Your scorecard, with weights set before you meet anyone.

Setting weights before the first meeting is the single cheapest protection against being persuaded by presentation quality.

Limitations and tradeoffs

Be honest with yourself about what any partner — and any measurement approach — can and cannot deliver.

Retailer data latency. Sales and inventory data frequently arrive days or weeks later, so fast media decisions often run ahead of retail truth.

Inconsistent attribution windows. Different platforms and retailers use different windows and scopes; totals will not reconcile without documented definitions, which is why measurement standards call for disclosure.

Halo and product-scope choices. Whether a measured outcome includes only the advertised SKU or the wider brand changes the reported result substantially.

Platform modeling. Modeled conversions and store-sales estimates are subject to eligibility, volume, and methodology conditions set by the platform, not by you.

DTC versus retail reconciliation. Two channels with different pricing, margins, and data structures rarely combine into one clean efficiency number.

Promotion and distribution confounding. Price events and new store counts move sales independently of media and will contaminate naive reads.

Privacy thresholds. Aggregation minimums and consent requirements limit granularity, sometimes below the level a proposed analysis assumes.

Attributed is not incremental. Attributed sales describe an observed association within one system. Incremental sales require a counterfactual. Any agency that conflates the two will eventually recommend the wrong budget.

The practical next step

If you are actively evaluating partners, the most useful first conversation is not a pitch — it is a scoping discussion about the decision you are trying to make.

A scoped diagnostic conversation with our team examines your commercial constraint, your current source-of-truth map across retail and DTC, the partner scope that constraint actually requires, the evidence gaps that would undermine any agency's reporting, and the smallest responsible test that could resolve the biggest open question. You leave with a scoped view of the problem and a written list of the evidence any partner should be required to produce.

We make no claims about retail lift, distribution, CAC, or ROAS outcomes — no agency responsibly can. What the conversation produces is clarity about the job, the evidence, and the decision. If that is useful, start with the growth services overview or get in touch through the contact page.