Blog

What Actually Keeps a Retailer on Retail Media Ad Tech After the Pilot

Sarah Mackinnon
July 28, 2026
Share
Jump to Title

The pilot goes well. CTR is up. The numbers are good enough that nobody on the call is arguing anymore. Everyone nods, the deck gets forwarded to finance, and the vendor gets treated like the decision is basically made.

It isn't.

A good pilot tells you a vendor's technology works once, on the version of your site that existed the week you tested it. It doesn't tell you what happens six months later, when your merchandising team wants something the vendor's roadmap didn't plan for, or when a specific demand partner starts causing a relevance problem nobody flagged in the RFP. That's the part of the evaluation most retailers skip, and it's usually the part that decides whether the platform is still in place a year from now.

TL;DR

  • A strong A/B test result is table stakes. It proves the technology works once, not that the vendor can keep up with you.
  • What actually determines long-term fit is how fast a vendor can ship something you specifically need, not something on a generic roadmap.
  • Vendors whose ranking logic is bundled inside their own ad server are structurally slower to respond to a single retailer's request, because any change has to work for every client on that system.
  • An independent optimization layer can ship a fix for one retailer's specific problem without waiting on a broader release cycle.
  • This is also why the commercial model matters: a flat plus variable structure only makes sense for a vendor that expects to keep shipping incremental things for you, not one built around a single large number.

The Pilot Isn't the Decision

Most retail media RFPs ask for a performance benchmark, and most vendors can hit one under test conditions. That's not a knock on the technology. It's just a fact about pilots: they're short, controlled, and designed to answer one question. Did unified ranking outperform the legacy setup on the pages we tested.

That's a real answer, and a useful one. It's also not the same question you'll be asking a year in, which is closer to: when we needed something changed, did it happen.

What "Flexible" Actually Means When You're the One Waiting on a Roadmap

Every vendor will tell you they're flexible. The honest way to check is to ask what happens when your team asks for something the vendor didn't already plan to build.

Here's a concrete version of that question, answered with a real example instead of a claim. Macy's flagged a relevance issue: ads from a specific demand partner were showing up in places that conflicted with how Macy's wanted its own product visibility to work. That's not a generic feature request. It's a retailer-specific problem that only shows up once you're actually live, using your own catalog, your own merchandising priorities, and your own demand mix.

Michael Krans has talked about how his team actually catches problems like that before they become a pattern. Macy's watches for what he calls pogo-sticking, a shopper clicking into a product page and immediately bouncing back out because the item wasn't what they expected. That behavior is the tell. It means the search experience broke a promise to the shopper, and it shows up in the data long before anyone would think to write it into an RFP. Krans has said that since Macy's began working with Pentaleap, that bounce behavior has been shifting in the right direction, and his team keeps watching it closely, along with the rate shoppers move from product page to bag to checkout, because relevance isn't a problem you solve once and move on from.

That's the point worth sitting with. Macy's isn't treating relevance monitoring as a one-time audit before signing a contract. It's an ongoing discipline, which means the vendor's ability to respond to what that monitoring turns up has to be ongoing too. A vendor that ships once and then goes quiet doesn't match a retailer that's watching this continuously.

Andreas Reiffen has described this as the pattern across retailers generally, not something specific to Macy's. His account is that most retailers who bring in real-time bidding technology start the same way: an e-commerce team unhappy with relevance, complaining about what's surfacing on the page. The instinct isn't to rip out the legacy ad server and replace it. It's to use the new technology to fix the relevance problem first, on top of what's already there, and only later use that same connection to bring in additional demand sources. The Macy's example fits that pattern exactly. It wasn't a special case Pentaleap built once. It's how the fix usually starts.

The response wasn't a roadmap conversation. Pentaleap shipped a targeted control, a way to exclude that specific demand source from specific placements, built in direct response to what Macy's needed. Michael Krans, VP of Retail Media at Macy's, put it this way:

"The partnership with Pentaleap helps us create a more open and flexible retail media ecosystem, one that ensures relevant ads and protects the user experience while expanding access to new demand sources that drive growth."

That's the pattern worth evaluating for. Not "can the vendor build features," every vendor can build features eventually, but "can the vendor build the feature your retailer specifically needs, at the speed your strategy actually requires."

Why This Is Structural, Not Just a Service-Quality Difference

This isn't really about which vendor has nicer account managers. It comes down to where the ranking logic lives.

If a vendor's optimization logic is built inside its own ad server, every change has to work across every client running on that system. A fix for one retailer's relevance problem has to be evaluated against every other retailer's setup before it can ship, because it's the same shared codebase serving all of them. That's a reasonable way to run a product. It's also a structural reason changes move slowly, no matter how responsive the team wants to be.

An independent optimization layer that sits between the ad server and the organic listings doesn't have that constraint in the same way. Because it's decoupled from any single ad server or demand source, a change built for one retailer's specific configuration doesn't need to be re-validated against every other client's setup first. That's what made the Macy's fix possible on the timeline it happened on. It's not a service commitment. It's an architectural one.

Michael Krans described the version of this problem that comes from being locked into a single vendor's stack. Talking through what retail media looked like before Macy's decoupled demand from ad serving, he explained that the retailer's speed used to be capped at whatever the vendor could get to. As he put it, "your roadmap was their roadmap." Any new placement, format, or fix waited on somebody else's development calendar, not the retailer's own priorities.

That's the exact failure mode a bundled architecture produces, described independently by a customer, not a vendor's marketing team. Once demand is decoupled from ad serving, a retailer isn't waiting on one company's release cycle to add a new source or resolve a relevance problem. Multiple sources can compete for the same placement, on the retailer's timeline instead of the vendor's.

Krans has described the effect of that decoupling in a specific way: it turns the ad server into an actual marketplace, where demand sources compete against each other for a placement in real time, instead of a system that just pushes pre-loaded campaigns out the door on a schedule. That distinction matters for this article's whole argument. A marketplace can absorb a new demand source, a new format, or a fix to a relevance problem without anyone waiting on a release. A delivery mechanism can't, because it was only ever built to move what was already loaded into it.

Andreas Reiffen, CEO of Pentaleap, has described the same shift from the architecture side. In the old model, he's said, a retailer had to pick one player that handled everything, ad serving, front end, demand, all bundled together. What changed is that the stack became modular and composable: a retailer can run its own front-end orchestration, choose its own ad-serving layer, and stitch in whatever demand sources it wants, each piece swappable without touching the others. That's not a different opinion about the same architecture Krans is describing. It's the same architecture, seen from the side that built it.

Why This Changes What You Should Actually Ask a Vendor

If post-pilot responsiveness is the real differentiator, it should show up in the RFP and the reference calls, not just the performance benchmark. A few questions worth asking directly:

  • Can you show us an example of a feature you built in direct response to one client's specific problem, not a general roadmap item?
  • How long did that take, from request to shipped?
  • Does a fix like that require changes to your core ad-serving logic for every client, or can it be scoped to one retailer?
  • What's your process when a request doesn't fit the standard roadmap?

These questions surface the structural difference faster than a benchmark ever will, because a vendor with logic bundled inside its own ad server will answer the last two questions very differently than one running an independent layer.

What This Means for Pricing

This same distinction is also the honest reason Pentaleap's commercial model looks different from the market's dominant revenue-share norm. A flat plus variable structure only makes sense for a vendor that expects an ongoing relationship built around continuing to ship things a retailer specifically needs, not one structured around extracting a percentage of a single large number and calling the relationship done. If the value is in staying responsive after go-live, the pricing model should reflect an ongoing partnership, not a one-time toll.

Krans has described this same shift in different terms: moving from a managed-service mindset to an infrastructure mindset. A managed service is priced and delivered as a one-time engagement. Infrastructure is priced as something ongoing, because it's built to keep evolving alongside the retailer rather than being handed over and forgotten. That's the logic a flat plus variable model is actually built on.

Why This Isn't Just a Vendor Question

It's worth remembering who actually benefits when a vendor stays responsive after go-live. Krans has described what he considers the real goal for Macy's search experience: media that feels like a personal stylist, not a billboard, where the distinction between organic and sponsored dissolves because the sponsored product is relevant enough that the shopper is glad to see it. That's not a vendor promise. It's a standard a retailer holds itself to, continuously, on every page. A vendor that can only ship once, at launch, can't help a retailer hold that standard for very long. A vendor built to keep responding can.

Key Takeaways

  • A good pilot proves the technology works once. It doesn't prove the vendor can keep up with your strategy afterward.
  • The real test is whether a vendor can ship something built specifically for your retailer's problem, on your timeline, not a generic feature from a shared roadmap.
  • This capability is architectural. Vendors whose logic is bundled inside their own ad server have to validate every change across every client on that system. An independent optimization layer doesn't share that constraint.
  • Ask for a specific example of a client-driven fix and how long it took to ship, not just a performance number.
  • The commercial model should match the relationship. A flat plus variable structure fits an ongoing partnership better than a revenue-share model built around one big number.
  • When ad serving and demand are locked into one vendor's stack, the retailer's roadmap becomes whatever the vendor's roadmap already was. Decoupling the two removes that dependency.
  • Responsiveness isn't just a vendor convenience. It's what lets a retailer hold its own relevance standard continuously, instead of resetting it every time a new problem surfaces.

Frequently Asked Questions

What should retailers look for in a retail media vendor beyond the initial pilot results? Look for evidence of post-launch responsiveness, specifically, whether the vendor has shipped a change built for one retailer's specific problem rather than a shared roadmap item, and how quickly that happened. Pilot performance shows the technology works under test conditions. It doesn't show how the vendor behaves once you're live and something unexpected comes up.

How do you know if a retail media platform will keep up with your strategy after launch? Ask for a concrete example, not a claim. A vendor that can point to a specific fix it shipped for a specific client problem, and describe how long it took, is showing you evidence rather than a promise. If the answer is vague or generalized, that's worth noting.

Why does it matter whether a vendor's ranking logic is inside the ad server or in a separate layer? Because it determines how fast a change can ship. Logic bundled inside an ad server has to be validated against every client using that same system before anything changes. An independent optimization layer that sits between the ad server and the organic listings can scope a fix to one retailer's configuration without that constraint.

Does a flexible commercial model matter as much as platform flexibility? Yes, and they're connected. A revenue-share model treats the relationship as a one-time transaction on a large number. A flat plus variable structure assumes an ongoing relationship where the vendor keeps shipping incremental value, which is the same behavior retailers should be evaluating for in the first place.

What happens when a retailer's ad serving and demand sources are locked into one vendor's stack? Any new placement, format, or fix has to wait on that vendor's own development calendar rather than the retailer's priorities. Michael Krans of Macy's has described this directly: under a legacy, bundled setup, a retailer's roadmap effectively becomes whatever the vendor's roadmap already was. Decoupling demand from ad serving removes that dependency and lets a retailer add or fix sources on its own timeline.

Stay Ahead with Retail Radar

Subscribe for cutting-edge insight into the latest retail media developments and trends

By submitting I accept the Privacy Policy.
Thank you! You are now subscribed to the Pentaleap newsletter.
Oops! Something went wrong while submitting the form.
A mail box
Thank you! You are now subscribed to the Pentaleap newsletter.