Come 2026, AI agents are going to systematically gut our traditional tracking methods. They’re already stripping UTMs and referrer data, which makes attribution a total mess. So what’s the plan? How do we run a real incrementality test when AI agents strip UTMs and referrers and we can’t trust the data coming in?
Key Takeaways
- Run controlled experiments with holdout groups. It’s the only way to directly measure campaign lift when the attribution data is shot.
- Use geo-based tests. You segment your audience by location, run different ads in different places, and isolate the real impact.
- Start collecting your own first-party data and get server-side tracking set up yesterday. You need a measurement system that doesn’t depend on client-side tracking.
- For the long game, get serious about media mix modeling (MMM). It’s how you’ll understand cross-channel effectiveness when you can’t see the granular details anymore.
Deconstructing “Project Horizon”: A Q4 2025 Campaign Analysis
In Q4 2025, we ran “Project Horizon” for a B2B SaaS client that sells compliance software. They gave us an ambitious goal: boost qualified leads by 20% over the previous year, focusing on big enterprise clients in finance. We had 12 weeks (Oct 1 to Dec 23) and a budget of $850,000 to spend across LinkedIn, Google Search, and programmatic. Right from the start, we knew AI agents were going to chew up our last-touch attribution, so we had to plan for it.
Strategy and Creative Approach
Our strategy was all about thought leadership and problem-solution content. On LinkedIn Ads, we ran carousels and videos with interviews from compliance experts talking about the real costs of getting it wrong, which set up the client’s software as the answer. For Google Search, we went after high-intent keywords like “financial compliance software” and “regulatory risk management solutions.” Programmatic’s job was to get whitepaper downloads and webinar signups using lookalikes from their current customer list. All the creative was designed to look sleek and professional, and we cut the jargon to make it easy for busy execs to digest.
Targeting Methodologies and Expected Challenges
We were super specific with LinkedIn targeting, going after senior managers and directors in compliance, legal, and risk at financial firms with 500+ employees. For Google Ads, we stuck to exact and phrase match keywords, with bids tied to conversion value. Programmatic used firmographics plus behavioral signals, like people who had just been reading industry news. We knew attribution was going to be a problem. Our own tests in Q3 2025 showed that 30-40% of traffic, especially from programmatic, was already showing up blind, no UTMs, no referrer, thanks to privacy browsers and AI agents. That told us our normal CPL and ROAS numbers would be completely unreliable unless we found another way to measure what was going on.
The Incrementality Framework: Beyond Last-Touch
Because we knew the data was going to be garbage, the entire measurement plan for Project Horizon was built around incrementality testing. Last-touch attribution was a non-starter. We predicted it would be worthless, and it was. So, we set up a geo-based holdout experiment to find out what our ad spend was actually buying us.
Implementation of Geo-Lift Testing
Here’s how we did it. We split the US into a test group and a control group. The test group included 15 big metro areas like NYC, Chicago, and SF, and that’s where we ran the full campaign. The control group was made of 15 similar metros, think Boston, Atlanta, Houston, where we either cut spend way back or shut it off completely for the 12 weeks of Project Horizon. The budget split wasn’t even: the test group got about 70% of the total budget, the control group got just 5% to maintain a whisper of brand presence, and the last 25% went to a national brand campaign that ran everywhere. This setup made sure any lift we saw was from Project Horizon itself. To make sure we were comparing apples to apples, we used census data and industry reports to match the geos on things like target audience density. For example, a 2024 Statista Research Department report shows NY and CA are huge for financial services revenue, so you can’t just throw them into different groups without a properly matched counterpart.
Data Collection and Analysis
The metrics we watched were the usual suspects:
- Impressions: Total views of ads.
- Click-Through Rate (CTR): Percentage of impressions leading to clicks.
- Leads: Defined as a form submission for a whitepaper, webinar, or contact request.
- Qualified Leads (SQLs): Leads that met specific criteria for company size, industry, and role, verified by our sales development team.
- Opportunity Creation: SQLs that progressed to a sales-qualified opportunity in the CRM.
We watched these numbers in both the test and control geos. For leads and opps, we just looked at the raw volume coming into the CRM. We tried to log lead source when we could, but the real truth was in the aggregate difference between the two groups. And yeah, the individual attribution data was a disaster. About 42% of new leads in our test markets came in with no source data, just as we’d feared from the AI agents. Trying to do a standard campaign-level ROI calculation on that data would have been a waste of time without the incrementality setup.
Results: What Worked, What Didn’t, and the True Impact
If you only looked at the raw platform data, you’d get the wrong idea entirely.
Campaign Performance (Test Group Only):
- Total Impressions: 125,000,000
- Average CTR: 0.85%
- Total Leads Generated: 11,250
- Cost Per Lead (CPL): $75.56 (based on attributed leads)
But the geo-lift analysis told the real story. In those 12 weeks, the test markets produced 2,100 more SQLs than the control markets. Since the control group had almost no ad spend, we can confidently say Project Horizon generated those leads. With each SQL being worth about $15,000 in LTV, those 2,100 incremental leads translate to a potential revenue lift of $31.5 million.
Incremental Impact:
- Incremental SQLs: 2,100
- Incremental Opportunity Creation: 420 (20% conversion rate from SQL to Opportunity)
- Incremental Revenue Potential: $31,500,000
- Campaign Ad Spend (Test Group): $807,500
- Incremental ROAS: 39.01x (Incremental Revenue Potential / Campaign Ad Spend)
That 39.01x incremental ROAS is a world away from what the simple $75.56 CPL would have told us. The gap just goes to show how essential this kind of testing is when your attribution is broken. If we hadn’t done this, we would have told the client the campaign was a mild success at best, not the massive win it actually was.
What Worked Well
- LinkedIn Video Ads: The videos with expert interviews were the clear winner, pulling a 1.2% CTR and great engagement. They really hit home with the target audience and got the funnel moving.
- Long-Form Content Gating: Gating our whitepapers and guides behind a form worked. Even with spotty attribution, we got high-quality leads because people saw enough value in the content to hand over their info, privacy settings be damned.
- Geo-Targeting Accuracy: The work we did picking the test and control geos paid off. It gave us a clean signal for the analysis.
What Didn’t Work as Expected
- Programmatic Display Attribution: Just like we thought, programmatic was an attribution black hole. It got hit hardest by the stripped data. We know it contributed to the overall lift because the geo test told us so, but it accounted for less than 10% of *attributed* leads. The geo-lift was the only way we could prove its value.
- Generic Search Terms: Broad search terms brought in traffic but almost no SQLs. We should have been more ruthless about cutting them and focusing only on long-tail, high-intent keywords.
Optimization Steps Taken
We didn’t just sit back and watch. We made a few key changes mid-flight:
- Shifted Budget to LinkedIn Video: We saw the video ads were working, so we moved $50,000 out of weak programmatic segments and put it straight into LinkedIn video. We immediately saw top-of-funnel engagement tick up in the test geos.
- Cleaned Up Google Ads Keywords: We killed a bunch of broad match keywords that weren’t converting and doubled down on exact and phrase match terms, especially stuff like “best [software category] for [industry].” That alone gave us a 15% bump in conversion rate for the search leads we could track.
- Started a Server-Side Tracking Pilot: We got a pilot going for server-side tracking on our landing pages to get cleaner first-party data. It wasn’t fully running for this campaign, but in the small segment we tested, it cut down our “direct/none” traffic by 25%. Seriously, every marketing team needs to get on this for 2026. If you’re still just using client-side tracking, you’re fighting a battle you’re going to lose.
If Project Horizon taught us anything, it’s that the old days of just trusting client-side, last-touch attribution are finished. You have to use more advanced measurement like incrementality testing to figure out if your campaigns are actually working, because AI agents and privacy tech are making the old signals useless. This really changes how you have to think about analyzing campaigns.
If you want to measure performance when UTMs and referrers are gone, you need two things: a solid incrementality testing plan, like a geo-lift study, and a serious focus on collecting your own first-party data. That’s how you’ll find out what’s really driving your business.
What are UTM parameters and referrers, and why are they important for marketing attribution?
UTM parameters are the little bits of code you add to a URL to track where your traffic is coming from, the source, medium, campaign name, etc. The referrer is just the URL of the webpage a user was on right before they clicked to your site. You need both to see which of your marketing campaigns are actually sending you people and sales.
How do AI agents strip UTMs and referrers?
Think of AI agents as bits of software inside things like ad blockers, private browsers, or search engine bots. They’re built to act like a person browsing the web, but their goal is to protect user privacy or index content. To do that, they often just delete tracking tags like UTMs and referrer data from the traffic, which is why so much of it shows up as “direct/none” in your analytics.
What is incrementality testing, and why is it becoming essential?
Incrementality testing is how you measure the actual, real-world impact of your marketing. You do it by comparing a group of people who saw your ads (the test group) to a similar group who didn’t (the control group). You have to do this now because the old attribution models fall apart when tracking data goes missing, like when UTMs get stripped. It’s the only way to prove your ad spend is actually bringing in new business you wouldn’t have gotten otherwise.
What are some common methods for conducting incrementality tests?
The most common ways are geo-lift testing, where you compare performance in different geographic areas, and A/B testing with holdout groups, where you just don’t show ads to a random slice of your audience. There’s also ghost bidding, where you enter an ad auction but don’t actually serve an ad, which helps you measure the baseline. The goal of all these methods is the same: to isolate the extra lift your marketing created.
Beyond incrementality testing, what other strategies can marketers use to combat lost attribution data?
Besides incrementality, you should be obsessed with first-party data collection, getting data directly from your customers. You can also set up server-side tracking, which sends data from your server to your analytics tools, getting around the browser restrictions. And finally, get good at media mix modeling (MMM). It gives you that high-level view of how all your channels work together to affect the business which is exactly what you need when the little details are fuzzy.