Attribution modeling (part 1) — Introduction

This article is an introduction to a series of publications on attribution modeling.

Many things have been said about attribution modeling. It was supposed to be a panacea, the single source of truth, and the answer to questions about allocating advertising budgets. New tools and methods of attribution analysis were created, and sophisticated models were built.

Later came voices saying that attribution is gone and there’s not much left to deal with. Meta shut down its Facebook Attribution tool, and in Google Ads and Analytics the number of available models was reduced. But there is one thing that must be said clearly.

Attribution is like the weather. It can’t not exist. It’s always there.

Regardless of the quality of the data and the tools available, the attribution of conversions and revenue will always be one of the fundamental challenges for every marketer.

What is attribution modeling?

Attribution modeling is assigning to individual marketing activities their share in the revenue a company earns. In digital marketing analytics it often comes down to reporting the conversions and revenue generated by ad campaigns.

Despite the passing years, the most commonly used, standard models for evaluating the effectiveness of online marketing assign the entire credit for driving a conversion to the last click on an ad or another link leading to the website. The user’s earlier visits to the site are ignored, even though it often happens that users visit a site many times from different sources before they convert:

The conversion tracking systems of advertising platforms (e.g. Google Ads, Meta Ads) attribute conversions only to that system’s own ads. It doesn’t matter whether the user also came to the site from other sources or not.

If our customer clicks a Google Ads ad, then clicks an ad on Facebook, then visits the site via a link in an article on a web portal, and after a few days comes directly and makes a purchase, then:

  • Google Ads will indicate that the conversion comes from a Google Ads ad,
  • Meta’s ad conversion tracking will attribute it to the Facebook ad,
  • the Google Analytics traffic acquisition report will indicate a referral from the portal’s domain as the source of this conversion.

What’s more, each of these reports will “claim” that its source is 100% responsible for driving the conversion.

Users who convert on their first visit to the site make up only a certain part of conversions. That part will be smaller the more serious a decision the given conversion represents. If the conversion is, for example, a file download, it will on average require fewer visits than a newsletter subscription.

A conversion is often preceded by more than one visit from the user.

When the lead is a purchase inquiry, there will be even fewer first-visit conversions. If the conversion is a purchase with a completed payment, the number of conversions that required more than one visit will be greater — all the more so the more valuable the product being bought.

The average number of visits before a purchase will also be affected by other factors, such as the strength of the brand, trust in it and customer loyalty, as well as how unique the product is and how much competition there is on the market — which may prompt users to look for reviews and compare offers.

In e-commerce, users make on average several visits before a transaction, and those visits often come from different sources. We’re talking about multichannel conversions — several different, overlapping marketing channels are responsible for driving the transaction:

It’s obvious that if several conversion tracking systems each claim a 100% share in driving a given transaction, they can’t all be right.

Assigning all the credit for driving a conversion to one of the overlapping channels is, as a rule, incorrect, and decisions about allocating advertising budgets to individual channels made on this basis may be flawed.

Sometimes changing the attribution model changes little

The error resulting from assigning the whole share in a conversion to one interaction won’t always be large. In the case of overlapping channels, the effect of the sources bleeding into one another can cancel out, so it can happen that a model assigning 100% of the conversion share to the last click won’t differ significantly from alternative models.

For example, let’s assume we have the following conversion paths:

Attribution pathNumber of conversions
1.Google > Google10
2.Google > Facebook10
3.Facebook > Google10
4.Facebook > Facebook10
Total40

In such a case, the attribution of conversions to sources across different attribution models will look as follows:

SourceLast clickFirst clickLinearTime decay
Google20202020
Facebook20202020
Total40404040

As you can see, each of these completely different attribution models shows exactly the same thing!

In the real world the distribution of paths won’t of course be this even, but the example above shows how the influence of individual attribution paths on the overall attribution result can cancel out and flatten the differences between different models. See also the article on comparing attribution models.

Often, however, the differences between individual models can be very significant and result in even a manyfold underestimation or overestimation of a given traffic source’s profitability.

Attribution modeling helps determine the importance of individual traffic sources and the actual value of each of the overlapping sources leading to a conversion.

Why does attribution modeling matter?

Multichannel marketing is a team sport, and as in every professional team discipline, what matters is the selection of players and rewarding them appropriately. Appropriately — that is, in line with their contribution to the team’s success.

In football, one of the key metrics is the number of goals scored. So the bonus for a won match could be divided among the goal scorers. A player who scored more goals would receive a multiple bonus:

The thing is, though, that the outcome of a match is made up of the play of all the team members. In that view, players in defense or midfield would be considered unproductive, and the goalkeeper, who very rarely gets a chance to score — would essentially be an unnecessary cost.

You could modify the reward rules and allocate part of the reward to the players who assisted on goals:

But this method too seems to discriminate against players in defense. In truth, all the players should have a share in the reward. The question is: what share?

If the answer to this question is so hard, then why not reward all the players equally? Such a method is simple, clear, and — you could say — fair.

Except that not all players are equally valuable, and strikers are not stars and pillars of the team for no reason. So the optimal, best, most effective reward method — one that will let you assemble the team’s line-up and motivate the players — is more complex.

A similar situation can occur in companies, where it often happens that the highest performance bonuses go to salespeople. The result of their work is, after all, tangible — it’s the sales revenue from the customers they serve.

The thing is, if the product is of poor quality and has a bad reputation on the market, customers don’t seek it out, and logistics is failing — then even the best salespeople won’t be able to generate results. That’s why you have to take care of the whole organization and reward all employees appropriately.

Great products sell themselves, so paradoxically, lowering the salespeople’s percentage commission and supporting the team can increase the salespeople’s earnings, because they’ll start selling more, faster, and more easily.

Finally, rewarding appropriately means we can’t waste budget on those team members who don’t contribute to the result or whose results are illusory. Such a situation is described by the following story:

Imagine a store that employs salespeople to go out into the field and look for customers – and that also invests in advertising.

Now picture this: a customer sees an ad, talks to a salesperson and comes to the store to make a purchase – and there, right at the door, another salesperson is waiting. He grabs the customer by the hand and claims that it was he who brought them in.

That salesperson’s work will look highly “effective”, yet in reality it will bring the store no additional benefit. His results will come at the expense of the other salespeople, who – frustrated and underpaid – will start to quit.

The advertising budget will shrink as well, and as a result sales will fall.

There’s probably no longer any need to argue that properly assessing effectiveness and appropriately rewarding team members has a significant impact on the results achieved. Marketing will be no exception here.

What will attribution modeling help us with, and what won’t it?

Attribution modeling is not a measurement of a campaign’s impact on the increase in sales, but a method of analyzing the available data from conversion tracking.

Conversion tracking, which reports only the correlation between clicks and other interactions and conversions, does not provide data that would let us determine the degree of the cause-and-effect relationship between them.

A true measurement of a campaign’s effects is only possible within a test with a control group (conversion lift), which requires intervening in how ads are served (see also the article on the data-driven attribution model).

Even when advertising systems begin to model attribution with randomized control-group tests, they will only be able to do so within their own system. And if this barrier were ever overcome (read: all martech market participants, including Google, Meta, Microsoft, Amazon, etc. start freely exchanging data, which for many reasons is unlikely), certain sources would still remain independent (e.g. organic search results, referrals from external sites — which can’t be blocked for the purposes of the test).

That’s why running a test that consists of turning off each traffic source for a tested group of users — and thereby measuring the effectiveness of all marketing channels — will in practice most likely never be possible.

It’s also worth remembering that tracking systems such as Google Analytics don’t give us 100% of the information about traffic sources, and attribution modeling using these tools is limited only to the data available in them.

Here are a few reasons for the incompleteness of the data:

  • Many visits have no identified traffic source. In Google Analytics such visits will appear as direct visits. A direct visit is not only typing the site address into the browser or clicking a bookmark (see the article on direct visits). Many tracking-blocking programs remove tracking tags, which further increases the number of visits identified as direct.
  • Conversion tracking systems mainly measure online sources, so users who came to the site after seeing a TV ad, a billboard, a flyer, or a print ad, or after visiting a traditional store in a physical location — won’t have those touchpoints identified.
  • Tracking interactions other than clicks is difficult. While a click makes it possible to pass an identifier and process data in a first-party context, tracking impressions requires solutions that pass data between sites. Cross-site tracking is currently frowned upon, and tech companies are trying to counteract it. Some media don’t allow an impression tag to be placed at all, because they don’t want to share data about users’ use of their service on a mass scale.

Despite these limitations, attribution modeling can provide a lot of valuable information and help save and/or earn a good deal of money. Even if the answers obtained aren’t 100% certain, it will be easier for us to answer questions such as:

  • How much do individual traffic sources influence conversions and what is their value compared with the spend on them?
  • To what extent do traffic sources overlap?
  • Do the overlapping channels support each other or cannibalize each other?
  • How much should increasing spend, decreasing it, or dropping a given traffic source affect the change in revenue and spend?
  • Which traffic sources should we invest more in, which less, and which perhaps drop entirely?

Attribution modeling in practice

In the following articles on this topic we’ll try to look at a range of practical issues related to attribution modeling.

If you’re interested in consulting on attribution modeling, we invite you to contact us.

Next article (part 2): Channel groupings

Worth reading: A guide to attribution in Google Analytics

Author

Date

Let's talk about your business.

Porozmawiajmy o Twoim biznesie