The architecture of the data trade

Few systems are as invisible – and at the same time as far-reaching – as the trade in user data. What we click on, search for or visit becomes a signal in a global market that sorts people into categories and targets information at them.
By Clara Schöttke (text), Paul Geiersbach (photo)

Admittedly, clicking ‘Accept’ is often quick and easy. Who cares about the sheer volume of data? And surely a few little searches get lost in it all – or do they?

Work is stressful, there’s no room left on your to-do list and your head is full. In your break you want to daydream for a moment, just to check what the weather’s like in Rome. So you pick up your phone and tap on the free weather app. Before you can type ‘Rome’ into the search box, a message pops up saying you need to enable location services to use the app. The button glows such a bright green that you instinctively click ‘Accept’ without reading the message. It’s 28 degrees in Rome, and the sun is shining.

In the afternoon, a friend tells you he has been diagnosed with depression. After the call, you wonder how best to support him and google the illness to understand it better. The very first result shows general information and advice in the preview. To read the content on the website, you have to take out a subscription or accept the site with advertising. Taking out a subscription would take far too long, and you only want a quick overview. So you click ‘Accept’.


Nobody can expect all the information about their behaviour to end up with hundreds of companies.

Wolfie Christl, tracking researcher
A landscape-format colour image. It is night; through a window you see a person sitting at a desk in front of a computer. The living-room light is on. 2023, Kassel

Moving around the internet unobserved? That’s harder than you might think. But what data is collected, and by whom?

What few people realise is that this click puts them into collections – into 651,463 categories, to be precise. Data trails and personal data are the new currency of the digital age. Users look at free content and in doing so accept that their personal data will be collected, analysed and traded. This often happens without their knowledge, and on an unimaginable scale.

It is data brokers who strike gold in these mountains of data. The categories into which personal data are sorted are called ‘segments’. They contain hard facts such as age or places visited, but also a great deal of information based on inferences and estimates.

Your search for the weather in Rome? You may have been put into the categories ‘holiday prospect’ or ‘interested in culture’. And thanks to your location sharing, it gets even more specific: you are a person in your particular city who is interested in Rome. 

And the Google search about depression you needed to help a friend? Suddenly you are categorised with a sensitive label such as ‘depression’, ‘depression-related’ or ‘mental health’.

A landscape-format graphic. White lines connect various logos; these strands and links form one large cloud. Screenshot 2023, Lightbeam

It’s not just the website we visit that collects our data. Lightbeam visualises the third-party requests a website makes – which other sites are quietly connected to it or contacted in the background.

A portrait-format colour photograph. Green grass photographed from above, the blades lying in different directions, some pressed flat against the ground. 2023, Kassel

Unlike online, the traces we leave in the physical world are visible. When we move, we – and others – can see the path we have taken.

The industry trades in segments just like any other goods. But these are not direct data. Advertisers don’t pay for the raw data of individuals with a pseudonymous ID. Instead, they pay to reach people in a targeted way through particular audience segments. The list of segments is long, ranging from simple descriptions such as ‘men over 40’ to creative labels such as ‘fragile seniors’ and sensitive ones such as ‘depression’, ‘LGBT’, ‘gambling addiction’ or ‘political decision-makers’. Highly sensitive data such as health information and sexual orientation are listed, too. Who is for or against abortion or Black Lives Matter – it’s all collected in audience segments. 

More categories

Age: 25-34
Age: 35-44
Age: 45-54
Age: 55-64
Age: 65+
Zodiac sign > Gemini
Sports fans
Jewelry lovers
Organic food products > transactions in the last year
advertising enthusiast with restricted cross
Home Ownership > more likely
Household Size > 2 persons
Neighbourhood-area > very good
Energy > Green user
Education > Academic
Social Status > low earners without orientation
Lifestage > singles, low to average income, young age
Insurance > individualistic about risk-taking
Social Status > minimalist high-income earners
Family Type > multi-generation household
Lifestage > multi-person Household, low to average incom
Healthy Living > Beauty & Wellness Enthusiasts
COVID-19 Business > Business Decision Makers
Older people in suburban communities
Well-off residents of commuter-belt communities
Affinity for private health insurance
Probability of payment default – highest
gambling and lottery
with children 0-3 years
FAZ
busy moms
Premium Consumers – aged 30-39
Municipal decision-makers
Bargain Hunter
Decision Maker
Furniture & Shopping Cart 500+ EUR & Desktop User
Oversized Women
Interest in Schlager Music
Wine Drinkers
in Finding a Relationship
very low income
highest income
golden ager
culture lovers
household budget manager
tv heavy consumers
pregnant
married
dating
football supporters fortuna duesseldorf
erotic
hedonists milieu
traditional milieu
liberal intellectual milieu
zdf tv watchers
very high creditworthiness
low creditworthiness
parents vaccinate A

The global data trade is driven by a small number of highly interconnected companies. They include traditional data brokers such as Acxiom and Experian, which derive extensive profiles from millions of individual data trails and market them as audience segments. The digital advertising system is also dominated by platforms such as Google and Meta, which build detailed behavioural profiles from search queries, location data and app use, and use them for personalised advertising.

The ad-tech sector plays a special role: companies such as Criteo – a French provider known for its retargeting and cookie-based tracking – and other real-time bidding networks collect data across websites and sort users into finely grained target groups. There are players using such mechanisms in Germany, too, such as the Schober Group in the address and marketing business, or analytics firms such as Zeotap, which work with anonymised mobile and customer data.

What these companies have in common is that their work mostly remains invisible: users rarely know how many players are involved in their data trails, or how detailed the resulting profiles are. While the GDPR tries to regulate these practices in Europe, comparable models are largely legal in the US. The result is a global system that turns everyday digital actions into tradeable segments – and so influences what information, offers and opportunities people get to see online at all.

A white wall with a network socket, two cables plugged into it. The fitting has broken away from the wall. 2023, Hannover

When troubleshooting, the first point of contact with the internet is often the physical cables.

How exactly does it work – and what happens to our data?

We all constantly leave data behind online, for example through cookies, website trackers, location data, credit card details or metadata. Even the way we type on a keyboard or scroll through a website is recorded. Free services such as search engines, dating games and weather apps in particular generate large amounts of data, some of it sensitive. So-called ‘third-party requests’ are also loaded in the background when you open them.

This endless stream of data – which city you check the weather in, or which part of an article you skipped while reading – can be used by data brokers. Companies obtain data from various sources, organise it and repackage it in order to track people across different devices. Data are offered to other companies in exchange for money or other economic benefits. That’s why the author of this article still sometimes comes across the dress she bought for her school-leavers’ ball in 2017 as an online advert.

What might such a process look like?

1. Data collection

Methods of collecting or obtaining personal data, including hidden means such as trackers placed on websites visited, or from open sources such as electoral rolls, social media or purchase databases, and profiles from other sources such as data brokers. 

2. Profiling

Users are divided into small groups or ‘segments’ based on characteristics such as consumer behaviour, personality traits, demographics, use of apps & websites, location data, creditworthiness, illnesses or sexual orientation. 

3. Personalisation
This involves designing personalised content for each segment. 

4. Targeting
Personalised content is distributed via online platforms to reach the target audience with these tailored, targeted messages.

The problem is obvious – and yet more far-reaching than you might first think. For one thing, collecting and passing on data is an intrusion into everyone’s privacy. For another, this kind of data trade deliberately exploits vulnerable groups and the weaknesses of internet users. Segments can be used in marketing to deliberately include or exclude groups of people from information, and so to discriminate against them. They decide who sees adverts for flats or jobs, and where misleading or fraudulent advertising – so-called ‘scamvertising’ – appears. In concrete terms, this can mean, for example, that you aren’t shown rental flats outside the income bracket you’ve been assigned.

A landscape-format colour photograph. Details of a book and branches in dark tones. 2023, Netherlands

Finding out what data has been collected, and getting a glimpse into the world of cookies and third parties, is hard. Asking questions and getting answers is difficult, complicated and full of obstacles.

A landscape-format colour photograph showing a close-up of a street corner with a building and road. 2023, Kassel

A major cable junction lies hidden beneath the tarmac at an ordinary street corner in Kassel. Hardly anyone realises that the internet, too, is just an enormous number of cables joined together.

A landscape-format colour photograph. A desert landscape seen from above, with a gentle sand dune that has caved in at the centre of the image. 2023, Netherlands

Glitches, lags, errors – when something goes wrong, the first fix is usually to switch it off and on again. But what happens when there’s a data leak? And where does our data disappear to?

The internet isn’t free – and we’re going round in circles

The idea that all content on the internet is freely accessible and free of charge is wrong, and it’s absurd that this belief persists. Instead of paying with money, you pay with personal data. On the one hand, companies advertise that they can offer an even better-tailored personalised search; on the other, they deny collecting data and stress anonymity. The author of this article has experienced this too: two years ago she searched for tractors on Instagram just for fun. For up to a year afterwards, she kept being shown tractors and farm machinery.

A landscape-format colour photograph. A close-up of a bare stomach; where the belly button would be, a thin, translucent, paper-like layer is stuck over it. 2022, Saarbrücken

In science fiction, clones have no belly buttons, because they don’t need them. What do our digital clones look like – the identities pieced together from our data and digital traces?

2023, Austria

Nutrients and data flowing through water or through the internet are what aquatic plants and data brokers live on.

Data companies sort people into groups and decide who gets to see what. And because of the segments, we’re all going round in circles. What we search for is stored in order to show us similar suggestions later. Breaking out of the categories is difficult, if not almost impossible. That is not only a danger to a democratic society; it also has a decisive influence on how we see the world and ourselves. By selecting and categorising, companies directly shape users’ subjective perception – and with it, their perception of world events.

They don’t just decide which categories people are sorted into; they trade the data, analyse it and can therefore also decide which person is shown which advert. They know their users – better than we think – and shape how they see the world and themselves.