The question: What happens when a system only learns from the outcomes of its own decisions?

Dataset & sample

Source
Wang et al., “LLM-Based User Personas for Recommendations at Scale”, arXiv 2606.12198 (Google, Google DeepMind)
Period
30+ days
Sample
Live A/B test on a commercial video platform serving billions of users, equal non-overlapping traffic per arm
Primary metric
Watch time, active users

In June, a team from Google and Google DeepMind published how they added an LLM to a video recommender that serves billions of users. The headline result is a small lift in watch time. The most important number is somewhere else, in a short paragraph in section 5, and it explains why your marketing budget is probably stuck in a loop.

What they built

For every user, an LLM reads their watch history and writes a short description of their interests in plain text. It has two parts:

  • Summarized interests: what you already watch. The familiar.
  • Exploration interests: new topics, related to the first ones, that the model reasons you might like. What you haven't been shown yet.

Both feed the existing recommendation system, which turns each interest into actual videos. They tested it live for more than 30 days, with equal, non-overlapping groups of users on each side: one with the LLM interests, one with the production system alone. The paper calls it "a large-scale commercial video recommendation platform" and never names it.

The number that matters

Videos that came from the exploration interests got 40.91% fewer impressions than videos from the familiar interests.

But once they were actually shown, they were 13.6% more likely to be watched.

Impressions (index, familiar interests = 100)
Impressions (index, familiar interests = 100)
GroupValue
Familiar100.0
New59.1

Source: Wang et al., arXiv 2606.12198

Chance of being watched once shown (index, familiar = 100)
Chance of being watched once shown (index, familiar = 100)
GroupValue
Familiar100.0
New113.6

Source: Wang et al., arXiv 2606.12198

Read that again. The new content worked better, and the system showed it less.

The LLM wasn't the bottleneck. It suggested the right things. What held them back was the next layer: the ranking models that decide what you see first. Those models are trained on past engagement, and new content has none. So it scores low, gets shown less, collects less evidence, and scores low again.

The authors are blunt about the system they started from. They describe the production recommender as having a severe recency bias: it leans toward whatever you engaged with most recently.

Subscribers only

Keep reading.
It’s free.

Get the rest of this note, plus what I’d do on Monday.

Free. No spam. Field notes and experiments.

Already subscribed? Enter the same email and we'll send you a sign-in link — or sign in here.