The question: What happens when a system only learns from the outcomes of its own decisions?
Dataset & sample
- Source
- Wang et al., “LLM-Based User Personas for Recommendations at Scale”, arXiv 2606.12198 (Google, Google DeepMind)
- Period
- 30+ days
- Sample
- Live A/B test on a commercial video platform serving billions of users, equal non-overlapping traffic per arm
- Primary metric
- Watch time, active users
In June, a team from Google and Google DeepMind published how they added an LLM to a video recommender that serves billions of users. The headline result is a small lift in watch time. The most important number is somewhere else, in a short paragraph in section 5, and it explains why your marketing budget is probably stuck in a loop.
What they built
For every user, an LLM reads their watch history and writes a short description of their interests in plain text. It has two parts:
- Summarized interests: what you already watch. The familiar.
- Exploration interests: new topics, related to the first ones, that the model reasons you might like. What you haven't been shown yet.
Both feed the existing recommendation system, which turns each interest into actual videos. They tested it live for more than 30 days, with equal, non-overlapping groups of users on each side: one with the LLM interests, one with the production system alone. The paper calls it "a large-scale commercial video recommendation platform" and never names it.
The number that matters
Videos that came from the exploration interests got 40.91% fewer impressions than videos from the familiar interests.
But once they were actually shown, they were 13.6% more likely to be watched.
| Group | Value |
|---|---|
| Familiar | 100.0 |
| New | 59.1 |
Source: Wang et al., arXiv 2606.12198
| Group | Value |
|---|---|
| Familiar | 100.0 |
| New | 113.6 |
Source: Wang et al., arXiv 2606.12198
Read that again. The new content worked better, and the system showed it less.
The LLM wasn't the bottleneck. It suggested the right things. What held them back was the next layer: the ranking models that decide what you see first. Those models are trained on past engagement, and new content has none. So it scores low, gets shown less, collects less evidence, and scores low again.
The authors are blunt about the system they started from. They describe the production recommender as having a severe recency bias: it leans toward whatever you engaged with most recently.
Subscribers only
Keep reading.
It’s free.
Get the rest of this note, plus what I’d do on Monday.
Already subscribed? Enter the same email and we'll send you a sign-in link — or sign in here.