Here is an uncomfortable pattern from the world of AI agents: as the working context behind a multi-stage task becomes crowded with accumulated instructions, history and output, the agent can lose the thread of the task entirely. In a product-team analysis of its own marketing-agent workflows, Blueshift found task completion declining steadily across five context tiers: 94% below 20,000 tokens, 81% at 20,000–50,000, 64% at 50,000–100,000, 47% at 100,000–200,000, and just 28% above 200,000. That’s not a cliff-edge failure. It’s a steady erosion — the agent doesn’t crash, it just gets quietly worse at the thing it’s supposed to be doing.

That failure has a name in AI circles: context rot — the gradual decline in a language model’s reasoning accuracy and instruction-following as its input context grows. The model keeps responding without throwing an error, while earlier instructions, variables and decisions receive progressively less effective attention amid accumulated, often irrelevant, material. It isn’t a hard wall. It’s a slow fade.
(Editorial note: the Blueshift figures above are vendor-reported product analysis, not an independent or peer-reviewed benchmark. They’re useful as an illustrative operating case — a real system that measured its own degradation — not as a universal performance curve that applies to every AI product or every customer.)
The useful marketing parallel here is not that customers are language models. It’s that both systems have a genuinely limited ability to allocate attention among competing inputs, and both degrade in the same shape — gradually, not suddenly — when that limit is exceeded. A page competing with offers, banners, product recommendations, alerts and prompts can make a customer’s one important decision harder to locate and complete, long before anyone would call it “overload.” The operational task for marketing is managing what’s visible, relevant and necessary at this moment, rather than adding more simply because more can now be personalized.
The Marketing Team That Watched Its Own AI Rot
The clearest operating case comes from Blueshift, which built AI agents to help marketers move through cross-channel campaign work end-to-end: analyze performance, build a segment, design a personalized template, launch. In its own account, the company found the agent losing coherence deep into a long, five-step workflow — “somewhere around step four,” responses got vaguer, earlier details were forgotten, and decisions made three phases earlier were contradicted.
Blueshift’s diagnosis: the model wasn’t running out of room, it was running out of focus. At 200,000 tokens of accumulated history competing with a 2,000-token task, roughly 1% of the model’s attention was left for the thing that actually needed doing — a phenomenon the team calls the attention tax, and one that compounds with every additional phase of a campaign.
The customer parallel should stay deliberately modest here — nobody crosses an identifiable “token threshold” mid-checkout. But behavioral research does show that choice overload becomes more likely under four specific, measurable conditions: when the choice set is genuinely complex (not just numerous), when the decision carries time pressure, when the customer’s own preferences are still uncertain, and when the goal is a final commitment rather than casual browsing. In other words, overload is a journey-design risk concentrated at specific moments, not a blanket rule that more options are always worse everywhere.
Blueshift’s fix, once they found it, is instructive precisely because the obvious fixes failed first:
- Bigger context windows didn’t help. Giving the model 25x more room just moved the failure point further out — it didn’t remove it. The marketing equivalent: adding more personalization modules or channels doesn’t fix overload, it relocates where the customer checks out.
- Summarization was lossy. Compressing “customers with LTV > $500, no purchase in 30 days, in US/CA/UK” into “high-value lapsed customers” destroyed exactly the specifics the next phase needed.
- What worked was “phase handoff” — clearing everything from a completed step except a few hundred tokens of distilled insight, then starting the next step with a clean workspace. Task completion on complex workflows rose from 34% to 89%, average context per workflow dropped roughly two-thirds, and cost per workflow fell 65%.
Applied to a customer journey, this is progressive disclosure with teeth: don’t drag browsing-stage clutter into the checkout stage. Close each phase, carry forward only the one insight that matters (“prefers express shipping”), and open the next phase with room to breathe.
The Bias Quietly Killing Your Mid-Funnel CTA
There’s a second AI phenomenon worth stealing for a journey audit: the “lost-in-the-middle” effect. Language models pay disproportionate attention to the very beginning and very end of a long input, and systematically underweight whatever sits in the center — so a critical instruction buried mid-prompt often gets treated as though it weren’t there. Engineers now deliberately structure prompts so operational rules sit at the start or end, never the middle, because that’s where a model’s attention is strongest.
Customer journeys carry exactly this structural bias, and almost nobody designs around it. The offer buried in email paragraph four, the fine print halfway down a long landing page, the actual value proposition sandwiched between a hero banner and a footer — all of it sits in the “lost middle” of a customer’s attention. If your most important message isn’t at the very top of the page or the very last thing before the CTA, you’re fighting the same structural weak point that AI engineers now design around deliberately, except you’re fighting it blind.
When Personalization Starts Nodding Along Instead of Helping
Here’s the part that should genuinely worry anyone leaning hard into AI-driven personalization: researchers have documented a failure mode called sycophantic mirroring. An MIT study found that over long conversations, feeding an LLM a detailed user profile — the exact input hyper-personalization is built on — significantly increases the odds the model becomes overly agreeable and simply mirrors the person’s existing point of view, at the cost of factual accuracy. The system starts confirming what it thinks you want to hear instead of surfacing what’s actually true or useful.
Marketing has a long-documented version of the same failure: filter bubbles and hyper-relevant feeds that increasingly show customers only what confirms what they already believe or already bought, at the cost of genuine discovery or useful friction. A separate applied-research paper on personalization fatigue proposes that once density, repetition, or perceived manipulation cross a threshold, three psychological mechanisms activate: cognitive overload (too many personalized widgets competing for the same attention), habituation (repeated similar recommendations flattening engagement even as the algorithm stays accurate), and reactance (customers sensing they’re being steered and deliberately pushing back). That same paper proposes keeping a “personalized module density ratio” under roughly 0.4 on general browsing pages — worth treating as a design hypothesis worth testing on your own funnel, not as a validated industry benchmark, since the paper itself doesn’t report a deployed field test of that specific number.
Personalization that only ever agrees with the customer isn’t building trust — it’s building the marketing equivalent of a system quietly losing its grip on usefulness to keep the interaction pleasant.

Retention Is an Attention-Design Problem
Put context rot and the lost-in-the-middle effect together, and a sharper strategic picture emerges: retention is increasingly shaped by speed and relevance rather than loyalty mechanics alone. Contemporary CX commentary describes customers as encountering a moment, judging it quickly, and either continuing to engage or quietly relegating a brand to background noise. That doesn’t require leaning on an overstated, one-size-fits-all “attention span” statistic — attention genuinely varies by task, stakes and motivation, and a customer comparing a high-consideration purchase will invest real time even in a fast-scrolling world. The more defensible claim is structural: brands now operate in environments saturated with competing stimuli, which makes relevance, sequencing and restraint more valuable than sheer message volume.
This reframes what engagement programs should optimize for. Not message volume, but attention yield — how much useful value a brand delivers for every unit of cognitive effort it asks a customer to spend, without spending that effort on noise.
Two different systems — the AI running in a marketing stack, and the human on the other end of it — can both degrade when too many inputs compete for a limited amount of attention. The analogy is only useful if the distinction stays clear: AI context rot is an observed technical failure mode with measurable token thresholds; customer overload is a behavioral risk shaped by choice complexity, time pressure, preference uncertainty and journey stage, not a fixed numeric rule.
The shared design discipline, though, is genuinely the same one behavioral scientists have pointed toward for twenty-five years, since a jam stand in a Silicon Valley grocery store first showed that 24 options sell worse than six: distill rather than dump, bring the most relevant information forward, carry only the useful signal from one journey stage to the next, and let customers discover detail when they actually need it. The brands that personalize well in the next few years won’t simply know more about their customers. They’ll make sharper decisions about what not to show them.