
Users often rate an AI experience more positively after several interactions. When the system includes memory, profile data, or tailored responses, an organization may quickly credit AI personalization for that improvement.
That conclusion can be premature. Perceptions may change during repeated use even without personalization, so the real measurement question is whether the personalized experience changed differently from a comparable non-personalized experience over the same period.
Adapted from an article by Vijay Rajan, Founder of Compeers AI.
The research, Tailored to you: longitudinal effects of personalising language models, is a Google DeepMind-authored arXiv preprint released in September 2026. It should be treated as preprint evidence, not as a peer-reviewed conclusion.
The preregistered randomized trial reported 992 successful completers and 4,960 conversations over five days. Participants took part in controlled, text-based relationship-advice conversations using assigned topics; the evidence does not indicate that every participant independently sought advice about a real personal relationship problem.

The researchers tested three conditions. The control had no personal information carried between sessions. Memory-based personalization used cumulative summaries of earlier conversations, while survey-based personalization used a persona summary generated from an intake survey.
Participants perceived the personalization manipulations. That matters because the lack of a different usefulness or competence trajectory cannot be dismissed simply by saying that users failed to notice the interventions.
Perceived usefulness and competence increased significantly across sessions in all three conditions. Neither memory-based nor survey-based personalization produced a significantly different longitudinal trajectory for these measures relative to the control.
The narrow implication is important: rising ratings are not sufficient evidence that AI personalization produced the change. Users may have become more familiar or comfortable with the interaction, but the study did not isolate familiarity, comfort, habituation, or any other psychological mechanism as the cause.
The finding also does not show that personalized AI is generally ineffective. The experiment evaluated two implementations in one controlled setting over five days. Different designs, tasks, data sources, or time periods could produce different results.
Nor did the study measure commercial outcomes such as conversion, retention, purchase behavior, task resolution, or brand preference. Statistical findings about perceptions in relationship-advice conversations should not be converted into business claims that the research did not test.
AI personalization can affect different outcomes in different ways. In this study, the memory-based condition received significantly lower creepiness ratings than the non-personalized control, but the reported effect was small at d = -0.10.
Conversation-log coding also identified more, and more rapidly increasing, discrete self-disclosure instances in the memory condition than in the control. This was another small effect, reported as d = 0.17.
That result needs careful wording. Self-reported disclosure depth or intimacy and the Self-Disclosure Index did not differ by condition, so the evidence does not support the broad claim that AI memory universally makes people share more personal information.
Lower creepiness should not be described as greater trust because the study did not directly measure a trust construct. Survey-based personalization also did not produce the same significant creepiness result as memory-based personalization, reinforcing the need to assess each personalization design separately.
The experiment involved controlled relationship-advice conversations, not customer service, commerce, sales, or branded marketing interactions. Customer interactions can involve different goals, histories, stakes, data expectations, and definitions of success.
The preprint therefore does not establish that AI personalization improves or harms customer satisfaction, resolution, conversion, retention, purchase behavior, or loyalty. Its practical value for marketers is as a research-design prompt: some apparent gains attributed to personalization might instead be changes that occur across repeated use.
Organizations should use that possibility to improve testing, not as a reason to adopt or abandon personalization universally.
The strongest practical test compares personalized and non-personalized experiences over identical periods. The goal is to determine whether personalization produces an incremental change beyond any pattern shared by both groups.
Avoid testing whether “personalization works” in the abstract. Ask a narrower question, such as whether conversation memory improves perceived relevance after several interactions or whether a profile-based experience reduces task effort.
Where feasible and appropriate, randomly assign comparable users to personalized and non-personalized experiences. Keep the core task, timing, model configuration, interface, and measurement intervals as consistent as possible.
Collect baseline data and repeat measures at equivalent intervals. A before-and-after improvement in only the personalized group cannot separate a personalization effect from time-related change, repeated exposure, or other shared influences.
Do not ask only whether the personalized group improved. Ask whether it changed differently from the control over time. That difference in trajectory is the more relevant evidence for attributing an effect to personalization.
Conversation memory, profile-based tailoring, survey-derived personas, contextual recommendations, and other approaches are distinct interventions. Pooling them under one label can hide meaningful differences in user reactions.
Interviews, open-ended feedback, or moderated research with real human respondents can explain why an experience felt relevant, useful, repetitive, intrusive, or creepy. Qualitative findings can also reveal whether participants understood what information the AI retained and how it was used.
Research involving remembered conversations or profile data should address consent, appropriate data use, user control, and potential regret about sharing. These considerations are part of the experience being evaluated, not administrative details outside it.
This framework is an editorial experimental principle informed by the preprint. It is not a customer-marketing result directly demonstrated by the relationship-advice study.
No single score can show whether AI personalization is successful. Teams should define a primary outcome before data collection where possible, then use secondary measures to explain or qualify the result.
Segment analysis can help identify whether reactions differ across relevant groups, but small subgroup results are easy to overinterpret. Segments should have adequate sample sizes and, where possible, a rationale specified before the analysis begins.
Compeers AI is an AI-native all-in-one platform for custom market research using real human respondents. Teams can use qualitative, quantitative, or mixed-method research to compare reactions to personalized and non-personalized concepts or experiences.
The platform supports planning, fieldwork, advanced analytics, analysis, visualization, reporting, and interactive research exploration across an end-to-end research workflow. AI accelerates execution, while human researchers remain involved in methodological decisions, review, interpretation, and recommendations.
Compeers can support the custom market research component of an evaluation. It is not a substitute for controlled product experimentation when an organization needs to establish effects on live behavioral or commercial outcomes.
Improvement over time is not, by itself, proof of personalization impact. Compare personalized and non-personalized experiences over the same period and examine whether their trajectories differ.
Personalization may still affect specific outcomes, including reactions to memory or observable disclosure behavior, even when usefulness ratings do not diverge. Investment decisions should rest on incremental evidence tied to the intended customer outcome, not on a favorable before-and-after trend alone.
Does AI personalization make an AI system more useful?
It can, but the cited preprint did not find a significantly different usefulness trajectory for its memory-based or survey-based conditions relative to the control. The result applies to those implementations and that five-day relationship-advice setting, not to every form of personalization.
Can repeated use make people rate AI more positively?
In the preprint, perceived usefulness and competence increased across repeated sessions in all three conditions. The study did not establish which psychological mechanism caused the increase.
Did the Google DeepMind study prove that familiarity caused higher ratings?
No. Familiarity is a plausible explanation to test, but the experiment did not isolate familiarity, comfort, or habituation as the cause of rising ratings.
Does AI memory make users share more personal information?
The memory condition produced more behaviorally coded self-disclosure instances than the control, with a small reported effect. Self-reported disclosure depth or intimacy did not differ by condition, so broader claims are not supported.
Does lower creepiness mean users trust an AI more?
No. Memory-based personalization received slightly lower creepiness ratings, but the research did not directly measure trust. The two concepts should not be treated as interchangeable.
Can this relationship-advice study be applied to customer experience?
Not directly. Customer interactions differ in their goals, stakes, data expectations, and success measures, so organizations should test the hypothesis in their own context.
What is the best way to compare personalized and non-personalized AI?
Assign comparable users to both experiences, hold key conditions consistent, and measure each group at the same intervals. Then test whether the personalized group changes differently over time rather than looking only for improvement within that group.
Which metrics should marketers use to evaluate AI personalization?
Use clearly defined perception, behavioral, business, and guardrail measures suited to the use case. Select the primary outcome in advance where possible, and do not treat one favorable secondary measure as proof of overall success.