Loading Motifs…
Pick each reminder from a pool of written variants by how well each has worked, minus a penalty for ones this person saw recently. Duolingo's version beat random rotation; the penalty's own share is small.
+0.5% daily activesDuolingo's bandit against random rotation over a two-week online test; new users' next-day retention rose 2.2%. Group sizes not reported (Yancey and Settles, 2020).
Product screens: Duolingo via Deconstructor of Fun ↗. Screens illustrate the product; the number is sourced in the evidence below.
Built from published experiments and company reports across several products. Nobody has checked the screens on a dated day, so read each number with the caveat printed beside it.
Pick each reminder from a pool of written variants by how well each has worked, minus a penalty for ones this person saw recently. Duolingo's version beat random rotation; the penalty's own share is small.
A reminder people haven't seen recently tends to work better, and one repeated every day becomes background; Duolingo found a bandit that settled on one template per user would get repetitive and desensitise people (Yancey and Settles, Duolingo).
This play also appears inside the Duolingo playbook. You can still use it independently.
The bandit beat random rotation on daily actives and retention.
STRONGWhat it does not show A bundle; the recency penalty's share of the gain is not separated. Run and reported by Duolingo; group sizes, baselines and confidence intervals are not reported. Opt-outs, disabled notifications and uninstalls are not reported.
A Sleeping, Recovering Bandit Algorithm for Optimizing Recurring Notifications Yancey and Settles · 2020
Five months after launch, the bandit still beat a random holdout.
STRONGWhat it does not show The measurement window and number of rounds are not reported. Reward is two-hour conversion, not retention.
A Sleeping, Recovering Bandit Algorithm for Optimizing Recurring Notifications Yancey and Settles · 2020
Offline, the recency penalty added only a little.
SUPPORTINGWhat it does not show Off-policy estimates from randomized logs, not a deployed comparison. Empirical Bayes shrinkage and softmax exploration were not used in these estimates.
Use the play for this decision alone, or combine it with an existing product strategy after checking for conflicts.
A two-hour conversion reward ignores opt-outs, disabled notifications and uninstalls, so a variant can score well while raising them.
Without forced variety, a bandit can settle on one variant when variants barely differ.
Scores drift as the policy improves, making variants appear to worsen over time.
Plus everything above, as plain text for your coding agent, with a short instruction on top: check the fit, say what the evidence doesn’t support, and adapt it rather than copy it. Paste it and you get an answer, no prompt to write.
Duolingo's peer-reviewed bandit paper (Yancey and Settles, 2020), with one outside trial pointing the other way · evidence reviewed 25 September 2026.
Explore company-inspired strategies that connect your product, experience, and growth.
Browse the playbooks ↗A Sleeping, Recovering Bandit Algorithm for Optimizing Recurring Notifications Yancey and Settles · 2020