Nobody is coming to check your engagement design, and that is the problem

Focused business professionals collaborating in a modern office setting with a digital workspace.

Ethical gamification has an odd problem, which is that nobody is checking it. I should say at the start that this is not quite true, and I would rather correct the title myself than have you do it. Somebody is checking. In January 2023 consumer authorities across the EU screened 399 online shops and found manipulative practices on 148 of them. A European Commission behavioural study published in May 2022 found at least one dark pattern on 97% of the most popular websites and apps used by EU consumers. Regulators are looking.

What they are looking at is the countable stuff: fake countdown timers, hidden costs, pre-ticked boxes, cancellation flows built to wear you down. Nobody is checking the part a product team spends its days on, which is whether the streak, the reward, the notification and the leaderboard serve the person using the product or only the metric. That is the part I call ethical gamification, and I mean something specific by it: engagement designed so that the person is better off for having been engaged, and would say so if you asked them.

There is no inspector for that, no standard to hold it against, and in most product teams nobody whose job it is. So the first check your engagement design ever gets is the one you run yourself. Most teams never run it.

Why a product team is the right place to start

Here is a constructed example, not a client. A product team of six ships a seven day streak to lift week one retention. In the first sprint the retention number goes up, the team is pleased, and the streak is declared a success.

Nothing in that story is wrong. The metric moved. What the metric cannot tell the team is why it moved. Did people come back because the product was useful to them on day four, or because they did not want to lose a number? Both look identical on the dashboard in week one. They look very different in month three, when the users who stayed for the number are the ones who feel faintly resentful every morning and say so in reviews.

The team did not decide to be manipulative. Nobody in the room was asked whether the mechanic served the user, because no step in the process asks it. Reviews cover accessibility, performance and brand. Engagement design sits in a gap between them.

What the research says ethical gamification is aimed at

The best evidence I know of on where gamification actually shifts people comes from a 2024 meta-analysis by Li, Hew and Du in Educational Technology Research and Development. It covered 35 interventions and about 2,500 participants, and it looked at learners, not product users, so I treat it as a lead and not a rule.

It found that gamification raised peopleโ€™s sense of belonging and sense of choice by far more than it raised their sense of competence. Points, badges and leaderboards mostly speak to competence. In other words, the mechanics most teams reach for first are aimed at the weakest lever, and the two that did the most work are the ones that are hardest to turn into a dark pattern, because they only function when the person feels they have a real choice and a real place.

I would argue that is the check a product team can run without a lawyer. Ask what lever each mechanic pulls. If the answer is fear of losing something, that deserves a second look.

Four questions for a Friday afternoon

I score engagement design on five dimensions: belonging, autonomy, honesty, inclusion and AI transparency. You do not need the full instrument to start. These four questions cover most of the ground for a single feature.

Would the user still choose this if they could see exactly how it works? If the honest answer is no, you have found a dark pattern, whatever the retention chart says.

What happens when they stop? A design that punishes absence, with a lost streak, a shamed leaderboard position or a nagging reminder, is telling you its real purpose.

Who cannot use it? A mechanic that only rewards speed, competition or constant availability excludes people by design, and it usually excludes the ones you most needed to reach.

Who can see the AI in it? If a model decides which nudge a person gets, the person should be able to tell. That is both good practice and, increasingly, a legal direction.

What to do with the answers

Write them down, with a name against each, and keep it to one page. The aim is not a report. The aim is a record that somebody asked, which is more than most teams can show.

If a feature fails the first question, do not switch it off in a panic. Look at what it was doing for the user before it started doing something for the dashboard. Often there is a good mechanic underneath, such as a progress view that reminds people what they already finished, wrapped in a bad one, such as a penalty for missing a day. Keep the first and drop the second, then watch what happens to the retention number. If it falls, you have learned how much of it was borrowed from fear, and that is worth knowing before a regulator or a journalist works it out for you.

If a feature passes, say so and move on. A check that only ever produces bad news will not survive a busy quarter. In my view the habit matters more than the verdict, because the teams that ask the question every release are the ones that get faster at answering it.

What AI changes

AI makes this problem bigger rather than smaller, and I want to be precise about why. A model is very good at spotting patterns. It is not good at the emotional and motivational drivers behind them, and it will confidently tell you it worked great when it did not.

Ask an AI tool to propose a mechanic for a learning or retention problem and it will suggest the leaderboard, because the leaderboard is the most common mechanic across the material it has seen. Frequency is not effectiveness, and the model cannot tell the difference. So a team that uses AI to personalise nudges without a check is multiplying whichever mechanic was already the most common. The weakest lever, at greater speed.

The regulation, said carefully

I am a designer, not a lawyer. Nothing here is legal advice. This is how I read the direction of travel and how I design around it.

The Digital Fairness Act is a proposal the Commission has said will address dark patterns, addictive design and unfair personalisation. As far as I can verify at the time of writing, it has not been published. The trackers I checked, last updated in December 2025, expected a proposal in the third or fourth quarter of 2026. If I were designing for it, I would not wait for the text. I would run the four questions above on the features that depend most on habit.

The check nobody else will run

Come back to the title. Regulators check the visible patterns on the shop floor. Nobody checks the design decisions that sit behind the product, and when one eventually does, it will be a journalist, a regulator or a user with a screenshot, rather than a colleague who asked a question on a Friday.

Run it yourself first. It costs an afternoon, and the answers are better coming from inside the team.


Sources checked on 5 October 2026

  • European Commission press release IP/23/418, 30 January 2023: 148 of 399 online shops screened with manipulative practices.
  • European Commission, Directorate-General for Justice and Consumers, Behavioural study on unfair commercial practices in the digital environment: dark patterns and manipulative personalisation, 16 May 2022 (Publications Office record): 97% of the most popular websites and apps used by EU consumers deployed at least one dark pattern.
  • Li, L., Hew, K. F. and Du, J. (2024). Educational Technology Research and Development, 72(2), 765 to 796. Checked against the Springer record.
  • Digital Fairness Act status: StreamLex tracker dated 19 December 2025 and Wikipedia, neither newer than late 2025. Check the Commissionโ€™s own pages before posting.
  • The AI argument is mine and rests on no external source.

Similar Posts