We chose to remove our AI tutor during quizzes. Here is why it is not a question of cheating.

When we explain that PimenkoAI's AI tutor automatically withdraws from Quizzes, Assignments and Workshops, the reaction is almost always the same: "That makes sense, otherwise learners cheat." It is true, and the figures confirm it. A Wiley survey of more than 2,900 North American students and instructors found that 86% of instructors believe cheating is more likely online than in person, while 46% of students acknowledge that AI makes cheating easier.

Except that this is not why we made that decision. Or rather: it is the less important of the two reasons.

The real reason is cognitive. A quiz is not only a measurement instrument; it is a learning act in its own right. An AI that helps during a quiz does not merely distort the grade. It removes the mechanism through which a quiz builds memory. More troubling still, neither the learner nor the instructor can notice this at the time.


The common assumption: "hiding AI during a quiz is about cheating"

This view is not wrong; it is incomplete. And it leads to poor design decisions.

If the problem is reduced to cheating, then the logical solution is surveillance: plagiarism detection, proctoring, browser lockdown, behavioural analysis. An entire market has been built around that. The underlying reasoning is that the learner is a potential cheat who must be contained.

This framing creates two problems. It establishes a relationship of mistrust with learners, which is rarely productive in professional training. More importantly, it misses the essential point: even a perfectly honest learner who uses AI during a quiz with the best intentions in the world, in order to "understand better", deprives themselves of a major cognitive benefit.

The question is not "Is the learner cheating?" It is "Are we taking away the effort that builds their memory?"


What a quiz really is: the testing effect

The testing effect (or retrieval practice) is one of the most firmly established findings in cognitive psychology. Its premise is simple: the very act of retrieving information from memory strengthens that memory more than rereading the same information.

In other words, testing yourself is not only a way to check what you know. It is a way to learn, and a more effective one than rereading.

The foundational experiment was published by Henry Roediger and Jeffrey Karpicke in 2006 in Psychological Science. Students read a text and were then divided into two conditions: some reread it, while others took a free-recall test. All of them then took a final test after different intervals.

The results after one week:

Condition Recall after 1 week
Tested group 56 %
Rereading group 42 %

In a second experiment, the gap widened further: students who studied once and then took three tests recalled 61% of the content after one week, compared with 40% for those who reread the text four times in succession.

One detail deserves emphasis because it changes how a quiz should be designed: in Roediger and Karpicke's experiment, the tested students received no feedback at all. There was no correction and no correct answer shown. Simply searching their memory was enough to produce the benefit.

This finding has been confirmed at scale. Yang and colleagues' meta-analysis, published in 2021 in Psychological Bulletin, aggregates 573 effect sizes from 222 studies, involving almost 48,500 students in real classrooms. The overall effect of quizzing on attainment is g = 0.499, a notable effect when set against the average for educational interventions (around 0.40). And the detail that matters here: the measured effect is g = 0.63 with feedback, compared with g = 0.60 without feedback. The difference is marginal.

The benefit does not come from the information received. It comes from the retrieval effort.

The same meta-analysis establishes a useful hierarchy: testing (0.499) outperforms rereading (0.330), which itself outperforms elaborative strategies such as concept mapping (0.095). The effect holds across all educational levels and 18 subject categories, and not only for factual knowledge: it also applies to conceptual understanding and transfer.


The trap of delayed benefit

This is the point that makes the question genuinely counterintuitive, and explains why the problem goes unnoticed in most learning environments.

Let us return to the 2006 experiment. After one week, the tested group wins by a wide margin. But what happens if the final test is administered five minutes after learning?

Interval Tested group Rereading group
5 minutes 75 % 81 %
2 days 68 % 54 %
1 week 56 % 42 %

After five minutes, rereading wins. The curves cross between the fifth day and the second day of the experiment.

This crossover is one of the most important findings in the learning sciences, and one of the least incorporated into the design of learning environments. It means that immediate performance and durable learning are two different quantities that can move in opposite directions.

Applied to a quiz with AI assistance, the mechanism becomes clear. A learner who receives help during the quiz gets a good grade. They feel that they have worked well, and that feeling is sincere: at that moment, they really do command the content because they have just processed it. The instructor, meanwhile, sees acceptable results and concludes that the sequence is working.

Two weeks later, in a real situation, the knowledge is not there. No one connects this to the quiz.

This is the very definition of invisible harm. It produces no warning signal at the point where it occurs, and the indicator everyone watches, the grade, moves in the wrong direction: upward.


What this means for AI in an LMS

An AI tutor available during a quiz intervenes at precisely the point when the learner should be making the retrieval effort. It short-circuits the mechanism.

The phenomenon is not limited to cases where the learner explicitly asks for the answer. Even AI that "helps them think", rephrases the question or gives a contextual hint reduces the retrieval load. Yet it is precisely this load that produces the memory benefit. Well-intentioned help is still help, and effort avoided is not effort made.

This finding is consistent with recent research on AI in education. The study by Bastani and colleagues, published in 2025 in the Proceedings of the National Academy of Sciences, followed nearly a thousand secondary-school mathematics students. Students with access to an AI assistant without pedagogical guardrails scored 17% lower on the final exam, taken without AI, than the group that had never had access to it. Yet during practice exercises, when AI was available, they performed considerably better.

The same pattern again: immediate performance up, durable learning down.

Research published in early September 2026 by David Stromberg, Victor Lei and Yanhui Wu gives this phenomenon unprecedented scale. The authors used 30 months of administrative records for 26,811 middle- and high-school students in a district of central China, comparing students according to when they adopted a generative AI tool. Six months after adoption:

Indicator Change
Time spent per assignment 64 min to 45 min
Assignment grade +18 %
In-person exam grade −20 %

In upper-secondary entrance examinations, the loss reaches 24% of the reference mean, and 18% in university entrance examinations. The most troubling result for our subject lies elsewhere. Among students who do not use AI, a good assignment grade predicts a good exam grade, as one would expect. Among users, the relationship reverses: the faster assignment grades improve, the more exam results deteriorate. The authors draw a direct recommendation from this: a rapid rise in assignment grades becomes a warning signal rather than a cause for satisfaction.

Two qualifications make this work even more useful. First, the penalty develops gradually: it is small at the outset and reaches its full extent only after six months for monthly examinations, and two years for entrance examinations. The authors conclude that short studies probably underestimate the real cost. Second, one AI user in five escapes the penalty entirely: those who continue to spend as much time on their assignments as before achieve the same results as non-users. The decisive variable is not access to the tool, but the effort retained.

This work is a CEPR working paper and has not yet been peer reviewed. Its convergence with Bastani's experimental findings, obtained using an entirely different method, nevertheless gives it considerable weight.


The honest counterargument: what about formative quizzes?

There is a serious objection to everything above, and ignoring it would weaken the argument.

Not all quizzes are assessments. Many are low-stakes practice exercises, specifically designed for the learner to make mistakes, discover where they stand and adjust. In that case, contextual help can make sense.

The nuance that matters is timing. Three configurations are defensible:

Help before the quiz. Revising with an AI tutor beforehand, in preparation for the attempt, presents no problem. The retrieval effort remains intact during the quiz.

Feedback after submission. Once an answer has been submitted, a detailed explanation of the error is useful, and the retrieval benefit has already been gained. This is probably the best place for AI in an assessment setting.

Help during an explicitly formative quiz, with no grade at stake and unlimited attempts. Here, the instructor may legitimately decide that support takes precedence over pure retrieval. That must remain their choice, not default behaviour.

What is not defensible is AI being available by default in every assessment, without the instructor having decided on it.


Our design choice

PimenkoAI's tutor is hidden by default in Quiz, Assignment and Workshop activities. It remains available in learning spaces: Forum, Glossary, H5P content and course pages.

We have invented nothing. This behaviour follows scientific findings established for twenty years, which we simply took seriously when defining the product's default behaviour.

Two clarifications about this choice.

It is configurable. An instructor who uses the Quiz module for low-stakes practice can reactivate the assistant for that activity. We do not claim to know their learning design better than they do. What we do claim is that the safe setting should be the default, and reactivation should be a conscious decision.

It is set per activity, not globally. A course can readily combine assisted practice quizzes with unassisted assessments. Indeed, that is often the right configuration.

One final word on how this kind of decision is made. Default behaviour is never neutral: the great majority of users do not change the settings they are given. Choosing what is enabled at installation is deciding on behalf of most people. That decision may as well be based on evidence.


See how the tutor's behaviour is configured per activity →


References