AI Tutor: The Difference Between an Assistant That Drives Progress and One That Does the Work
For the past two years, the "AI tutor" has become the inescapable sales pitch of every e-learning platform. Every LMS, every authoring tool, every ed-tech startup has its pedagogical chatbot. The promise is compelling: an assistant available 24/7, capable of answering learners' questions, explaining difficult concepts, guiding them through their learning journey.
The problem is that most of these tools do something else entirely. They reduce friction. They make the course more comfortable. And scientific research is increasingly clear on what that produces: learners who feel supported during training but perform worse when the tool is removed.
Before integrating an AI tutor into your training program, there are five questions to ask. Not the vendor's glossy marketing materials — about its actual behavior.
What AI Tutors Typically Do (and Why It's Insufficient)
The vast majority of AI assistants for training do three things: answer questions, summarize content, explain concepts. It's useful. It's not tutoring.
A real tutor — human or AI — doesn't just provide answers. It organizes the learner's effort. It knows when to help and when not to help. It calibrates its intervention according to acceptable difficulty level: enough for the learner to move forward, but not so much that they become entirely dependent on it.
The distinction between "answering a question" and "helping someone find an answer" is fundamental. It's what separates a search assistant from an educational tool. And it's precisely what most AI chatbots integrated into LMS platforms fail to do.
What Science Says About Learning: Key Findings That Change the Perspective
Research in cognitive science converges on one point: cognitive effort is a necessary condition for learning. Not painful effort, not suffering — but resistance, the obligation to search for yourself, to make connections, to get things wrong and correct them.
Manu Kapur, a researcher at ETH Zurich, theorized this concept as "productive failure": learners who struggle with a difficult problem before receiving an explanation retain the underlying concepts better than those who get the explanation first. The struggle is part of the process.
This is counterintuitive for training designers, and even more so for chatbot developers: the AI that answers quickly, that gives the right answer immediately, that avoids frustration — that AI can actively harm learning.
Two studies published in 2025 and 2026 document this effect with precision that should capture the attention of anyone integrating an AI tool into a training program.
The Data That Matters: Three Studies You Should Know
Bastani et al., PNAS 2025. Hamsa Bastani and colleagues conducted a randomized controlled trial with nearly a thousand high school students in mathematics at Budapest British International School.
During exercises with AI available, both equipped groups showed clear progress: +48% for the group with a standard assistant, +127% for the group with a guardrailed assistant. The guardrail was simple: providing teacher-designed hints rather than direct answers.
Then access was removed and students took the exam alone. The group with the standard assistant scored 17% lower than the group that never had AI. For the guardrailed group, this penalty largely disappeared.
The precise point deserves stating plainly: guardrails didn't improve final performance; they prevented it from degrading. It's not the AI that's the problem — it's AI without pedagogical design.
Kestin et al., Scientific Reports 2025. A Harvard team conducted a randomized controlled trial with 194 physics students, comparing a lesson with an AI tutor designed on explicit pedagogical principles to a lesson with active learning. Median gains for the AI group were more than double, with a median time of 49 minutes versus 60. The effect size was estimated at 0.73 to 1.3 standard deviations by quantile regression.
Two caveats the authors note themselves. Linear regression gives 0.63, and their preference for quantile regression stems from a ceiling effect at post-test that undermines precision of the upper bound. More importantly, the protocol compared an AI lesson at home to a lesson in class: the modality changed along with the tool. Reading this result as "AI beats the teacher" would be overinterpretation.
Liu et al., 2026. A team from Carnegie Mellon, Oxford, MIT, and UCLA conducted three randomized trials with 1,222 participants on mathematical reasoning and reading comprehension tasks. Result: just 10 to 15 minutes of interaction with an AI that delivers direct solutions is enough to degrade autonomous performance and persistence when facing difficulties. The authors call these systems "myopic collaborators," optimized for immediate and complete answers, incapable of saying no.
Worth noting because the nuance matters for application: participants were adults recruited for online experiments, not students in degree programs. This actually makes the finding more directly relevant for professional training.
All three studies point in the same direction: how you design the AI tutor determines whether the learner learns or relies on it. There is no neutral position.
Five Criteria for Evaluating an AI Tutor
1. Does it ask questions before answering?
A pedagogically serious tutor doesn't give the answer right away. It first asks what the learner has tried, what they already understand, where exactly they're stuck. This step isn't cosmetic: it activates retrieval practice, one of the best-documented learning mechanisms in research. An assistant that answers directly without this step is a comfort tool, not a tutor.
2. Does it offer progressive, configurable levels of help?
The appropriate level of help varies by learner, concept, and point in the learning path. A good AI tutor must be able to move from a gentle hint to a complete explanation, according to the policy defined for the course. And this configuration must be in the hands of the instructor, not imposed by default by the vendor.
3. Does it withdraw during assessment moments?
This is the question almost nobody asks when buying a tool. Yet it's decisive. An AI tutor that remains available during homework or tests isn't testing learner competency: it's testing their ability to use the AI. The Bastani data illustrates this clearly: negative effects on exams appear precisely in conditions where learners continued using the AI outside the tutored context.
4. Are its answers grounded in course content or general internet?
An AI tutor connected to a general LLM can answer any question — including questions unrelated to the course, with answers that are plausible but inaccurate in your context. A tutor grounded in course content (via a knowledge base or RAG mechanism) answers based on what the learner is supposed to learn. The difference isn't technical: it's pedagogical. It ensures consistency between the course and the assistance.
5. Can the instructor control the tool's behavior per activity?
A global setting for the entire platform isn't sufficient. Depending on the type of activity — discovery exercise, practice, summative assessment — the tutor's behavior should be different. This granular control must be accessible to the instructor without needing to go through a system administrator or contact technical support.
What We Built Based on These Criteria
The Tutor mode in PimenkoAI was designed around these five criteria.
It operates with four explicit sub-modes: understanding a concept, progressing step-by-step, revising, finding a resource in the course. The active sub-mode is chosen by the learner and determines the type of support they receive. Help levels (N0 to N3) are configurable per course.
The Tutor is automatically hidden in assessment spaces: Quizzes, Assignments, and Workshops. It remains available in learning spaces (Forums, Glossary, H5P). This isn't an option — it's the default behavior.
When a learner asks the assistant to complete an assignment for them, the exchange is redirected toward pedagogical help. The response isn't "I can't do that": it's "here's how to help you do it yourself."
A limitation is clearly displayed in the interface: semantic search across course resources (full RAG) is currently being finalized. PimenkoAI indicates this rather than presenting an answer as grounded in a source it hasn't consulted.
References
-
Bastani, H., et al. (2025). Generative AI without guardrails can
harm learning: Evidence from high school mathematics.
PNAS, 122(26), e2422633122.
Online: https://www.pnas.org/doi/10.1073/pnas.2422633122 -
Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti,
G. (2025). AI tutoring outperforms in-class active learning: an
RCT introducing a novel research-based design in an authentic
educational setting. Scientific Reports, 15, 17458.
Online: https://www.nature.com/articles/s41598-025-97652-6 -
Liu, G., Christian, B., Dumbalska, T., Bakker, M. A., &
Dubey, R. (2026). AI Assistance Reduces Persistence and Hurts
Independent Performance. arXiv:2604.04721.
Online: https://arxiv.org/abs/2604.04721 - Kapur, M. (2024). Productive Failure: Unlocking Deeper Learning Through the Science of Failing. Jossey-Bass.