Post

Individual AI Tutoring

Individual AI Tutoring

It was the fifth (or sixth?) time he impatiently checked his phone. He sat at Blue Lobster, a fancy restaurant he was visiting for the first time in his life.

“One evening here is three months’ worth of a RAG subscription,”

the young man mumbled disappointedly to himself as he glanced through the menu. Although a few minutes late, she did arrive, with a look of surprise on her face.

So, Leo, tell me, do you take girls here often? This place doesn’t feel exactly cheap.

What does she mean by “do you take girls here often”?! I consulted my RAG about the “not exactly cheap” part, of course, but does she mean I’m promiscuous? That this is just a little theater I’m acting out to get into her panties? This behavior is outside the RAG script! Now what?

Yeah, about that. Please excuse me, I have to… use the toilet. I’ll be right back.

On the way, Leo remembered that the dating guide app, Reasonably Apt Girlfriend (RAG), had recently started offering real-time dating advice provided via audio, complete with an amazing 50-millisecond response time. And who wouldn’t want that? Wouldn’t you? RAG turned out to be Leo’s lifesaver that evening.

ragged Dick Do we need apps telling us what to do all the time? Image source: Picryl

Leo relied on RAG to bypass the awkwardness of dating. Solution dispensers like Photomath do the exact same thing for mathematics: they strip away all the necessary friction, even more important in education than in romance.

The One AI To Rule Them All

Today, I experimented with solving a very simple rational equation using AI. My adventure wasn’t as successful as Leo’s.

I haven’t played with Photomath for quite some time, so in good faith, I thought I would take a picture of my entire handwritten solution to see whether the app could handle it now. It couldn’t, and it continues to fail at this.

That’s the main issue I have with apps like Photomath. They are essentially just solution dispensers without any overarching pedagogical design. Dispensers like these cannot help but give you the full solution with all the steps, stripped of any beneficial friction. That’s what machines do, after all.

whole solution, step by step Photomath is supposed to be good for verifying one’s solution. My solution got discarded immediately. Image source: MŽE

What you’re left with is a cage in which you’re shown how handy Photomath is as it solves all of your exercises for you.

dispenser mode, working as expected Dispenser mode in action, ready more than ever to shine. Image source: MŽE

I couldn’t help but try out the “AI-powered help” advertised on the main screen above. Once I was taken to a Google search, my handwritten solution from the camera was immediately flagged as wrong. The explanation was so misguided yet written so confidently that I actually needed to double-check whether I had made a mistake.

Google AI mode, part one It’s frightening how confidently the response is written. Devoid of logic and missing any citations. Is this garbage what students will be shown regularly from now on? Image source: MŽE (Emphasis added by author.)

Let’s take a look at what researchers actually have to say about GenAI in education.

Pseudoscientific Intermezzo

Some researchers, to put it mildly, should seriously review their methodology. For example, Mensah et al. (2025) aimed to explore:

the impact of Photomath on the mathematics achievement of pre-service teachers in algebra, with a focus on the mediating roles of conceptual understanding and interest

They did this by letting 257 college undergraduates play around with Photomath and then asking them how they felt afterward. Based purely on this self-reported questionnaire, they inferred causal claims such as:

Educators and institutions are encouraged to integrate Photomath into their teaching strategies to enhance students’ conceptual understanding of algebra.

This is questionable at best.1 Eerily, Kusi et al. (2025) published a study in a different journal with nearly identical settings and suspiciously similar conclusions.2

Let’s look around for some more honest research instead.

Language Models Are Horrible Masters

Liu et al. (2026) conducted a multi-faceted study on the behavioral patterns that immediately emerge when students use AI chatbots. Participants were first given an AI chatbot capable of answering questions correctly, but the chatbot was then removed for the final three exercises, which students had to solve on their own.3 Throughout the test, any question could be skipped, and wrong answers were not penalized.

Consistently, students with access to the AI chatbot were less likely to try answering questions on their own compared to the control group without the chatbot. Furthermore, the AI group was significantly more likely to skip the final three unassisted questions than the control group.

There is another subtle nuance here worth highlighting. The AI group appeared more engaged during the first part of the test simply because its participants were delegating problem-solving to the chatbot, rather than using it for consultation. The paper touches on an interesting explanation for this behavior: using an AI chatbot drastically decreases the perceived expected time it should take to solve a problem. In other words:

If a clanker can do it in under one second, why should I jump through all these hoops? Where’s the fun in that?

Bastani et al. (2025) similarly found that when students work with ChatGPT-based interfaces, their short-term performance increases by 50 to 127 percent, depending on the specific UI. Once access is removed, however, students’ performance drops significantly (a 17% reduction in grades) compared to the control group.

You might argue that if students were simply educated on the downsides of cognitive offloading, these negative side effects could be mitigated. But who is responsible for teaching students about the dangers of GenAI? Nobody really has a clear answer given how rapidly language model products evolve. Currently, designing any robust AI safety policy feels like guesswork to me.

I would also argue that teachers are dealing with three unprecedented problems they haven’t encountered prior to AI:

1. Teachers

Teachers themselves have to figure out how AI works and what techniques are effective. Because there is no authoritative playbook, integrating AI in schools remains messy. Meanwhile, out-of-school usage happens without pedagogical supervision. In their free time, students regularly interact with apps designed for engagement that inadvertently contribute to emotional dysregulation and reduced self-discipline. How can we expect them to spontaneously use addictive GenAI products in ways that align with what parents and teachers want?

2. Workload

Teachers are expected to guide kids on how (and how not) to use GenAI, recognizing that students use these tools outside of school regardless. But how can teachers offer that guidance when they are already overwhelmed by their existing workload?

3. Quality Assurance

Even if a teacher successfully uses GenAI to generate learning materials, the output is rarely usable out of the box. The same goes for homework submitted by students. Much like the Google AI example demonstrated at the start of this post, quality control is now also about detecting confident nonsense.

Even though the current landscape feels bleak, I remain hopeful and optimistic. Here is why.

Language Models Are Exceptional Servants

Kestin et al. (2025) suggest that designing an AI tutor aligned with human pedagogy yields better outcomes than standard in-class instruction without GenAI. That is because a single teacher cannot provide personalized feedback to an entire room of students simultaneously, and one-on-one tutoring has long been recognized as one of the most effective methods of knowledge transfer.

Historically, individual tutoring was a luxury—an interventional measure rather than a systemic, preventive one. Sophisticated language models change that dynamic. However, we still need to dictate how the technology behaves, not the other way around. Technology should adapt to humans, not humans to technology.

Specifically, the paper highlights several principles for designing effective AI tutors:

  • using progressive disclosure,
  • enabling self-pacing,
  • providing timely and targeted advice,
  • managing cognitive load,
  • encouraging active learning, and
  • promoting a growth mindset.4

Capinding (2023) points out another fascinating phenomenon: using educational technology can decrease the variance in student performance by helping struggling outliers catch up to the class average.5

ducklings Even though in-class education still has its strengths, it’s nowhere near as effective as one-on-one tutoring, especially for students with specific needs. Image source: Picryl

Third Space Learning (TSL), a US company developing its own AI tutor, makes a compelling case for the ethical, pedagogically aligned use of generative AI that complies with curriculum standards.6 What I find most refreshing about TSL’s philosophy is their inverted focus: teachers come first, technology comes second.

I tend to agree. We need to start distinguishing between using personalized GenAI solution dispensers and making GenAI reliable and aligning it with human pedagogy to deliver individual tutoring.

The fact that there are already people making this distinction is precisely why I remain optimistic about the future of education. TSL’s position shows genuine courage, particularly in an environment dominated by the “AI is omnipotent” narrative spreading across the internet faster than the covid pandemic.

I’ve learned this distinction firsthand from tutoring dozens of students over the years. Real tutoring focuses on reasoning and fostering a long-term passion for learning. Well-designed AI tutors must prioritize these same principles rather than optimizing for short-term incentives.

Happy learning!


  1. To be completely honest, the Journal of Pedagogical Sociology and Psychology doesn’t strike me as a top-tier peer-reviewed outlet. But that is precisely my point: when researchers struggle to evaluate technology properly, how can we expect the public to? ↩︎

  2. If your job is to shed light on high-stakes societal issues, you should be transparent about your intentions and methodology. If you’re writing research as a hobby, feel free to praise Photomath all you want. But when work is published under academic credentials, my tolerance for poor methodology decreases substantially because pseudo-research undermines the work of serious scholars. ↩︎

  3. The research evaluated student behavior during basic arithmetic and reading comprehension tasks, both showing similar patterns. ↩︎

  4. Although “growth mindset” can sound like a buzzword, I think of it as feedback depersonalization. The focus remains on the student’s strategy rather than framing errors around innate ability or disability. ↩︎

  5. While phrasing like this can sound like it was lifted from an HR textbook, the underlying point holds. Responsible technology integration can be empowering; human expertise should be the primary differentiator, not the tech itself. ↩︎

  6. Amorio and Paglinawan (2026) surveyed high school math teachers on how they use GenAI and found that despite numerous hurdles, teachers frequently find creative ways to overcome them. I highly recommend reading it if you have the time. ↩︎

Music Fun Fact

Loading a fun music fact...