Blog · chatgpt

ChatGPT for studying: how to use it without sabotaging your memory

ChatGPT for studying: how to use it without sabotaging your memory

You paste in the lecture PDF, ask for a summary, ask for examples, ask it to explain again, simpler. The session is more productive than ever: in an hour, ChatGPT chews through a chapter that would have taken an afternoon. Days later comes the exam, or the simple attempt to explain the subject to someone, and almost nothing stuck. That feeling that studying with AI produces a lot and holds on to little is not your paranoia: it is exactly what the first controlled studies on the subject are measuring.

The direct answer: using ChatGPT to study works, but in the opposite direction to the most common use. When the AI produces the effort (summarizes, answers, writes for you), the immediate gain comes with a drop in retention. When it demands the effort from you (asks questions, corrects your attempt, points out the holes in what you explained), it becomes the cheapest private examiner in history. The line between the two uses is a single one: who sweats during the session. Outsourcing the effort is outsourcing the memory.

Below: the three studies that measured the effect (one of them Brazilian), why comfortable use fools you so well, the five-prompt protocol that reverses the direction, and what OpenAI's study mode solves.

Disclosure of interest: this blog belongs to Sift, the app we built to schedule reviews. This post evaluates ChatGPT as a study tool; all the numbers below have a primary source linked.

Does ChatGPT help or hurt? What three studies measured

For years, the question "does AI harm learning?" was answered with opinion. Between 2024 and 2025, three research groups put numbers on the table, with very different designs and limitations. Read together, they tell a coherent story: the problem was never AI on the study desk; it was AI that hands over the finished answer.

The Brazilian RCT: faster study, lower retention 45 days later

Barcaui (2025), published with peer review in Social Sciences & Humanities Open, randomized 120 business undergraduates in Brazil between studying AI concepts with ChatGPT or with traditional methods. On a surprise test 45 days later, the ChatGPT group scored 57.5%; the traditional group, 68.5% (Cohen's d = 0.68, a medium-to-large effect size).

The detail that matters most: the ChatGPT group studied on average 3.2 hours, against 5.8 for the traditional group. The study got faster and retention dropped, in the same motion. It is perceived productivity paying the bill with memory.

Honest caveats: it is a single RCT, done only with business students, and around 30% of participants did not show up for the surprise test (85 of the 120 completed it). A serious data point, not a closed verdict.

The PNAS experiment: the finished answer hurts, the hint does not

The largest of the three is that of Bastani, Bastani, Sungu and colleagues (PNAS, 2025): a field experiment with nearly 1,000 math students at a high school in Turkey. With access to GPT-4 during practice, practice scores rose 48% on the standard, ChatGPT-like interface (called GPT Base), and 127% on a tutor version with guard rails, which gave hints instead of handing over the answer (GPT Tutor).

Then came the exam, without access to the tool. The GPT Base group did 17% worse than the classmates who never had access. And the finding the headlines tend to cut: in GPT Tutor, the harm was largely mitigated. The study's own title warns: generative AI without guard rails can harm learning. With them, the picture changes.

The MIT study: those who write with an LLM engage less with their own text

Kosmyna and colleagues (MIT Media Lab, 2025) had 54 people write essays while wearing EEG electrodes, in three groups: with an LLM, with a search engine, and with the brain alone. Brain connectivity scaled down as external support increased: the strongest and most distributed networks appeared in the no-tool group; the weakest, in the LLM group. In the first session, 83.3% of the LLM group failed to correctly quote a sentence from their own essay, written minutes earlier, while the other two groups quoted with ease.

Mandatory context before you pass this number around the study group: it is a preprint, still without peer review, with 54 participants (18 in the final session, precisely the one most cited), and it measures engagement in the task, not permanent damage. The authors themselves maintain a public FAQ asking that the work not be cited as proof that AI "wrecks the brain," and there is already a published critical methodological comment (Stankovic et al., 2026). What the study supports is more modest, and still uncomfortable: when the tool writes, you engage so little that minutes later you do not recognize your own text.

Two of these experiments are broken down in video, number by number:

Why the brain loves it and the memory hates it

None of this is a new phenomenon created by AI. The mechanism has decades of literature behind it: fluency is not learning. When the material goes down smoothly (because you reread it, or because ChatGPT summarized it), the brain confuses "I recognize this" with "I know this."

The classic experiment is Roediger & Karpicke (2006, Psychological Science). Five minutes after studying, rereading the text four times won: 83% recall, against 71% for those who read it once and tested themselves three times. A week later, the ranking flipped: rereading plummeted to 40% and those who had tested themselves held at 61%. The cruelest data point in the study: restudying increased the students' confidence, without increasing long-term retention. High confidence with low memory has a name (illusion of fluency) and it is the default mental state of anyone leaving a session of ready-made summaries.

Robert Bjork calls the conditions that make the session harder now and the memory more durable later desirable difficulties (Bjork & Bjork, 2011). The discomfort of trying to remember is the learning itself: whoever cuts the discomfort cuts the memory. And common ChatGPT use consumes exactly that discomfort. Every summary requested, every answer received before your own attempt, is a retrieval your memory failed to perform.

Put plainly: asking ChatGPT for a summary is the rereading of 2026. In the review by Dunlosky and colleagues (2013), which ranked ten study techniques, summarizing, rereading, and highlighting already sat on the low-utility shelf, while testing yourself and spacing reviews sat on the high one. AI did not change the ranking; it just automated the wrong side of it. Anyone who wants to see the size of the bill that forgetting charges on any passive studying will find the full forgetting curve.

How to use ChatGPT to study: five prompts that reverse the direction

The same tool that consumes the effort also knows how to demand it. The five prompts for studying in ChatGPT below put the sweat back on the right side of the desk, and they all work on the free plan, today.

1. Produce first, ask for correction after. Close the material, write your explanation of the concept (on paper or in the chat itself), and only then ask: "Here is my explanation of X. Point out errors, inaccuracies, and what I missed." The first attempt is yours; the AI comes in as the grader, the role it is unbeatable at.

2. Ask for questions, not answers. "Ask me 10 questions about this chapter, one at a time, without showing the answer until I try." It is the testing effect with the boring part outsourced (coming up with questions) and the part that builds memory preserved (answering them).

3. Ask for holes, not praise. After explaining a topic, demand: "What cases break what I said? Where is my reasoning weak?" Models tend to agree with the user; instructing them to attack your explanation corrects the bias and turns the session into an interrogation.

4. Turn material into a test. Paste in your notes and ask for exam questions, or question-and-answer pairs for flashcards. A good card is a question that forces retrieval, and the friction of making cards by hand is exactly where most people give up on Anki.

5. Space the rounds. Questions generated today are worth more answered on the right date. With total study time held fixed, the right interval between sessions increased final recall by 64% (d = 1.1) in the study by Cepeda and colleagues (2008), with more than 1,350 participants; the rule of thumb is to space about 20% of the time until the exam. The how and the why are in the spaced repetition guide.

What never to delegate to AI

In any protocol, three deliverables remain yours:

The summary. Summarizing was already a low-utility technique when you were the one summarizing (Dunlosky et al., 2013); outsourced, nothing is left. The value of a summary was never in the final text: it is in the triage your head does to write it.

The first attempt. Solving the exercise, explaining the concept, writing the first version. That is where the memory works; everything after is correction. The GPT Base group in the PNAS experiment lost on the exam precisely because it skipped this step.

The retrieval. If the information has already passed through your head, the act of pulling it back is the training. Asking ChatGPT "what was the formula again?" before trying to remember is paying someone to work out for you.

Does ChatGPT's study mode solve it?

In July 2025, OpenAI launched Study Mode, available on the Free, Plus, Pro, and Team plans: instead of answering directly, the model leads through guided questions, hints, and self-assessment, in a Socratic method. In the grammar of this post, it is the right direction: it is the guard rails from the PNAS experiment turning into an off-the-shelf product.

Two caveats keep the verdict honest. There is still no published efficacy study on study mode, and a product description is not evidence. And the mode is optional: the finished answer is still a click away, and no interface protects against someone who wants a shortcut. The protocol above works in any mode, because the discipline lives in your prompts; no product installs it for you.

What to do today: studying with AI in five steps

  1. Write before opening the chat. Your explanation, material closed, five minutes.
  2. Paste it in and ask for a merciless correction. Errors, inaccuracies, what you missed.
  3. Ask for 10 questions with no answer key. One at a time, the answer only after your attempt.
  4. Turn the errors into new questions. What you got wrong today is the script for the next session.
  5. Schedule the return. Each question answered today needs a date to meet again, spaced.

The first four steps ChatGPT covers well. The fifth is where manual execution dies: deciding what comes back, calculating when, redoing the queue with every hit and miss. That back-office work is Sift's job: it turns what you study into a question you answer, instead of a summary you read, and brings each one back on the date when answering it pays off most, with the forgetting curve measured per user.

Frequently asked questions

Is studying through ChatGPT reliable?

For facts and figures, trust while distrusting: models get things wrong with the same fluency they get them right, and the confidence of the text is no sign of accuracy. Reliability rises when the AI's role is to examine: if it asks and corrects based on the material you pasted in, the source is still your material.

Is it wrong to study through ChatGPT?

No, and none of the studies cited say so. Studying with ChatGPT is good when it makes you produce; the PNAS experiment (2025) showed harm in the version that handed over the finished answer and broad mitigation in the version that gave hints. What is wrong is letting the AI sweat in your place.

Is ChatGPT good for studying for a competitive civil-service exam?

It is a good examiner of black-letter law: turning statutes into questions, demanding literal definitions, correcting your memory attempts. What it does not solve is the review schedule that a huge syllabus demands; that part is the backbone of a well-built study plan for the exam.

Does studying with AI work?

It works when the AI increases your retrieval effort (by asking, correcting, pointing out holes) and tends to come at a cost when it replaces that effort (summarizing, answering, writing for you). It is the consistent pattern across the three recent studies cited above, each with its own limitations.

ChatGPT is not going to leave your study desk, and it does not need to. The question to ask before each prompt is short: "will this force me to think more or to think less?" Sift, by design, only knows how to do the first thing: it takes what you study, hands it back as a question, and sets the reunion for the date when your memory needs it most. The sweat stays yours; the schedule stops being.

How we verified this article: every statistic links its primary source: the study, the year, the exact result. When a number has no study behind it, it doesn't make the text. Read our methodology.