Before using AI feedback, check it against the student's actual work
A teacher-focused workflow for using AI to flag possible error patterns while keeping evidence checks, instructional judgment, and final feedback with the teacher.
AI can help a teacher scan a set of responses and flag patterns that may deserve attention. But spotting a pattern is not the same as deciding what a student understands, what score they deserve, or what feedback they should receive.
A safer division of labor is narrow: ask AI to point to something worth checking and quote the exact words in the student’s response that support the flag. The teacher then verifies the evidence, discards weak suggestions, and writes the final feedback.
The examples below are fully synthetic. They are not outputs from a recorded test of a particular AI model.
Start with three visible pieces of student work
Suppose the question is: “Why does water collect on the outside of a glass of ice water?”
Three synthetic answers:
- Response A: “Water from inside the glass seeps through the glass because there is more water inside.”
- Response B: “Water vapor in the air meets the cold surface, cools, and condenses into droplets.”
- Response C: “The water appears because cold air can hold more water vapor, so the vapor turns into drops.”
Response C is a useful boundary case. It mentions air, vapor, and droplets, so a quick keyword scan might make it look close to the expected explanation. The causal claim still needs careful reading.
Give AI only the part that can be checked
Instead of asking, “Grade these three answers,” constrain the task:
Read the three student responses. For each one, identify no more than two points the teacher should check. Every point must include an exact quote from the student's response as evidence. If the response does not provide enough evidence to decide what the student understands, say “insufficient evidence.” Do not assign a score and do not write feedback to send to the student.
The desired output is not a grade sheet. It is a short list of evidence-backed questions for the teacher to inspect. For Response A, for example, a useful flag could point to the phrase “seeps through the glass” so the teacher can examine the student’s idea about where the outside water comes from.
Return to the original response before writing feedback
For every AI suggestion, do three checks. First, find the quoted words in the student’s actual response. If the quote is missing or inaccurate, discard the suggestion. Second, ask whether the suggestion describes what the student actually wrote or guesses at the student’s intention. A guess should not be treated as a finding. Third, compare the point with your own learning objective and assessment criteria.
Only after those checks should the teacher decide what to say. Rather than copying a label such as “misunderstands condensation,” the teacher might ask Response A: “If the glass has no holes, where else could the water on the outside come from?” That question comes from the teacher’s instructional judgment, not from handing the decision to AI.
The hard case: a plausible conclusion with faulty reasoning
Response C shows why keyword matching is not enough. The answer contains words associated with the topic, but the teacher still needs to inspect the relationship between the claims.
You can make that distinction explicit in the AI request:
For each response, separate the conclusion from the reasoning. If the conclusion appears plausible but the reasoning conflicts with the reference material I provide, flag it for the teacher to check. Do not rewrite the student's answer.
If the AI misses the problem in Response C, that is evidence that it cannot serve as the final filter. If it flags the issue correctly, that still does not prove it will do so reliably on other work. This constructed example is for checking the boundary of the workflow, not for claiming an error rate.
When this workflow should stop
Do not upload real student work to a tool that your school has not approved when the material includes names, student IDs, grades, health information, family circumstances, or other sensitive data. You can begin with synthetic examples or properly de-identified material to see whether the workflow is useful at all.
The workflow also fails its purpose if AI repeatedly misquotes students, infers intentions that are not in the work, or makes verification take longer than simply reading the responses yourself. In that case, remove the AI step rather than preserving it for its own sake.
UNESCO’s guidance on AI in education emphasizes a human-centered approach, human agency, and appropriate safeguards around data and educational decisions. That maps to a practical boundary here: AI may help direct a teacher’s attention, while the teacher remains responsible for interpreting the work and deciding on feedback.
Primary sources
- AI and the future of learning — UNESCO — UNESCO’s overview of a human-centered approach to AI in education and the role of educators.
- Guidance for generative AI in education and research — UNESCO — Primary guidance supporting the cautions about human agency, data protection, and context-appropriate use of generative AI.
The useful boundary is simple: AI can flag where to look; the teacher decides what the evidence means and what to say to the student.