I used to spend three hours every Sunday grading essays. Now it's about forty-five minutes, and the feedback my kids get is more specific than it used to be. Here's the workflow.
What AI is good at
- Grouping. Paste a batch of student responses (with names removed) and ask for the three most common errors. You'll get a real answer.
- Drafting comments from a rubric. "Write a 3-sentence comment for a response that hits criteria 1 and 2 but misses criteria 3." That's a boring, structured task, exactly what the model is best at.
- Suggesting next steps. "What is the single most useful thing this student could work on next?" beats "good effort" every time.
What AI is bad at
- Reading student voice. It will flatten it. Your kids' weird, funny, half-formed ideas are the point. Don't let the model iron them out.
- High-stakes decisions. Grades that go on report cards, comments to parents about behavior, anything that touches an IEP, you sign off, every time.
- Anything with a student's name in it. Don't paste that into a consumer tool. Anonymize first.
The Sunday-night workflow
1. Anonymize the batch (find/replace names to "S1, S2, S3").
2. Ask the model to group responses by error pattern.
3. For each pattern, ask for a 3-sentence comment draft aligned to the rubric.
4. Read every draft. Rewrite the ones that don't sound like you. Delete the ones that miss the kid.
5. Paste back into your gradebook.
The rule of thumb: the model drafts, you decide. The 45 minutes is what's left after you've cut everything that isn't right.


