All field notes
ClassroomJuly 30, 2026·6 min read

The Untimed Test Problem: Using AI to Create Flexible Assessments for Mixed-Ability Classes

When some students finish in 15 minutes and others need 45, assessments fail fast learners and exhaust slow ones. AI can generate parallel versions so everyone shows what they know.

By Ms. Park, Middle School Math & Honors

The problem with a classroom assessment:

Three students finish in twelve minutes. Three need forty. Two are still writing after time's up. One didn't even start.

You can't give everyone the same test for the same amount of time unless they all have the same processing speed and the same prior knowledge. If your class has mixed ability, the test is either too easy for some kids (they're bored, they finish early, they've checked out mentally) or too hard for others (they run out of time, they panic, their grades reflect anxiety, not understanding).

AI can solve this, partially, by generating parallel versions of the same assessment so fast learners get a harder version of the same concept and slower learners get a version that doesn't require as much reading or as many steps.

This isn't differentiation theater. This is actually different assessments for actual different readiness levels.

What parallel assessments look like

Same concept: solving two-step equations

Level 1 (shorter, fewer steps): "Solve: 2x + 3 = 11. Show your work."

Level 2 (same concept, more complexity): "Solve: 3(x + 2) = 24. Show your work. Explain what step you did first and why."

Level 3 (same concept, applied to context): "A class is buying shirts. Each shirt costs $8. After adding $5 for shipping, the total is $61. How many shirts are they buying? Write an equation and solve."

All three are assessing the same skill: solving two-step equations. But Level 1 finishes faster, Level 2 takes normal time, Level 3 takes longer because it adds context and explanation.

How to generate these with AI

You write the original assessment, then ask the AI:

"Here's an assessment on [topic]. Generate two other versions of this assessment that assess the *same skill* but at different complexity levels. One should be faster for a student who processes quickly but needs to confirm understanding. One should be deeper/longer for a student who needs the extra step of applying it to a real context. Keep the same rubric frame."

You get back three versions. One for each readiness level.

Then you give each student the version they're ready for.

The real move: it's still the same assessment

Here's the thing: the student doesn't feel like they got a different, easier test. They got *the* test. It just happens to be calibrated to their level.

This only works if:
- All versions hit the same learning objective
- The rubric is the same, not "easier students need less"
- Easier doesn't mean better grades for less work; it means the work matches their level

A student on the easier version should still get a B or an A by meeting the rubric for their version. They didn't get dumbed down. They got a version that lets them show understanding without getting paralyzed by the format.

Where this gets complicated

Three problems you have to solve:

Problem 1: They know they got a different test
They'll compare answers with peers, they'll realize the tests aren't the same. Solution: frame it as "you got the version matched to your learning level, just like your reading group has books at your level." It's normal. It's not shameful.

Problem 2: The grading has to be parallel too
If the easy test is easier to get an A on, students on the hard test feel punished. Solution: same rubric, same point scale, same percentages. "Solve and explain" is worth the same whether they're explaining two steps or three.

Problem 3: You have to track who got what
You can't use the results fairly if you don't know which version they took. Solution: note it on their paper or in your gradebook. Then you're comparing like to like.

Where this gets faster

Our [assessment blueprint tool](/assessment-blueprint) lets you build the main assessment, and our [differentiation plan tool](/differentiation-plan) helps you generate the parallel versions aligned to the same standard, so you're not building three separate tests from scratch.

The test of a flexible assessment that actually worked

Ask yourself: Did every student show what they know, or did some students fail the *format* while understanding the *concept?* If students in your class understand something but couldn't show it in the time/format you gave them, the assessment was the problem, not the student.

Parallel versions fix that.

Keep reading