Skip to content

All articles

Agile Classrooms Blog

How I Use Rubrics To Get Better LLM Work

Use a simple rubric loop to define quality before prompting, score AI output against visible criteria, and revise the weakest dimension first.

Article cover for How I Use Rubrics To Get Better LLM Work, showing a visible scoring rubric used to improve AI-supported work.

You prompt an LLM for something useful. It replies. You scan the output, feel the gap between the words and the work, and then you either accept it because it sounds polished or ask for a vague revision like "make it better."

That is not a quality system. It is a hope loop. Use rubrics instead.

Define quality before you ask for help. Inspect the result clearly. Revise the weakest part, not just prompt by feel.

Why Rubrics Help Before You Prompt

A rubric forces a question most prompts skip: What does good actually look like? Not "better." Not "more polished." But good in this specific situation, for this specific audience, doing this specific job.

This matters because a tool cannot reliably hit a quality target you have not defined. If the work needs to be clear, specific, audience-aware, concise, and useful for a next decision, name those dimensions. Use the same target to shape the request and inspect the result.

For a simple rubric, use four to six dimensions. Each gets a short description and a score. The score isn't magic, but it shifts the conversation from "I don't like it" to "clarity is a 4, specificity is a 2, and the next revision needs to fix specificity without wrecking the structure."

Why Rubrics Help After The Output

Work can sound competent while missing the actual job. A draft can be fluent yet generic. A caption can be short yet context-free.

A cover concept can look attractive yet fail to capture an article's tension.

Rubrics make those failures visible. They let the human in the loop inspect work against criteria, not just vibes. This doesn't remove judgment; it gives judgment a surface to work on.

The loop is simple:

  1. Define quality dimensions.
  2. Set a minimum score.
  3. Produce the work.
  4. Score the output.
  5. Identify the weakest dimension.
  6. Revise against that dimension.
  7. Score again.

The critical move is the weakest-dimension revision. Stop asking for "make it better." Instead, demand a targeted improvement: make the example more concrete, sharpen the audience fit, clarify the CTA, or make the claim verifiable.

A Real Article Rubric Example

Here is the kind of rubric I use when working on an article. This is not meant to be fancy. It is meant to make quality visible enough that the next edit pass has a target.

For my article rubrics, I score each dimension from 1 to 5:

  • 1 = missing or actively harmful
  • 2 = present but weak
  • 3 = usable but needs work
  • 4 = strong enough
  • 5 = excellent

For an article draft, my rubric might look like this:

  • Clarity: Can a reader understand the main idea without rereading?
  • Structure: Does the article build in a useful order, with each section earning its place?
  • Specificity: Does it use concrete examples, actual constraints, and real tradeoffs instead of generic advice?
  • Voice fit: Does it sound like me, or does it slip into bland generic language?
  • Usefulness: Does the reader leave with a move they can actually try?
  • Engagement: Is there enough tension, payoff, or recognition to keep someone reading?

Then I use a simple quality threshold: the overall score should be at least 4.5, and no single rubric item should be below 4. If one item is below 4, that becomes the next improvement target.

For example, if a draft scores 5 on clarity, 4.5 on structure, 3 on specificity, 4 on voice fit, 4.5 on usefulness, and 4 on engagement, it is not ready yet. The problem is not the whole article. The problem is specificity.

So the next revision request should be specific:

Revise this article to improve specificity from 3 to at least 4. Add concrete examples, make the tradeoffs more visible, and do not change the structure unless needed.

That is the practical value. The rubric does not just judge the work. It tells me where the next improvement should happen.

Educators Already Know This Move

If you're an educator, none of this feels foreign. You already use rubrics to define expectations, assess work, and give students feedback they can act on. The same pattern works with AI.

Do AI like you teach. Define what quality means. Score the output.

Name the weakest criterion. Ask for a revision that improves that criterion. The familiar teaching move becomes a better AI workflow.

This also builds confidence. Much AI use feels opaque because the output can sound confident even when the work is weak. A rubric gives the teacher a familiar way to inspect the output without needing to become a technical AI evaluator.

Product And Project People Know This Too

Product and project people follow a similar pattern, even with different language. Acceptance criteria, Definition of Done, review gates, and quality checks all try to prevent "done" from meaning "someone just produced an artifact."

Rubrics are especially useful when work is ambiguous. Most Definitions of Done and many acceptance criteria are pass/fail. That matters.

A story either has the export or it does not. A review either happened or it did not. But a lot of important work lives on a spectrum: clear enough, useful enough, specific enough, trustworthy enough, ready enough.

Acceptance criteria tell you if the artifact exists; rubrics tell you if it is usable. A pass/fail gate can confirm the artifact exists. A rubric can show exactly how to improve it until it is good enough.

If a discovery summary scores 4.5 on clarity but 2.5 on evidence quality, the team does not need a vague debate about whether it is "good." They know the next improvement target.

Try One Rubric Loop

The next time you use an LLM for work that matters, stop starting with the prompt. Start with the quality standard.

Pick four to six dimensions. Set a minimum score. Ask for the work.

Score the result. Then revise the weakest dimension first. If the work still misses the threshold after a few passes, that's useful information.

It means the human needs to rethink the brief, the source material, or the judgment call.

What would change in your next AI-supported task if you had to define "good" before you started?

Scoring work against visible criteria is the same habit that makes classroom evidence useful while learning is still in progress. For a related look at using evidence to drive the next attempt, read Formative Assessment Turns Learning Into Iteration.

🏅 Earn 0.25 SEUs/PDUs for reading this! Renew your PMP, CSM, or CSPO certification.

Subscribe for support and tools

Templates, routines, and ideas for building student-led classrooms.