Rubric Writing Guidelines
The standard, one section at a time. Search it, run the checker on your own rows, tick off the gates. From the one-page quick reference:
The ten sections
Three questions that decide most rows
Never
- Test a number the agent was handed§6
- Write a check the judge cannot see in the submitted work§2
- Make a Critical out of the golden’s layout, wording or ordering§7
- Let one row lean on another — every row stands alone§3
- Count the same mistake twice§5
- Add a second sentence, a contrastive tail, or a restatement of the verdict§4
First time? Work through the onboarding module ↗ — eight exercises, one per idea.
All sections§1 of 10
What you're doing and why
TL;DRAn AI judge grades the agent’s deck and model against your rubric, line by line — it sees the response and your rubric, nothing else. You write before you see how the agent did, and sub-task 1 is the pilot.
You are writing the answer key
An AI is given a finance brief and produces a deck and a model. Somebody has to decide whether what came back is any good — the same way every time. So it cannot be a judgement made fresh each time. It has to be written down in advance.
What you write down is a rubric. Then an AI judge opens the AI agent's deck and model, works through your rubric one line at a time, and marks each line pass or fail.
Where your rubric sits
Two things to notice.
You write the rubric before you see how the agent did. Agent outputs are generated in the background while you write. That is deliberate: if you write against an attempt you have already read, you end up grading that attempt rather than the task.
Sub-task 1 is the pilot. Its pass rates set the pattern every other sub-task follows. Time spent getting ST1 right is not perfectionism — it is the cheapest time in the project.
The six words we use precisely
| Word | What it means |
|---|---|
| Prompt | The brief the AI was given. One sub-task, one prompt. |
| Response | What came back: the deck, the model, or both. |
| Golden | The expert answer built in Phase 1 and checked before you start. It tells you which facts to expect. It does not tell you one single correct answer. |
| Criterion | Each line in the rubric. Each check represents one thing the prompt asks for, or clearly implies, that can be judged on its own. |
| Rubric | All the criteria for one sub-task. Ours run 5 to 20 rows. |
| Judge | The AI that grades the response against your rubric. It sees the response and your rubric. Nothing else. |
The guiding principle
If your rubric says the work passed and a partner reading the same deck says it didn't, the rubric is wrong — not the partner.
All sections§2 of 10
Before you write a row
TL;DRBefore any rows: read the prompt and golden together, list what the prompt asks for, write the intent in one sentence, sketch two or three other good builds — and check what the judge can actually see.
The habit that causes almost every defect
Criteria should always be written by looking at the prompt and golden in conjunction, never just one or the other. It is important to decide what the task is testing.
Everything downstream follows from that.
- Rows that grade the golden's layout as is.
- Rows that grade nothing.
- Rows that grade something nobody asked for.
1. List what the prompt asks for
Go through the prompt and write out every distinct thing it requires. Not the wording — the requirements.
Two failures come from skipping this. You grade something the prompt never asked for, and you leave something crucial it did ask for ungraded.
2. Write the intent in one sentence
Figure out, in your own words: what capability is this sub-task testing?
Example
Prompt: To support the take-private thesis for Northvale, the IC has asked us to build the market context and the operating analytics. Build out the 'market build' tab sizing off the market inputs, and the 'KPI' and 'KPI benchmarking' sheets that show how Northvale performs in the industry. This will form the base evidence for the growth and margin story underpinning the investment case.
What it is testing: can the agent size the market from the given inputs and turn Northvale's own financials into KPIs that stand up against the industry — evidence a growth and margin thesis can rest on.
| Crucial | Not crucial | Never test |
|---|---|---|
| The market size is built up from the inputs, not asserted | Top-down or bottom-up | The market inputs themselves — the agent was handed them |
| The build is live: change an input, the market size moves | Which KPIs are chosen, as long as they speak to growth and margin | |
| KPIs are calculated from Northvale's financials and tie back to them | Which peers make up the benchmark set | |
| Northvale and the peers are measured on the same KPI definitions | £m or £bn, tab layout, row order |
Notice what the crucial column contains. Not one figure. Every row is a property the work either has or does not have, and each one survives a completely different build. That is what a criterion should look like when the prompt leaves the choices open.
3. Sketch two or three other good builds
Ten minutes, roughly. What would a different analyst have produced from this same prompt?
When the only build you have ever looked at is the golden, every incidental choice in it starts to read as a requirement — a two-panel layout, a particular peer set, a specific ordering.
Anything that would fail those other builds is not Critical.
4. Check what the judge will actually receive
The judge opens the agent output(s). That is all.
It does not see the prompt. It does not see the input file. It does not see the golden.
If a check depends on any of those, the judge has nothing to look at and will guess. Decide this now, before you write — it is far quicker than deleting rows later.
All sections§3 of 10
The six columns
TL;DRSix columns, nothing else. Every row makes sense on its own; Section / Sheet matches the golden’s slide titles and tab names exactly, so the judge can find them; results are reported by Category. Five to twenty rows, spent on what matters most.
The file
Every rubric contains six components. Nothing else.
| Column | What it holds | Allowed values |
|---|---|---|
| # | Row number, for reference and reporting | Integers. No 4a, no 4.1 |
| Section / Sheet | The slide title or Excel tab, word for word to match the golden | The exact title, or Universal |
| Scope | Which part of the deliverable this checks | Deck / Model / Both |
| Rubric | The check itself | Full sentences that stand on their own. Two to four points at most |
| Category | What kind of check it is | Structural / Numeric / Logic / Methodology / Formatting / Consistency & Coherence |
| Tier | Whether this row gates — will the agent output fail if this criterion fails | Critical / Non-Critical |
The judge grades one row at a time and never sees how another row went. A row that leans on another has nothing to check, so it guesses. Every row has to make sense on its own.
Section is how the judge finds the slide
Write the real title, exactly as it appears. Not "Slide 1". Not "the returns slide". The real title is the golden's: Section / Sheet matches the golden's slide titles and tab names, word for word.
The judge does not have a page number to work from. Given "Slide 1" it grades whichever slide it lands on, and the result is noise.
Universal is for checks that apply to the entire deliverable regardless of which slide — a house style rule, a footnote convention. If your check names specific slides, it is not Universal. It belongs on those slides, or the judge has nothing to go and look at.
Category is not admin
Results are grouped and reported by Category. A numeric check filed as Structural vanishes from the numeric results the client actually reads.
| Category | The question it answers |
|---|---|
| Structural | Is the element there? |
| Numeric | Is the number right? |
| Logic | Does the claim hold? |
| Methodology | Was it built the right way? |
| Formatting | Does it follow the stated presentation rules? |
| Consistency & Coherence | Do the same figures agree wherever they appear? |
Set the scope carefully
Make sure the deliverable tests everything that the prompt asks for. If the agent output is expected to produce only a deck, do not include scope = model. If the agent output asks for both, ensure there are rubrics testing both.
How many rows
Five to twenty per sub-task. At least five Criticals.
Twenty sounds generous until you count what a sub-task actually covers. If a sub-task spans two model tabs and four slides — six areas, it may have roughly three checks each.
Spend the rows on what matters most, not evenly.
All sections§4 of 10
Writing the criterion
TL;DROne sentence saying what must be there, with the expected answer in the row. Name what you point at, four or five numbers at most — no second sentence, no contrastive tails.
The shape
One sentence saying what must be there. Never include a second that rules out a wrong reading.
Wrong
The Returns slide states an IRR and a money multiple for the base case. In other words, both figures appear for the base case specifically, not only for whichever case the selector is left on.
Right
The Returns slide states an IRR and a money multiple for the base case.
The first sentence is the check. The second should not exist because a judge will not pass a slide showing the upside case.
Put the expected answer in the row
The Returns slide states a base case IRR of 25.1% (acceptable range 23.8% to 26.4%).
If the judge has to go and work out what the right answer is, it will sometimes come back with a different one — and then the same submission passes in one run and fails in the next.
Be specific, but do not hand over the answer
"The slide analyses the margin trend" cannot be graded. Two people will read it two ways. Say what has to be on the slide.
But there is a line. This goes too far
...explains the margin expansion (mix shift, pricing power, operating leverage)
That bracket tells anyone holding the rubric exactly what to write. Be specific about what must appear, not about the thinking you are trying to grade.
This is what good looks like
The commentary attributes the FY2025A EBITDA margin movement to at least two specific operational drivers, each named and quantified.
The judge counts named drivers and checks each carries a number. Anyone can pass it with a correct analysis; nobody can pass it by copying the rubric.
Name what you are pointing at
The judge reads one row at a time, with no memory of the others. A row saying "that", "it" or "they" gives it nothing to point at, so it guesses.
As delivered
...and attributes that return level to the group's decentralised structure, asset-light model and serial acquisition strategy.
Rewritten
The slide states approximately 20% average ROCE over the cycle and attributes that ROCE level to at least two of the group's asset-light model, decentralised structure and serial acquisition strategy.
Two fixes there. The pronoun is replaced by the thing itself. And "at least two of" means a slide giving two good reasons instead of three still passes — which it should.
Four or five numbers per row at most
With the year written next to each one.
A judge asked to verify nine figures verifies two or three and assumes the rest. An assumed check is not a check.
Two cheap tests before you move on
Read it aloud. If you cannot say in one sentence what a passing slide or model would contain, it is not finished. Twenty rows takes about five minutes.
Cover half of a criterion. Would anything now pass or fail differently? If not, that half is decoration.
Watch for tails like "rather than...", "instead of...", "as opposed to..." — they read as though the criterion is grading two things at once, and they push the judge around. When you need to rule something out, say what must be true instead.
Watch for a second sentence that just restates the verdict: if the date must be 15 October 2025, "any other date fails" adds nothing.
Do not use AI to draft, write the criteria manually — it saves hours of rework and catching errors after.
All sections§5 of 10
How small is one criterion
TL;DROne test decides whether a check is one row or two. The rest of the section is what going wrong in each direction looks like — and why the same thing is never checked twice.
The test
Would you ever want to pass one part and fail the other?
If no, it is one criterion. If the report should be able to tell them apart, it is two.
That is the whole rule. The rest of this section is what going wrong in each direction looks like.
Split too far and the rubric stops working
Four rows
- C1 — the chart covers FY2021A to FY2025A
- C2 — the series is sales by business area
- C3 — the chart is indexed, not absolute
- C4 — the index base is FY2021 = 100
One row
- C1 — The chart shows sales by business area for FY2021A to FY2025A, indexed to FY2021 = 100.
A chart indexed to FY2022 instead of FY2021 tells a different growth story. Against the four-row set it still scores three out of four. More rows, less use.
Split too little and the judge leans towards pass
The headline financials present FY2021A to FY2025A sales of SEK 21,715m, SEK 27,016m, SEK 31,835m, SEK 32,544m and SEK 32,229m, alongside at least one profitability measure (such as EBITA, EBITA margin or ROCE) shown for each of those years.
Two rows
| Category | Rubric |
|---|---|
| Numeric | The headline financials present FY2024A sales of SEK 32,544m and FY2025A sales of SEK 32,229m. |
| Structural | The headline financials show at least one profitability measure for each of FY2021A to FY2025A. |
Two things happened there.
Three of the five figures went. That looks like lost coverage and is not. Two checks that actually happen beat five that may not.
"Sales figures wrong" and "no profitability measure shown" became separate rows, because they are different problems with different fixes, and the report should say which one occurred.
A long criterion feels rigorous. It is the opposite: a judge facing a wall of conditions it cannot sort out leans towards a false pass.
Never check the same thing twice
One Critical failure already fails the sub-task, so double counting changes nothing about pass or fail. It corrupts the report. "Failed #2 and #3" reads as two separate things going wrong.
Giving every fact one owner helps. If one row owns the segmental revenue mix, the row covering the EBITDA margin states the margin and its drivers without restating the mix.
All sections§6 of 10
Numbers
TL;DRTolerance follows where the number came from, not how confident you feel: extracted — rounding only; calculated from provided figures — about ±1%; worked out by the analyst — about ±5%; handed to the agent — never tested. Write every range out in full.
Tolerance follows where the number came from
Not how confident you feel.
| Where the number comes from | Tolerance |
|---|---|
| Extracted from the VDR or source documents | Tight — rounding only |
| Calculated from figures the prompt fixed or the VDR provided | About ±1% |
| Worked out by the analyst: ratios, valuation, returns | About ±5% |
| Handed to the agent by the prompt or the input model | Never tested at all |
| Choices left open by the prompt | Fix them in the prompt or input file, or widen everything downstream |
Never test a number the agent was handed
If the prompt fixes WACC at 9.4%, do not check that the model says 9.4%. It will. The agent was told.
Check what 9.4% produces
The DCF states an enterprise value of 1,098 (acceptable range 1,087 to 1,109).
A row checking that a handed figure was copied across passes automatically. It measures transcription, not capability, and it inflates the set. Twenty rows where six pass for free should be a fourteen-row set.
Write the range out in full
Not "within ±5%". The number and both bounds, in the row
...an FY2025A EBITDA margin of 18.6% (acceptable range 17.7% to 19.5%).
The judge should never have to do arithmetic to find out what passes, everything should clearly be laid out.
Tight is not the same as rigorous
People default to narrow ranges because narrow feels careful. On a figure an analyst worked out, a narrow range fails work that is right.
A range set too wide passes bad work. Set too narrow it fails good work. Neither is the safe option, which is why the ladder decides it.
The strongest checks may not carry numbers
- Sources equal uses.
- The case selector works.
- The model moves when a driver changes.
- Historical columns hold values, and forecast tabs pull from them by formula.
These are impossible to fake and take a second to grade. Most people writing a first rubric never think of them. Use them.
All sections§7 of 10
Critical or Non-Critical
TL;DRA Critical gates: one failure fails the whole sub-task. A row is Critical only if it is Mandatory, Objective and Stable — miss one and it is Non-Critical. At least five per sub-task, and a strong attempt should still fail about one in five.
What "gate" means
A criterion gates when failing it fails the whole sub-task on its own — no matter how many other rows passed.
That is what the Tier column decides. Critical gates. Non-Critical does not: it is recorded, reported and counted, but on its own it cannot sink the sub-task.
So a 20-row rubric where 19 rows pass and one Critical fails does not score 95%. It fails.
Three tests, all three required
1. Mandatory: Every correct answer has it. Work that fails this is wrong, not just different.
2. Objective: You can check it directly. Visible in the deck or the model, with the expected values written into the criterion.
3. Stable: Multiple judges would agree. Separate LLMs-as-a-Judge would reach the same verdict for the same output on a given criterion.
Miss any one of the three and it is Non-Critical. "When in doubt, leave it out."
Is the work wrong, or just different?
This is the question that decides most rows.
As delivered — marked Critical
The slide presents a two-panel company overview with exhibits of FY2025 sales by business area and FY2025 sales by geography, and a headline financials block.
Rewritten
The company overview presents an exhibit of FY2025 sales by business area, an exhibit of FY2025 sales by geography, and a headline financials block.
The prompt never asked for two panels. The golden just happened to have two. As a Critical, a submission showing the same three exhibits stacked — or in three panels — fails, and unfairly fails the entire sub-task with it.
Never make a Critical out of a choice the golden happened to make: e.g. which evidence it used, which multiple, how it laid the slide out. Another sensible answer has to be able to pass.
Stated in the prompt, but still Non-Critical
If something is written in the prompt, someone put it there for the agent, and the rubric should have a view on it. That does not make it Critical.
E.g. A brief asks for navy primary and muted red for risks. Grade it — as a Formatting row, Non-Critical. A deck that gets every number right and uses the wrong accent colour is not wrong. It is imperfect.
The line: does its absence make the work wrong, or just less polished?
How many, and how hard
At least five Criticals per sub-task.
For calibration: a strong AI attempt should still fail roughly one Critical in five. If it clears all of them, the set is not discriminating and the pass rate will tell you nothing.
Be careful about inflating the Critical count. Under a gate, every extra Critical makes the sub-task harder to pass — five Criticals at that difficulty gives around a one-in-three task pass rate; ten makes it roughly one in ten. Mark as Critical only what genuinely deserves to fail the whole thing.
All sections§8 of 10
Test what you wrote
TL;DRA good criterion flips: PASS on work that is right, FAIL on the same work once you deliberately break it. Testing only that the golden passes proves nothing.
What "flip" means
A good criterion flips: PASS on work that is right, FAIL on the same work once you break it.
Testing only that the golden passes proves nothing — a row saying "the slide exists" passes the golden too. You learn whether a criterion works by deliberately damaging the build and checking the verdict flips the other way.
If it does not flip, the row is not checking anything, and it will pass every submission you ever put through it.
The flip test
Three steps. Run it on every Critical.
- Write the criterion. Decide for yourself whether the golden satisfies it.
- Have a judge grade it. If it agrees with you, carry on. If not, suspect your wording first.
- Break the build deliberately — hardcode the figure, delete the exhibit — and grade it again. This time it has to fail.
Why step 2 alone is not enough
Step 2 only shows that good work can pass. A criterion that passes everything is worth nothing. Step 3 is what tells you the row can discriminate.
| Version | Criterion | What happens |
|---|---|---|
| v1 | The summary revenue line ties to the revenue build. | A model with a hardcoded summary passes. The two figures do match. |
| v2 | The summary revenue line is a formula pointing at the build tab, and moves when a driver changes. | The hardcoded model fails. A live one passes. |
"Ties to" meant far less than the writer thought it did. Only breaking the build on purpose exposed that.
When you and the judge disagree, the criterion may be wrong
Not the judge. Before you argue with the verdict, reread the row and ask what else it could mean. Most disagreements are a word doing less work than you assumed — "ties to", "reflects", "is consistent with", "supports".
All sections§9 of 10
Before you hand over
TL;DRFirst the five-minute read, then the five gates: the golden passes every Critical, two readers read each row one way, fakes fail, three runs agree on 90% of criteria or better, and it is still hard enough.
First, the five-minute read
Twenty rows takes about five minutes and catches most of what goes wrong.
Then the five gates
| Gate | What it means | |
|---|---|---|
| 1. Grade the golden against your own set | It must pass every Critical. If it does not, either the golden is wrong or your rubric is too strict — and you have to understand which | |
| 2. Two-reader test | If two reviewers could read a criterion two ways, reword it | |
| 3. Try to break it | Ask what an empty file or a confident-looking fake would score. Both should fail the number and method rows | |
| 4. Check three runs agree | Grade the Criticals three times and compare. We want the three runs to agree on 90% of criteria or better | |
| 5. Check it is hard enough | A strong AI attempt should still fail about one Critical in five |
Inter-judge agreement
Grading is not one opinion. The same Criticals get graded three times, and we look at how often the runs land on the same verdict. That share is inter-judge agreement, and we want it at 90% or better.
Low agreement is not the judge being unreliable. It is a row that can be read more than one way, so different runs read it differently. Agreement is a measure of your wording, not of the model.
Two things fix it. Name what you are pointing at, and put the expected value in the row.
All sections§10 of 10
A worked example
TL;DROne invented sub-task graded end to end: the thirteen-row rubric, the decision behind each row, and what is deliberately not there.
As an example, everything here is invented: company, figures, prompt.
The sub-task
Northvale Group plc — take-private IC memo.
In the provided Excel, build Group Revenue by Segment and Group EBITDA by Segment on the Charts tab.
Then build two slides.
- Slide 1, Business Model & Historical Operating Performance: revenue and growth chart, EBITDA and margin chart, FY2025A segmental revenue mix, at least six commentary points.
- Slide 2, Historical Profitability & Capital Structure: PBT and net income with margins, Net Debt / EBITDA, reported FCF / EBITDA, at least six commentary points.Every commentary slide carries a one-line takeaway and a source footnote. Navy primary, muted red for risks. Sources are in the VDR.
Intent: can the agent pull Northvale's five-year segmental history out of the VDR, carry it into a working model, and present the growth and profitability story at IC standard.
The rubric
| # | Section / Sheet | Scope | Rubric | Category | Tier |
|---|---|---|---|---|---|
| 1 | Charts | Model | The Charts tab presents Group Revenue and Group EBITDA by Segment for FY2021A to FY2025A across the Clinics, Diagnostics and Retail & Products segments. | Structural | Critical |
| Charts | Model | The segment values in the Group Revenue by Segment chart sum to group revenue in each of FY2021A to FY2025A. | Consistency & Coherence | Critical | |
| The decision: No figure at all. Impossible to fake, takes a second to grade. If the segments do not sum, the numbers are wrong somewhere | |||||
| Charts | Model | The EBITDA margin series is a formula calculated from the revenue and EBITDA figures on the Charts tab, and moves when a segment value is changed. | Methodology | Critical | |
| The decision: Tests how it was built, not what it says. A hardcoded margin row passes a "the margin is 18.4%" check and fails this one | |||||
| 4 | Business Model & Historical Operating Performance | Deck | The revenue chart states FY2023A revenue of £544m and FY2024A revenue of £581m. | Numeric | Critical |
| Business Model & Historical Operating Performance | Deck | The revenue chart states year-on-year revenue growth for each of FY2022A to FY2025A, with FY2025A growth of 5.3% (acceptable range 5.2% to 5.4%). | Numeric | Critical | |
| The decision: 5.3% is calculated from figures the VDR provided, so ±1% | |||||
| Business Model & Historical Operating Performance | Deck | The FY2025A segmental revenue mix states Clinics at 60.0% (acceptable range 59.4% to 60.6%), Diagnostics at 24.0% (23.8% to 24.2%) and Retail & Products at 16.0% (15.8% to 16.2%). | Numeric | Critical | |
| The decision: The three percentages are one fact — you would not pass Clinics and fail Diagnostics. Calculated, so ±1% | |||||
| Business Model & Historical Operating Performance | Deck | The commentary block contains at least six points. | Structural | Non-Critical | |
| The decision: Two rows, one per slide, so the report can say which commentary block fell short. Six points was asked for, but five good points is not wrong — Non-Critical | |||||
| Historical Profitability & Capital Structure | Deck | The slide states FY2025A Net Debt / EBITDA of 3.5x (acceptable range 3.3x to 3.7x). | Numeric | Critical | |
| The decision: A ratio the analyst works out, so ±5%. That is why it reads 3.3x to 3.7x and not 3.5x flat | |||||
| 9 | Historical Profitability & Capital Structure | Deck | The slide states FY2025A reported FCF / EBITDA of 55% (acceptable range 52% to 58%). | Numeric | Non-Critical |
| 10 | Historical Profitability & Capital Structure | Deck | The slide states FY2025A PBT margin of 8.1% (acceptable range 8.0% to 8.2%) and FY2025A net income margin of 6.1% (acceptable range 6.0% to 6.2%). | Numeric | Non-Critical |
| Historical Profitability & Capital Structure | Deck | The commentary block contains at least six points. | Structural | Non-Critical | |
| The decision: Two rows, one per slide, so the report can say which commentary block fell short. Six points was asked for, but five good points is not wrong — Non-Critical | |||||
| Universal | Deck | Every commentary slide carries a one-line takeaway and a source footnote. | Formatting | Non-Critical | |
| The decision: The prompt asked for them, so they are graded. Their absence makes the deck less polished, not wrong — Non-Critical | |||||
| Universal | Deck | Risk callouts are shown in muted red and the primary palette is navy. | Formatting | Non-Critical | |
| The decision: The prompt asked for them, so they are graded. Their absence makes the deck less polished, not wrong — Non-Critical | |||||
Why the rows look like that
What is not here, and why
No Logic row. This sub-task asks for figures and charts, not for a claim. Nothing to test.
No row about layout. The golden puts the segmental mix bottom-right. A deck that puts it top-left is a different deck, not a wrong one.