A Polished Submission Can Conceal the Learning
Separate the quality of the artefact from the capability attributed to the learner.
Begin With the Claim, Not the Assignment Format
Ask what the mark is supposed to mean before deciding how students will submit.
| Conventional task | Intended learning claim | What the final product can show | What may still be missing |
|---|---|---|---|
| Literature-review essay | Compare evidence, recognise disagreement, and justify a synthesis. | A structured argument and selected references. | How sources were found, why some were rejected, and whether the learner can defend the synthesis. |
| Mathematical modelling report | Choose assumptions, construct a model, test it, and interpret its limits. | A model, calculations, graphs, and a written interpretation. | Why assumptions were chosen, whether alternatives were evaluated, and whether the model transfers to changed conditions. |
| Programming project | Design, implement, test, and explain a solution. | Running code and documentation. | How failures were diagnosed, whether dependencies are understood, and whether the solution can be modified under a new constraint. |
| Policy brief | Select evidence for an audience and make a defensible recommendation under uncertainty. | A concise recommendation and supporting evidence. | How trade-offs were evaluated and whether the recommendation changes when an assumption changes. |
The BMLabs Evidence-of-Learning Chain
Six linked decisions turn a vague AI rule into a defensible assessment design.
Build the chain in order
- 1
State the learning claim
Name the knowledge, reasoning, performance, judgement, or transfer capability being assessed. Specify the conditions and standard.
- 2
Set the AI boundary
Decide which cognitive work AI may support, which work must remain attributable to the learner, and whether critical AI use is itself an outcome.
- 3
Choose complementary evidence
Select an appropriate combination of product, process, performance, and practice evidence. Every item should answer a specific question about the learning claim.
- 4
Add an independent checkpoint
Where the inference remains weak, add a proportionate oral explanation, live decision, changed-input problem, annotated reasoning step, practical demonstration, or supervised component.
- 5
Test fairness and feasibility
Examine accommodations, language demands, access, privacy, infrastructure, student workload, staff workload, class size, and disciplinary context.
- 6
Define the judgement and review rule
State what evidence is sufficient, how contradictory evidence will be interpreted, and what information will be collected to improve the task after use.
Detection-led response versus evidence-led design
| Topic | Detection-led response | Evidence-led design |
|---|---|---|
| Starting question | Did the student use AI? | What capability must the student demonstrate, and what evidence would make it visible? |
| Main mechanism | A detector score, stylistic suspicion, or a general prohibition. | A stated AI boundary plus complementary evidence tied to the learning claim. |
| Declaration | Treated as sufficient proof of compliance. | Supports transparency but does not replace task design or evidence. |
| Uncertainty | Converted into suspicion about authorship. | Addressed through proportionate additional evidence and cautious interpretation. |
| Shelf life | Depends heavily on current tool behaviour. | Remains useful because it is organised around learning outcomes and inference. |
Decide What AI May Do Before Students Begin
Permission should be task-specific, intelligible, and connected to the learning claim.
Decision framework
Choose an AI boundary
Which description best matches the role of AI in this component?
This tool supports planning; it does not create institutional permission. Current university, programme, accessibility, privacy, and disciplinary rules take precedence.
Worked Example: Redesigning a Mathematics Modelling Assignment
Keep AI available for appropriate exploration while making assumptions, judgement, calculation, and transfer visible.
| Stage | Student activity and AI boundary | Evidence generated | Reasonable inference |
|---|---|---|---|
| Problem framing | The student identifies the decision, audience, variables, exclusions, and data limits before using AI. | A short problem frame and variable map. | The student can define the modelling problem and recognise missing information. |
| Assumption exploration | AI may suggest alternatives. The student records suggestions used, rejects at least one with a reason, and identifies evidence needed. | An assumption ledger with rationale, consequence, and confidence. | The student can evaluate suggestions rather than merely accept them. |
| Model construction | AI-assisted code or algebra is permitted, but variables, transformations, units, and test cases are annotated. | Annotated model, test cases, and verification record. | The student has inspected key operations and connected the representation to the problem. |
| Model comparison | The student compares two approaches using fit, interpretability, data demands, sensitivity, and intended use. | A comparison table and justified model choice. | The student can exercise evaluative judgement under disciplinary constraints. |
| Final report | AI may assist with organisation and language where policy permits. Sources, calculations, limitations, and material assistance are checked and declared. | The complete modelling report and recommendation. | The student can integrate evidence into a disciplinary product, subject to the limits of product evidence. |
| Changed-input checkpoint | The student receives a new occupancy assumption and explains which parts of the model and recommendation change. | A short recalculation, annotated sketch, or equivalent explanation. | The student can transfer selected reasoning to a changed condition and explain a consequential decision. |
| Reflection and review | The student identifies one useful AI suggestion, one rejected or corrected suggestion, and one capability needing practice. | A concise evaluative reflection linked to the evidence record. | The student can examine the role of AI, while the reflection remains a self-report rather than proof. |
Keep Learning Visible Without Creating Surveillance
Collect the smallest amount of process evidence that materially strengthens the judgement.
Local checklist
Assessment launch checklist
Progress is stored only in this browser.
Questions educators commonly ask
Should every AI-permitted assessment include an oral defence?
No. Add an oral component only when spoken explanation provides relevant evidence and the format is fair. A short transfer task, annotated correction, practical demonstration, or sampled interview may be more appropriate.
Is an AI-use declaration enough?
A declaration supports transparency but does not by itself show that the intended learning occurred. Pair it with task evidence.
Can an AI detector identify which work needs checking?
Detector output should not be treated as proof of authorship or misconduct. Build ordinary assurance into the assessment and follow current institutional procedures when a concern arises.
Must students submit every prompt?
Usually not. Complete logs can be intrusive, burdensome, and weakly connected to the learning claim. Request only the record needed for transparency or evaluation.
How can this work in a very large class?
Use programme-level planning, common rubrics, brief transfer tasks, tutorial sampling, structured process artefacts, and targeted follow-up. Do not add stages the teaching team cannot apply consistently.
What if the product and checkpoint conflict?
Use the pre-announced judgement rule, consider whether both pieces target the same outcome, examine accessibility and context, and offer a proportionate additional demonstration where policy permits.
The compact design standard
- State the capability before choosing the submission format.
- Specify the AI boundary at the level of task components and cognitive purposes.
- Use complementary product, process, performance, or practice evidence.
- Add an independent checkpoint only when it materially strengthens the inference.
- Test privacy, access, accessibility, workload, and disciplinary fit before launch.
- Interpret evidence proportionately and review the design after use.
