Learn Lab

Allow the Tool. Keep the Learning Visible.

A practical framework for deciding what AI may support, what students must still demonstrate, and how several pieces of evidence create a defensible judgement.

12 min read Published 20 July 2026Reviewed 20 July 2026
On this page

The report is coherent, well referenced, and technically correct. The student receives a high mark. Yet the teacher still cannot answer the most important question: what can this student now explain, decide, or do without the report in front of them? Generative AI makes that tension visible, but it did not create the underlying assessment problem. A submitted product has always been a proxy for learning. When AI can contribute substantially to that product, educators need a more explicit chain from the intended learning to the evidence on which a judgement is based.

Reader problem

A teacher may permit generative AI but still assess only a polished final submission. The product can contain substantial machine contribution, leaving weak evidence about which knowledge, reasoning, decisions, or disciplinary capabilities the student can demonstrate.

Expected outcome

The reader can redesign an assessment so that permitted AI assistance is transparent and purposeful while student explanation, judgement, disciplinary reasoning, and transfer remain visible through complementary evidence.

Author

Dr. Bivash Majumder

Assistant Professor in Mathematics, Prabhat Kumar College, Contai

Higher-education mathematics teaching, assessment experience, academic research, and research communication, combined with current peer-reviewed assessment-validity research and authoritative guidance on generative AI in education.

Editorial basis

Original contribution

The six-link BMLabs Evidence-of-Learning Chain, a worked mathematics modelling redesign, an AI-boundary decision tool, and a proportionate assessment-launch checklist. The chain is an editorial synthesis for planning and reflection, not a validated measurement instrument.

Risk category: education-technology. Review interval: 6 months.

A Polished Submission Can Conceal the Learning

Separate the quality of the artefact from the capability attributed to the learner.

Imagine a student submitting an elegant policy brief, laboratory discussion, coding project, or mathematical report. The work may deserve attention as a product. It does not, by itself, show how the student framed the problem, which alternatives were considered, whether errors were recognised, or whether the same ideas can be applied when conditions change. Assessment becomes fragile when the submitted object is treated as if it transparently contains the learner's knowledge. Generative AI increases the possible distance between product quality and independent capability, so the intended inference must be made explicit.

Begin With the Claim, Not the Assignment Format

Ask what the mark is supposed to mean before deciding how students will submit.

A learning outcome is the capability a course intends students to develop. An assessment task is the situation created to elicit evidence of that capability. The submitted artefact is one result of performing the task. The assessor then interprets that evidence and makes a judgement. These elements are connected but not identical. If the outcome is evaluative judgement, a final essay may display a conclusion without revealing how competing evidence was weighed. If the outcome is mathematical modelling, a final equation may hide how assumptions were chosen. Begin with a sentence such as: on completion, the student must be able to do X, under Y conditions, to Z standard.
From conventional task to visible learning claim
Conventional taskIntended learning claimWhat the final product can showWhat may still be missing
Literature-review essayCompare evidence, recognise disagreement, and justify a synthesis.A structured argument and selected references.How sources were found, why some were rejected, and whether the learner can defend the synthesis.
Mathematical modelling reportChoose assumptions, construct a model, test it, and interpret its limits.A model, calculations, graphs, and a written interpretation.Why assumptions were chosen, whether alternatives were evaluated, and whether the model transfers to changed conditions.
Programming projectDesign, implement, test, and explain a solution.Running code and documentation.How failures were diagnosed, whether dependencies are understood, and whether the solution can be modified under a new constraint.
Policy briefSelect evidence for an audience and make a defensible recommendation under uncertainty.A concise recommendation and supporting evidence.How trade-offs were evaluated and whether the recommendation changes when an assumption changes.
Scroll horizontally on a small screen when necessary.

The BMLabs Evidence-of-Learning Chain

Six linked decisions turn a vague AI rule into a defensible assessment design.

Build the chain in order

  1. 1

    State the learning claim

    Name the knowledge, reasoning, performance, judgement, or transfer capability being assessed. Specify the conditions and standard.

  2. 2

    Set the AI boundary

    Decide which cognitive work AI may support, which work must remain attributable to the learner, and whether critical AI use is itself an outcome.

  3. 3

    Choose complementary evidence

    Select an appropriate combination of product, process, performance, and practice evidence. Every item should answer a specific question about the learning claim.

  4. 4

    Add an independent checkpoint

    Where the inference remains weak, add a proportionate oral explanation, live decision, changed-input problem, annotated reasoning step, practical demonstration, or supervised component.

  5. 5

    Test fairness and feasibility

    Examine accommodations, language demands, access, privacy, infrastructure, student workload, staff workload, class size, and disciplinary context.

  6. 6

    Define the judgement and review rule

    State what evidence is sufficient, how contradictory evidence will be interpreted, and what information will be collected to improve the task after use.

Detection-led response versus evidence-led design

TopicDetection-led responseEvidence-led design
Starting questionDid the student use AI?What capability must the student demonstrate, and what evidence would make it visible?
Main mechanismA detector score, stylistic suspicion, or a general prohibition.A stated AI boundary plus complementary evidence tied to the learning claim.
DeclarationTreated as sufficient proof of compliance.Supports transparency but does not replace task design or evidence.
UncertaintyConverted into suspicion about authorship.Addressed through proportionate additional evidence and cautious interpretation.
Shelf lifeDepends heavily on current tool behaviour.Remains useful because it is organised around learning outcomes and inference.

Decide What AI May Do Before Students Begin

Permission should be task-specific, intelligible, and connected to the learning claim.

Decision framework

Choose an AI boundary

Which description best matches the role of AI in this component?

This tool supports planning; it does not create institutional permission. Current university, programme, accessibility, privacy, and disciplinary rules take precedence.

A usable task statement answers six questions before work begins. Which AI functions are allowed? Which are prohibited? What information must not be uploaded? What record or declaration is required? Which parts will be demonstrated independently? How will the work be judged? Avoid instructions such as use AI responsibly without examples. Permission can differ within one assessment: AI may be allowed for generating alternative assumptions, restricted for a changed-input calculation, and deliberately assessed when the student critiques an AI-generated explanation.

Worked Example: Redesigning a Mathematics Modelling Assignment

Keep AI available for appropriate exploration while making assumptions, judgement, calculation, and transfer visible.

Consider a second-year assignment asking students to model water use in a university building and recommend one conservation intervention. The original design requires only a final report containing assumptions, equations, a graph, and a recommendation. An AI system can propose variables, produce code, draft interpretations, and polish the recommendation. The report remains useful, but it provides weak evidence about choosing assumptions, comparing models, checking calculations, interpreting uncertainty, and revising a conclusion when conditions change. The redesign creates a sequence in which important decisions leave observable evidence.
A redesigned evidence sequence
StageStudent activity and AI boundaryEvidence generatedReasonable inference
Problem framingThe student identifies the decision, audience, variables, exclusions, and data limits before using AI.A short problem frame and variable map.The student can define the modelling problem and recognise missing information.
Assumption explorationAI may suggest alternatives. The student records suggestions used, rejects at least one with a reason, and identifies evidence needed.An assumption ledger with rationale, consequence, and confidence.The student can evaluate suggestions rather than merely accept them.
Model constructionAI-assisted code or algebra is permitted, but variables, transformations, units, and test cases are annotated.Annotated model, test cases, and verification record.The student has inspected key operations and connected the representation to the problem.
Model comparisonThe student compares two approaches using fit, interpretability, data demands, sensitivity, and intended use.A comparison table and justified model choice.The student can exercise evaluative judgement under disciplinary constraints.
Final reportAI may assist with organisation and language where policy permits. Sources, calculations, limitations, and material assistance are checked and declared.The complete modelling report and recommendation.The student can integrate evidence into a disciplinary product, subject to the limits of product evidence.
Changed-input checkpointThe student receives a new occupancy assumption and explains which parts of the model and recommendation change.A short recalculation, annotated sketch, or equivalent explanation.The student can transfer selected reasoning to a changed condition and explain a consequential decision.
Reflection and reviewThe student identifies one useful AI suggestion, one rejected or corrected suggestion, and one capability needing practice.A concise evaluative reflection linked to the evidence record.The student can examine the role of AI, while the reflection remains a self-report rather than proof.
Scroll horizontally on a small screen when necessary.

Keep Learning Visible Without Creating Surveillance

Collect the smallest amount of process evidence that materially strengthens the judgement.

Local checklist

Assessment launch checklist

0 of 12 complete

Progress is stored only in this browser.

Questions educators commonly ask

Should every AI-permitted assessment include an oral defence?

No. Add an oral component only when spoken explanation provides relevant evidence and the format is fair. A short transfer task, annotated correction, practical demonstration, or sampled interview may be more appropriate.

Is an AI-use declaration enough?

A declaration supports transparency but does not by itself show that the intended learning occurred. Pair it with task evidence.

Can an AI detector identify which work needs checking?

Detector output should not be treated as proof of authorship or misconduct. Build ordinary assurance into the assessment and follow current institutional procedures when a concern arises.

Must students submit every prompt?

Usually not. Complete logs can be intrusive, burdensome, and weakly connected to the learning claim. Request only the record needed for transparency or evaluation.

How can this work in a very large class?

Use programme-level planning, common rubrics, brief transfer tasks, tutorial sampling, structured process artefacts, and targeted follow-up. Do not add stages the teaching team cannot apply consistently.

What if the product and checkpoint conflict?

Use the pre-announced judgement rule, consider whether both pieces target the same outcome, examine accessibility and context, and offer a proportionate additional demonstration where policy permits.

The compact design standard

  • State the capability before choosing the submission format.
  • Specify the AI boundary at the level of task components and cognitive purposes.
  • Use complementary product, process, performance, or practice evidence.
  • Add an independent checkpoint only when it materially strengthens the inference.
  • Test privacy, access, accessibility, workload, and disciplinary fit before launch.
  • Interpret evidence proportionately and review the design after use.

Conclusion

An assessment does not become trustworthy because AI is banned, declared, or detected. It becomes more defensible when the learning claim is precise, the AI boundary is intelligible, several appropriate forms of evidence are combined, and the judgement remains proportionate to what those forms can show. The goal is not to make student work machine-free. It is to make learning visible enough for a fair academic decision.

Limitations

  • The BMLabs Evidence-of-Learning Chain is an editorial synthesis for planning and reflection; it has not been independently validated as a psychometric instrument.
  • A mathematics modelling example cannot represent assessment needs in laboratories, clinical education, fieldwork, studio practice, law, languages, or every institutional context.
  • Products, process records, oral explanations, supervised tasks, and transfer exercises are imperfect proxies. Combining them can strengthen an inference but cannot establish learning or authorship with certainty.
  • Oral, timed, supervised, and technology-dependent checkpoints may create accessibility, language, privacy, infrastructure, or equity barriers unless alternatives and accommodations are designed.
  • Staged tasks and process review can increase student and staff workload. The design must be proportionate to the learning claim and feasible at class or programme scale.
  • Generative AI capabilities, service terms, privacy arrangements, and institutional rules change rapidly. Educators must verify current local policy.

References and evidence

  1. 1. Assuring Quality Learning in a Gen AI-Integrated Future: The Role of Adaptive Capabilities

    Tertiary Education Quality and Standards Agency · 2026-06-24 · accessed 2026-07-21

  2. 2. Enacting Assessment Reform in a Time of Artificial Intelligence

    Tertiary Education Quality and Standards Agency · 2025-09-24 · accessed 2026-07-21

  3. 3. Identifying What Our Students Have Learned: A Framework for Practical Assessment Validation

    Assessment & Evaluation in Higher Education · 2026-02-01 · accessed 2026-07-21

  4. 4. Talk Is Cheap: Why Structural Assessment Changes Are Needed for a Time of GenAI

    Assessment & Evaluation in Higher Education · 2025-05-15 · accessed 2026-07-21

  5. 5. OECD Digital Education Outlook 2026: Exploring Effective Uses of Generative AI in Education

    OECD Publishing · 2026-01-19 · accessed 2026-07-21

  6. 6. Reimagining the Artificial Intelligence Assessment Scale: A Refined Framework for Educational Assessment

    Journal of University Teaching and Learning Practice · 2025-12-03 · accessed 2026-07-21

  7. 7. Adapting Assessment in the Age of Generative AI: The Assessment Adaptation Model

    Tertiary Education Quality and Standards Agency · accessed 2026-07-21

  8. 8. Guidance for Generative AI in Education and Research

    UNESCO · 2023-09-07 · accessed 2026-07-21

Editorial disclosure

AI assistance was used for source discovery, source comparison, structural drafting, and language editing. The live BMLabs schema, renderer, and audits were checked, authoritative sources were verified, and final claim selection, educational judgement, and publication responsibility remain with Dr. Bivash Majumder.

AI assistance status: research-assistance.

Corrections

  • 2026-07-21 · editorial

    Initial schema-version-2 pre-publication candidate created from the approved editorial plan. No post-publication correction has been made.

Report a factual problem through the BMLabs corrections page.

Share this article