Challenge 06

Reflective Improvement Workflow

This challenge develops reflective improvement using explicit judging and revision criteria. Participants demonstrate that they can separate construction, judging, and revision and evaluate whether the loop improves the artefact.

Interaction literacyMechanism literacyLevel 4: ReflectionLLM-as-judge, conditions of satisfaction, and separate judging sessions
Difficulty

What is not yet controlled

Direct generation leaves weaknesses, blind spots, and unsupported claims unexamined.

Approach

What is compared

Define explicit critique criteria, build a critique-and-revision loop, compare it with a direct-output baseline, and show whether iterative reflection improves quality, consistency, or reasoning.

Position in Module#

  • Unit: 6
  • Friday topic: Reflection Loops: Critique, Revision, and Better Reasoning
  • Literacy perspective: Interaction, secondarily Mechanism
  • Collaboration level: Reflection
  • Fundamentals connection: Generative AI Fundamentals
  • Challenge type: Reflective Improvement Workflow

This file is the detailed source of truth for the unit challenge. It specifies an evidence-based experiment that participants can transfer to their own professional, disciplinary, or organizational context.

Purpose and Demonstrated Mastery#

This challenge develops reflective improvement using explicit judging and revision criteria. Participants demonstrate that they can separate construction, judging, and revision and evaluate whether the loop improves the artefact.

Clarify in the submitted work:

  • the collaboration move participants practice: Reflection
  • the literacy perspective through which they manipulate, observe, and interpret AI behaviour: Interaction, secondarily Mechanism
  • the generative or agentic AI fundamental that explains why the experiment matters: Generative AI Fundamentals
  • the caveat or limitation participants should keep visible: an LLM judge can be biased by wording, order, verbosity, style, or knowledge of which output it produced. Running the same judgement with reordered candidates, anonymized outputs, or a separate judging session can reveal instability in the evaluation.

Challenge Focus Logic#

  • Difficulty: Direct generation leaves weaknesses, blind spots, and unsupported claims unexamined.
  • Approach: Define explicit critique criteria, build a critique-and-revision loop, compare it with a direct-output baseline, and show whether iterative reflection improves quality, consistency, or reasoning.
  • Artefacts: protocol document; prediction table; primary artefact or output; run log comparing baseline and reflective conditions; selected evidence excerpts or screenshots; concise interpretation; one practical reflection rule for similar tasks.

Assignment#

Choose a task, use case, document, decision process, or workflow from your own context that fits the unit focus. Define a baseline condition and a target condition that deliberately changes the relevant variable, workflow, responsibility logic, or design choice.

Before running the main comparison, predict what should change and why. Then produce an inspectable artefact, compare outcomes, and interpret whether the target condition improved control compared with the baseline.

Experimental Design#

Define before main execution:

  • task/use case and intended output
  • baseline configuration or approach
  • target configuration, workflow, or intervention
  • manipulated variable(s) or design change(s)
  • what remains constant across comparison
  • expected effects per manipulation/change
  • evaluation criteria or conditions of satisfaction
  • caveat to watch for: an LLM judge can be biased by wording, order, verbosity, style, or knowledge of which output it produced. Running the same judgement with reordered candidates, anonymized outputs, or a separate judging session can reveal instability in the evaluation.

The manipulated variable(s) and evaluation criteria should follow from the relevant literacy lens:

  • Interaction: request or exchange design
  • Mechanism: settings, model behaviour, variability, or representation limits
  • Data: grounding, source, retrieval, chunking, or bias conditions
  • System: workflow, role, tool, state, or control-point design
  • Human-Context: trust, responsibility, legitimacy, decision logic, acceptance, or accountability

Required Artefacts#

Required artefacts for this challenge:

  • protocol document
  • prediction table linking changes to expected effects
  • primary artefact or output produced by the challenge
  • run log comparing baseline versus target conditions
  • selected evidence excerpts, screenshots, traces, retrieval results, judge reports, stakeholder feedback, capability maps, or workflow diagrams as appropriate for the unit
  • concise interpretation plus one practical rule for similar tasks

Unit-specific artefact focus:

protocol document; prediction table; primary artefact or output; run log comparing baseline and reflective conditions; selected evidence excerpts or screenshots; concise interpretation; one practical reflection rule for similar tasks.

Evidence and Observation#

Participants should capture evidence that makes the comparison inspectable.

Define:

  • what should be observed: reasoning improvement and consistency
  • what counts as evidence for the selected use case
  • which qualitative or quantitative criteria should be used
  • how uncertainty, failures, caveats, or unexpected results should be documented

Interpretation#

Participants should explain:

  • whether the target condition improved control compared with the baseline
  • what evidence supports that conclusion
  • how the literacy perspective explains the observed difference
  • how the relevant fundamental or caveat appeared in practice
  • what practical rule they would carry into future work

Assessment Orientation#

Assess based on:

  • meaningfulness of the baseline versus target comparison
  • quality and coherence of protocol design
  • credibility of variable/control isolation
  • suitability of artefacts and evidence
  • evidence-grounded interpretation and conceptual understanding
  • explicit connection to the relevant literacy, collaboration level, and fundamental
  • actionable final insight

Teacher-Configurable Options#

  • exact submission format
  • work mode: individual or group
  • minimum scope: runs, materials, artefact depth, or workflow implementation
  • environment constraints: shared tool or alternatives
  • whether a workflow must be implemented or may be tightly specified
  • optional extensions