Challenge 03

Parameter Control Experiment

This challenge develops Mechanistic Literacy through controlled parameter or model-condition comparison. Participants demonstrate that they can predict, observe, and explain output variability in accessible mechanistic terms.

Mechanism literacyLevel 1: Instruction with mechanistic controlTokens, tokenisation, attention, sampling, and variability
Difficulty

What is not yet controlled

Similar prompts still produce different results when participants do not understand or control generation conditions.

Approach

What is compared

Hold the task stable, vary parameters systematically, predict effects before execution, isolate variables credibly, and interpret differences in accessible mechanistic terms.

Position in Module#

  • Unit: 3
  • Friday topic: Why Models Behave Differently: Parameters, Variability, and Control
  • Literacy perspective: Mechanism
  • Collaboration level: Instruction, with mechanistic control
  • Fundamentals connection: Generative AI Fundamentals
  • Challenge type: Parameter Control Experiment

This file is the detailed source of truth for the unit challenge. It specifies an evidence-based experiment that participants can transfer to their own professional, disciplinary, or organizational context.

Purpose and Demonstrated Mastery#

This challenge develops Mechanistic Literacy through controlled parameter or model-condition comparison. Participants demonstrate that they can predict, observe, and explain output variability in accessible mechanistic terms.

Clarify in the submitted work:

  • the collaboration move participants practice: Instruction, with mechanistic control
  • the literacy perspective through which they manipulate, observe, and interpret AI behaviour: Mechanism
  • the generative or agentic AI fundamental that explains why the experiment matters: Generative AI Fundamentals
  • the caveat or limitation participants should keep visible: the lowest unit the model directly processes is the token, not the human-visible character. Character-counting, letter-position, reversal, and word-length tasks therefore work against the model's native representation unless the task is decomposed, tool-supported, or covered by memorized textual patterns in training data.

Challenge Focus Logic#

  • Difficulty: Similar prompts still produce different results when participants do not understand or control generation conditions.
  • Approach: Hold the task stable, vary parameters systematically, predict effects before execution, isolate variables credibly, and interpret differences in accessible mechanistic terms.
  • Artefacts: protocol document; prediction table; run log with configurations, output evidence, and observations; selected excerpts or screenshots; concise interpretation; one practical control rule for similar tasks.

Assignment#

Choose a task, use case, document, decision process, or workflow from your own context that fits the unit focus. Define a baseline condition and a target condition that deliberately changes the relevant variable, workflow, responsibility logic, or design choice.

Before running the main comparison, predict what should change and why. Then produce an inspectable artefact, compare outcomes, and interpret whether the target condition improved control compared with the baseline.

Experimental Design#

Define before main execution:

  • task/use case and intended output
  • baseline configuration or approach
  • target configuration, workflow, or intervention
  • manipulated variable(s) or design change(s)
  • what remains constant across comparison
  • expected effects per manipulation/change
  • evaluation criteria or conditions of satisfaction
  • caveat to watch for: the lowest unit the model directly processes is the token, not the human-visible character. Character-counting, letter-position, reversal, and word-length tasks therefore work against the model's native representation unless the task is decomposed, tool-supported, or covered by memorized textual patterns in training data.

The manipulated variable(s) and evaluation criteria should follow from the relevant literacy lens:

  • Interaction: request or exchange design
  • Mechanism: settings, model behaviour, variability, or representation limits
  • Data: grounding, source, retrieval, chunking, or bias conditions
  • System: workflow, role, tool, state, or control-point design
  • Human-Context: trust, responsibility, legitimacy, decision logic, acceptance, or accountability

Required Artefacts#

Required artefacts for this challenge:

  • protocol document
  • prediction table linking changes to expected effects
  • primary artefact or output produced by the challenge
  • run log comparing baseline versus target conditions
  • selected evidence excerpts, screenshots, traces, retrieval results, judge reports, stakeholder feedback, capability maps, or workflow diagrams as appropriate for the unit
  • concise interpretation plus one practical rule for similar tasks

Unit-specific artefact focus:

protocol document; prediction table; run log with configurations, output evidence, and observations; selected excerpts or screenshots; concise interpretation; one practical control rule for similar tasks.

Evidence and Observation#

Participants should capture evidence that makes the comparison inspectable.

Define:

  • what should be observed: variability and control
  • what counts as evidence for the selected use case
  • which qualitative or quantitative criteria should be used
  • how uncertainty, failures, caveats, or unexpected results should be documented

Interpretation#

Participants should explain:

  • whether the target condition improved control compared with the baseline
  • what evidence supports that conclusion
  • how the literacy perspective explains the observed difference
  • how the relevant fundamental or caveat appeared in practice
  • what practical rule they would carry into future work

Assessment Orientation#

Assess based on:

  • meaningfulness of the baseline versus target comparison
  • quality and coherence of protocol design
  • credibility of variable/control isolation
  • suitability of artefacts and evidence
  • evidence-grounded interpretation and conceptual understanding
  • explicit connection to the relevant literacy, collaboration level, and fundamental
  • actionable final insight

Teacher-Configurable Options#

  • exact submission format
  • work mode: individual or group
  • minimum scope: runs, materials, artefact depth, or workflow implementation
  • environment constraints: shared tool or alternatives
  • whether a workflow must be implemented or may be tightly specified
  • optional extensions