Before checking a solution, identify the task

Verification needs a defined problem: its givens, requested result, variables and assumptions. Is the task asking for an answer or a particular form? If an image is misread, a checker can accurately verify a problem the learner never supplied.

Our experience design separates interpreting input from judging it. When notation or intent is ambiguous, confirming what was understood belongs in the experience rather than disappearing inside a lengthy explanation. Accuracy begins with identifying what was understood before selecting a check or phrasing assistance.

Matching, equivalence and task completion

SymPy documents that structural equality differs from symbolic equivalence. STACK explains that equivalence checks on a list do not necessarily assess step size or a sensible order, and that the final requested form needs a separate check. SymPy · STACK.

For Warda these become separate questions. Is the expression understood? Is a transformation acceptable under the task’s assumptions? Has the learner reached the requested result? Equivalence alone does not establish completion, while a different written form is not sufficient reason to reject a valid answer.

Domain restrictions are part of the meaning

Cancelling a denominator may affect the expression’s domain. An inverse operation may need a condition to preserve the intended solutions. We therefore do not want a checker comparing appearance alone, or an assistant overlooking assumptions because the next line looks familiar. Task meaning must travel through verification.

Unable to decide is a legitimate result

A checker may not support a form, or the input’s meaning may remain unresolved. “Cannot establish this yet” must be a possible outcome. It is not an incorrect learner answer. The next action might be clarification, confirmation of recognised input or appropriate review.

We do not want failed verification to turn silently into approval from a fluent model. The model can suggest an interpretation or explanation, but a suggestion does not become a verified judgement simply because it follows a tool call. Our design preserves this distinction in both learner messages and later learning evidence.

Judgement rules in Warda’s design

Judgement begins by identifying the task type. Matching a multiple-choice response to its key differs from checking an open mathematical transformation or interpreting why the learner chose it. Each responsibility has its own criteria; a fluent explanation does not become verified judgement through confidence alone.

Warda’s approach connects judgement to what was actually checked: input meaning, assumptions, transformation validity and the requested result. Insufficient evidence calls for clarification or appropriate review. Educational explanation then connects the outcome with an understandable reason and another attempt. Mathematical validity and educational usefulness are connected responsibilities with distinct questions.

Sources and context

  1. SymPy — Gotchas and pitfalls

    Documentation distinguishing structural equality from mathematical equivalence; not an announcement of Warda’s engine choice.

  2. STACK — Equivalence reasoning assessment

    Documentation on equivalence assessment and its limits, not an evaluation of Warda or a decision to adopt STACK.

Sources describe research, specifications or documented product behaviour, as identified above. They did not evaluate Warda or establish its effectiveness.