Are we measuring a model or a system?
A model name does not describe the complete experience. Role instructions, tools, content, session context and interpretation of service outcomes all intervene. Performance on a problem set cannot alone judge the path a learner takes through the application.
We need tasks representing platform responsibilities: contextual explanation, appropriate tool use, exam boundaries and unclear inputs. A single task result also differs from consistency across a conversation that reveals more information. That is why our research examines the system rather than only ranking model names.
Rules need explicit checks
Multiple-choice marking and attempt persistence have defined expectations. Warda’s engineering approach connects these responsibilities to engine, rule and service tests alongside agent behaviour evaluation. A rule check concerns the result and state rather than the accompanying response’s fluency.
A quality claim needs evidence tied to a version, environment and task. Checking a result differs from examining session boundaries or attempt persistence. Defining what each check means makes evaluation useful for repair instead of a general score that cannot explain what succeeded.
Explanations need criteria and review
A correct explanation can be too long to use or assume knowledge the learner does not have. Our quality definition combines criteria for accuracy, curriculum fit and a clear next action, while distinguishing when direct help is useful from when space for an attempt matters.
Anthropic discusses evaluating multiple aspects of an agent’s trajectory and final state. We draw on the distinction between what an agent says and what the system accomplishes. Educational use additionally needs appropriate learning criteria and specialist review. Engineering reference.
A comparison preserving measurement limits
Success on the first does not establish the second, and liking an explanation does not establish the third. We therefore do not promise improved grades or skill mastery from a technical test. Citing research about another product also cannot substitute for a study of Warda.
Evaluation belongs in the product cycle
Warda’s evaluation approach connects system responsibilities, response quality and the educational purpose of an activity. Reviewing a marking result differs from reviewing an explanation, and both differ from examining independent application. Each question needs a suitable task and evidence proportionate to the claim.
The methodology identifies where a path failed and whether content, tools, explanations or interaction design need change. Review then returns to the task to examine the change’s effect. Product improvement gains an examinable reason, while educational quality remains broader than model performance on a problem set.
Sources and context
- Anthropic — Demystifying evals for AI agents
An engineering reference on agent evaluation; it does not measure Warda’s educational impact.
Sources describe research, specifications or documented product behaviour, as identified above. They did not evaluate Warda or establish its effectiveness.