Conversation is not the sole source of truth

A model can interpret a question, explain an idea and select a suitable tool. Its generated description of a learner’s performance is not, by itself, a sufficient foundation for a grade or mastery estimate. Wording changes with context and model behaviour; educational judgement needs known inputs and an inspectable rule.

Warda therefore separates explanatory language from the decision path. For a multiple-choice response, the application loads the question and checks the attempt through a defined grading service. Language can then explain the result. A persuasive explanation does not acquire authority to replace that judgement.

One service, two entry points

The educational operation belongs to an application service, exposed to both the product interface and an agent tool. Whether a learner requests a result through the interface or the tutor invokes it during dialogue, the operation reaches the same implementation. A different entry point does not create another copy of the rules.

This makes the rule inspectable independently of model phrasing and reusable in another interaction without rebuilding it inside instructions. The interaction can change while ownership of the decision remains clear. Agent tools coordinate access and present results; they are not an independent grading engine.

An attempt is a meaningful event

Keeping only the most recent result makes learning history difficult to interpret. A new attempt must be distinguished from a repeated request; an independent response must retain the relevant context of assistance. Structured events describe the interaction rather than relying on later interpretation of unrestricted conversation.

This does not mean putting every sentence into each event. Events have defined types and fields, separate from conversation history. That separation supports interpretation and reconstruction; it does not establish a general statement about all platform storage. Structured evidence and dialogue serve different purposes.

What happens when a request repeats?

Request identity and event ordering are part of meaning, not cosmetic optimisations. The event and the message needed for subsequent processing belong to one transactional save boundary. Downstream handling must tolerate repeated delivery. Combining that boundary with an Outbox addresses the gap between saving a change and announcing it.

The engineering cost includes ordering, concurrency and event versions. It is justified where we need to explain an attempt and reconstruct a view from recorded evidence. It does not imply that every platform feature should use event sourcing.

The contract from interface to history

The operation begins with a request belonging to the learner’s session. The service checks session and question ownership before passing defined inputs to the grading rule. It does not accept a ready-made result because a model called it correct, or assume that a client-selected question is answerable by that learner. Persistence then distinguishes a new request from a repeat.

The contract can be examined at several levels: grading inputs, application ownership rules and persistence under retries or conflicts. Passing an isolated grading check does not establish correctness of the entire path. Correct judgement, correct attribution and correct recording are separate obligations; failure in any one can misrepresent learning evidence.

What makes the boundary testable?

Useful checks ask whether the judgement remains stable when explanatory language changes, whether a repeated request is counted once, and whether later readers can interpret an attempt without model guesswork. These are checks on contracts between components, not a competition in answer fluency.

For Warda, this design gives AI room to support dialogue while assigning measurement and history independent responsibilities. It explains our approach to reviewable behaviour. Architecture alone is not evidence of better student learning or superiority over another product.

Sources and context

  1. Warda architecture record — judgement and learning events

    Internal source: architecture principles and grading/event paths, presented here without operational data.

  2. Microsoft Learn — Transactional Outbox

    Pattern reference; its Cosmos DB example does not describe Warda’s components.

Internal records document Warda’s decisions and engineering contracts. External references, where included, explain a pattern or specification; design quality and effects on student learning are different questions.