Separate the channel from the decision
Voice has specific responsibilities: capturing speech, turn timing, interruption and playback. Curriculum search, answer verification and skill evidence are educational responsibilities. Separating them means that a voice provider change need not redefine learning rules.
Our architecture vision uses educational services reached through defined tools. The agent coordinates the interaction, each service owns its responsibility, and deterministic engines produce the calculations or judgements assigned to them. A convincing voice gives no authority to choose a grade or alter a criterion.
Shared services, different channel context
A lesson question should reach the appropriate curriculum identity and source through either channel. An objective assessment task uses the same marking service. The explanation therefore remains connected to its evidence regardless of presentation.
Shared services do not make a voice protocol a copy of text chat. Playback and what the learner heard may differ from generated text. Our design distinguishes common educational responsibility from transport and timing requirements.
An interruption changes what was heard
A learner can interrupt a generated explanation before hearing its end. The conversation should not proceed as though the entire response was delivered. Voice design needs to align context with actual delivery, cancel remaining playback and interpret the new question accordingly.
A disconnected session also does not establish an answer, and an absent response does not establish a lack of understanding. Event type and assistance context matter before interpreting an interaction as educational evidence.
The server owns tool responsibility
OpenAI documents a server-side control channel alongside audio, with application tools, authorization and business rules executed on the server. This illustrates the separation between conversation transport and educational responsibility without prescribing a provider for Warda. Official documentation.
A tool call or result needs an appropriate task context and permission. Curriculum and marking service results remain the explanation’s reference rather than material for the model to reinvent. This extends our services behind the conversation approach.
One attempt should not become two pieces of evidence
If a learner explains aloud and completes an answer in writing, those can be parts of the same attempt. Educational events need task identity and assistance context so that changing channels or reconnecting does not duplicate an evidence update. This belongs to service responsibility rather than a language agent’s judgement.
A question’s grade and a learning event’s meaning also remain distinct from conversation text. That separation allows evidence to be interpreted through its rules without treating every utterance as a measured answer. The aim is consistent movement between voice and text, with judgement grounded in the right attempt rather than message count.
Evaluate the whole interaction
Evaluation starts with a spoken question, passes through interpretation, tool selection and displayed evidence, and ends with what the learner heard and can do. Correct text can have unclear playback; a good transcript can lead to the wrong source. Testing one layer cannot establish quality across the interaction.
Mathematics adds consistency between the visible expression, its speech and the meaning being verified. All subjects also require attention to waiting, correction and returning to text. Read role-specific model evaluation for how these responsibilities inform comparisons.
Sources and context
- OpenAI — Server-side controls
Technical documentation for server control and tool execution; not evidence of learning outcomes.
Sources describe research, specifications or documented product behaviour, as identified above. They did not evaluate Warda or establish its effectiveness.