Policies & agreements

Responsible AI Policy

What the model is allowed to do, what it is never asked to do, exactly what leaves our systems and exactly what comes back — whether it is equally useful to every child in a class, and the figures we will not publish until we can show the working. Written for a principal, a privacy officer and a teacher.

Version 1.1 Issued 26 August 2026 Next review 26 August 2027 AI Accountability Owner Founder and CTO
See exactly what crosses

Free to read. No account, no login.

The short version

A summary, not a substitute — the sections below are the policy.

The AI never decides a mark, and the teacher owns every judgement

Correct or incorrect is decided by our own code against an answer key. The model is told the verdict; it is never asked for one (Terms clause 4.2).

The AI never talks to a child

No chatbot, no tutor, no hint — and no messaging of any kind in the student app. A child's route to an adult is their teacher, in the room.

No identifier crosses

One answer per request, with a year-level band rather than a year number. Not a name, not our student number, and not a pseudonymous token.

A page the model cannot read costs the child nothing

“Unreadable” is a first-class answer, never treated as wrong and never guessed at. It is the fairness property the rest of section 4 rests on.

No training, and versions pinned

The provider is contractually barred from training on school data, zero retention is required of them, and we never point at “latest”. A model change carries 14 days' notice.

We publish no accuracy or fairness figure

Not until there is a stated evaluation set, method, date and model version behind it. Any figure quoted to you without all four is not ours (Terms clause 4.5).

Working through a departmental AI checklist? Section 4 is fairness and non-discrimination, sections 5 to 8 are the boundary in full, and section 15 is the list of things this service is not — a diagnosis, a prediction, or a comparison between children.

On this page

Every commitment on this page is one we also make in the Privacy Policy or the Terms of Service. Where a commitment is contractual, the clause number is cited so a school can hold us to it.

Where we cannot claim something, we say so instead of leaving it out.

What this policy covers, and what governs

Reason Education is a maths assessment tool for primary schools. A child answers questions and does their working out on a canvas in the student app. Our own code marks the answer. AI Analysis is the one feature in the service that calls a language model: it describes the method a child used and where that method broke down, for a teacher to read alongside the child's actual work.

This policy covers that feature, and only that feature, as operated by entity name — to be confirmed (ACN acn — to be confirmed, ABN abn — to be confirmed), trading as Reason Education. There is no other AI in the product.

What actually binds us

This page describes how AI Analysis behaves. The contractual commitments sit in the Terms of Service — what the service does and does not do (clause 4), AI Analysis (clause 13), consent (clause 14), changes to the model (clause 15), sub-processors (clause 20) and accuracy (clause 26.3) — and in Privacy Policy section 7. If a statement on this page and a statement in those documents ever differ, those documents govern and this page is the thing that is wrong.

Keeping this accurate

This policy is version-controlled and dated. It is reviewed:

  • at least annually, on the next review date shown at the top of this page;
  • whenever the model version, the prompt, the output contract, the payload, or the country the model is served from changes; and
  • whenever our quarterly scan of AI safety guidance finds something that changes a control — see section 16.

A review that finds no change still gets a dated entry, because “we looked and it was still true” is the part an assessor cannot otherwise verify.

  • v1.1 First issue.

What the AI does, and what it does not do

The whole of it, in two lists.

What AI Analysis does

  • Describes the method a child appears to have used, from the image of their working out.
  • Says where that method broke down, where it can see a breakdown.
  • Applies a tag from a frozen list a human wrote — 23 error patterns. The model chooses from the list; it does not invent a category, and no label in the list describes a person rather than a page.
  • Returns how well it could read the page — high, medium, or unreadable. Not a percentage.
  • Works on one answer at a time, in a single request with no memory of the last one.

What AI Analysis does not do

  • It does not decide a mark. Marking is arithmetic against an answer key — see section 3.
  • It does not talk to a child. There is no chatbot, tutor or hint of any kind, and there is no messaging of any kind in the student app (Terms clauses 4.3 and 4.4).
  • It does not compare one child to another. There is no benchmarking, no ranking and no cross-school analytics — see section 4.
  • It does not predict future performance.
  • It makes no claim about any condition, disorder or disability, and none about a child's wellbeing, effort, attitude or behaviour.
  • It does not make or recommend a decision about a child — placement, grouping, referral or intervention.

These are commitments, not gaps

The six items above are not features we have not got to yet. They are the shape of the product, they are written into Terms clause 4.3, and removing one of them would be a change to the agreement rather than a release note.

The design decision the rest of this rests on

A child in this product cannot reach another person. No chat, no messaging, no comments, no forum, no display name, no avatar a child picks, no shared space of any kind. A child cannot see that another child exists, cannot search for one and cannot be found by one. Most online safety harm to children is one person reaching another, and we removed the mechanism rather than moderating it.

“The AI never talks to a child” is a consequence of that decision, not a separate promise. We hold the product against the eSafety Commissioner's Safety by Design principles — service provider responsibility, user empowerment and autonomy, transparency and accountability — and we state where a principle is not met yet. That is a description of our product against a set of principles, and we do not describe it as certification or conformance to a standard (see section 14).

The human decision: marking is arithmetic, the teacher decides

Terms clause 4.2 — marking is arithmetic

Correct or incorrect is decided by code against an answer key. The AI never decides a mark. Note the wording: never decides a mark, not never touches one — it is shown the mark as an input. AI Analysis describes the method a child used and where it broke down, for a teacher to read alongside the child's actual work. The teacher decides. Every curriculum judgement recorded in the service is a teacher's judgement, not the model's.

Four different questions get asked of a child's work, and four different things answer them.

Who answers what, in a service that uses a model.
The questionWho answers it
Is this answer correct?Our own code, against an answer key. The model is told the verdict as part of the request; it is never asked for one.
What method did the child use, and where did it break down?The model. This is a description for a teacher to weigh, not a finding. It can be wrong.
What does this mean for this child's learning?The teacher. Every curriculum judgement in the service is recorded as a teacher's judgement, against Victorian Curriculum 2.0, and is decided on correctness alone.
What happens next for this child?The school. AI output must never be the sole basis for a decision about a child (Terms clause 26.3).

This is the ordering that matters most, and it is the one a product like ours is most often tempted to blur. A model that is never asked for a verdict cannot give a wrong one, and a teacher who is shown a description rather than a score still has something to disagree with.

The oversight is structural rather than procedural. Nothing model-derived reaches a report, a parent conversation or a child without passing a teacher release gate, and that gate sits inside the query rather than in code that could forget to apply it. An AI reading can never render as a positive: there is no “all good” to click through, because silence from a model nobody has read is not evidence. An unread AI flag counts in neither direction and is reported as still outstanding rather than quietly omitted.

Fairness: whether the AI is equally useful to every child

A model that reads some children's pages less well than others is not a neutral tool. It is one that gives some children a teacher's full attention and others a shrug. That is the fairness question in this product, and it is narrower and more concrete than the word usually is: whether the AI's usefulness is spread evenly across the children in a class.

The children most likely to be read poorly are the ones with the least slack — a child whose letterforms are unconventional or whose pressure is uneven, a child working in more than one language or forming digits to a different convention, and any child whose page is faint, crowded, over-traced or restarted, which is most children on their worst day.

What the design already does about it

  • “Unreadable” is a first-class answer. It is never treated as wrong, never inferred, never guessed at, and it influences nothing. A page the model cannot read costs the child nothing at all. This is the single most important fairness property in the product.
  • The mark does not depend on the model. Correct or incorrect is decided by our own code against an answer key. A child whose handwriting reads poorly is marked exactly as a child whose handwriting reads well.
  • Nothing accumulates. “Unreadable” is a state of one page, not a fact about a child. It never totals, never trends, and never appears in a growth history.
  • No comparison of one child to another, or to a norm. No ranking, no percentile, no “below average”, no “compared to peers” — and the vocabulary guard in section 6 blocks comparison language outright, so the wording cannot drift back in.
  • No AI-sourced statement can move a curriculum standard or place a child in a group. A curriculum judgement is decided on correctness alone.

A trade we made on purpose, and state rather than hide

The scan that stops a page being sent when it carries a classmate's name works from the roster of the actual class sitting the paper, and it exempts a short list of names that are also ordinary mathematical wordssum, mark, rose, may, bill and others.

That makes the guard weaker for the child actually called Rose, and better for the other twenty-nine in the room, because a guard that fires on the word “sum” would be switched off inside a week and take everybody's protection with it. It is unequal by design. It is written into our source in those words, so that whoever changes the list next understands what they are changing, and the list is reviewed annually with this policy.

The thing we cannot measure, and why we are not fixing it by collecting more

Bias of this kind is usually measured by slicing results across groups. We hold no attribute to slice by. No gender, no language background, no disability status, no demographic field of any kind — those key names are refused outright at the boundary, and the product never collected them in the first place.

We are not going to start collecting them in order to measure fairness. Holding sensitive attributes about children so that we can audit ourselves is a worse trade than the one it fixes, and we would rather say so than quietly widen what we hold.

What we do instead:

  • We compare the unreadable rate across the cohorts we can see — by class, by run and by school — and report it monthly. It is a coarse proxy, and it is a real one, because classes differ in the ways that matter.
  • Where a school wants a finer cut, the school holds those attributes and we do not. We give the school its own per-run figures so that it can run that comparison itself, under its own obligations.
  • A material difference between classes is a fairness finding, not a hard class. It is a signal about capture, handwriting or the model, and it is investigated as one — under section 13, where it is a named category of AI incident.
  • The comparison runs from launch. It is a condition of the first production model call, not an improvement scheduled for afterwards.

What we have not done, and will not claim

We have measured nothing, and we publish no fairness claim and no bias figure. We will not until there is a stated evaluation set, method, date and model version — the same four conditions that govern any accuracy figure (section 14).

The evaluation set any future figure would rest on is synthetic by design, and cleaner than real children's pages — real handwriting is fainter, more crowded and more often ambiguous. A model that scores well on it has not been shown to score well on a Year 2 desk. That limit travels with any figure we ever publish, and it is why the pages are authored by several different hands and deliberately made hard. Once a school is using the product, a second, separately scored set of consented real pages is the check on how optimistic the synthetic one is.

What this section is not

Everything above describes handwriting, capture and legibility. It does not describe children. Nothing in it identifies, screens for or infers any condition, difficulty or disability, and no field exists in which such a claim could be recorded.

What crosses to the model, and what never does

Stated up front

The AI is designed to receive and process personal information. A photograph of a child's handwritten working out is personal information at the model, and we say so rather than answering otherwise. Everything below describes how we bounded that, not how we avoid it (Terms clause 13.1).

The payload is assembled from an allow-list — built up from named fields, never produced by stripping identifiers out of a full record. The difference matters: a field added to our database next year is not sent by default, because nothing is sent unless it was named. One answer is sent per request, not one paper, so a model cannot compare question 3 with question 9 and start describing a person rather than a page.

What crosses to the model, what is refused, and what comes back One answer per request crosses: the pseudonymised image of the working out, the question, the answer key and the child's answer, whether our own code marked the answer correct, the curriculum unit, standard code and the strategy taught, and a year-level band. Nothing else crosses — not a name, not our student number, not a database id, not a pseudonymous token, and never a year number. The call is made to Anthropic in the United States, with training barred, zero retention required and the version pinned. Exactly four fields come back: the method the child used, where it broke down, a tag from a frozen human-written list of twenty-three, and how well the model could read the page — high, medium or unreadable. Anything else is discarded whole. WHAT CROSSES — ONE ANSWER PER REQUEST the image of the working out, pseudonymised the question, the answer key, the child's answer whether our own code marked it correct the unit, the standard code, the strategy taught a year-level band — “middle primary” NEVER CROSSES a name, or our student number a database id, or a pseudonymous token a year number — never “Year 3” THE MODEL — ANTHROPIC, UNITED STATES no training · zero retention · version pinned EXACTLY FOUR FIELDS COME BACK 1 · the method the child used 2 · where it broke down 3 · a tag from a frozen list of 23 4 · high, medium or unreadable

Anything the model returns beyond those four fields is discarded whole, and the response is scanned before it is stored — see section 6 and section 7.

The same boundary, as a list

The allow-list, in both directions.
What crosses, every requestWhat never crosses
The image of the handwritten working outA first name or a last name
The question, the answer key, and the answer the child gaveThe student number we generate
Whether our own code marked the answer correct or incorrectAny database id
The curriculum unit, the standard code, and the strategy the class was taughtA pseudonymous token of any kind
A year-level band — “middle primary”A year number — never “Year 3”
Nothing else. One answer per requestAnything not on the allow-list, including any field added to our database later

Why not even a pseudonymous token

A pseudonymous token is still an identifier. It would make every request for one child linkable to every other, at the provider, without ever carrying a name. So there is no token, no session id and no thread. An earlier design carried one and it was removed. Each request stands alone, and the correlation happens in a variable inside the calling function on our own server, which never leaves it (Terms clause 13.2).

The word we use is “pseudonymised”

That is the word, and we use no stronger one. We know what we put into the payload. We cannot know what a child drew on the page (Terms clause 13.3). “Anonymised” would be a claim about the image, and it is not one we are in a position to make.

What comes back, and what is discarded

The output contract is fixed. The model returns four fields and no more:

  1. A description of the method the child appears to have used.
  2. Where it broke down — the point in that method where the work goes wrong.
  3. A tag from a frozen list a human wrote — 23 error patterns. The list is ours, it is finite, and the model selects from it. It cannot add a category, and a response naming one that is not on the list fails the contract.
  4. How well it could read the pagehigh, medium or unreadable. Three values, not a percentage.

Anything else is discarded whole

Not trimmed, not summarised, not repaired, not stored for review — discarded. If the model returns advice, a mark, a recommendation about a child, or free text beyond the four fields, none of it reaches the teacher and none of it reaches our database. We do not keep the parts we liked: a model that broke the contract in one field has told us something about how far to trust the other three, and a repaired response is a response we wrote and then attributed to a model.

An unreadable response may not also carry a tag. If the model could not read the page it does not get to guess what went wrong on it, and the contract refuses the combination.

The vocabulary guard

A second check reads the words themselves. More than eighty terms are refused across six categories — diagnosis and disability, wellbeing and behaviour, ability and progress, comparison, generalisation over time, and speaking to a child rather than about a page. A hit discards the whole response.

The reason it exists is worth stating plainly: a model that volunteers “this looks like dyscalculia” has, at that moment, generated health information about a child. The guard is what makes “we do not handle health information” a statement rather than a hope. It also scans our own tag list, so no label may describe a person rather than a page.

Everything discarded is counted, not silently dropped. Each quarantined response records the reason, the terms that fired and the model and prompt versions in force, and the rate is reviewed weekly — see section 10. Counting what you refuse is what turns a refusal into a monitoring signal rather than a hole.

The confidence value is the model's report on its own reading. It is not an accuracy figure, it is not evidence that the description is right, and it is shown to a teacher as what it is, with a visible provenance chip rather than a tooltip so that a principal reading over a teacher's shoulder can tell AI from arithmetic without hovering anything. We publish no accuracy figure at all — section 14 says why.

The response is also scanned against the class roster before it is stored, and quarantined if it echoes a name. That is control 4 in section 7.

Names on the page, and the limit of our controls

Children write their names on their work. That is what every child has been taught to do since Foundation, and no instruction will stop it entirely. Four controls apply (Terms clause 13.4):

  1. Design — the canvas has no name field, no header region and no title page. Nothing invites a name.
  2. Instruction — a child-readable hint on the canvas, and a matching note to the teacher in the run instructions, held in one place so the screen and the adult say the same thing.
  3. Pre-send scan — every piece of text in the payload is checked against the roster of the class actually sitting the paper. A hit means the page is not sent: it goes to teacher-only review, and it is logged. This control is blocking, and it runs before the call is made.
  4. Post-call scan — the model's response is checked against the same roster. A hit means the response is quarantined before it is stored, so the name is never written into our database. The name itself is not written into the quarantine record either: storing the personal information you just detected recreates the disclosure inside your own logs.

What we will not claim

We will never say the AI can never see a name. Control 4 fires after the call has been made — it is detection, not prevention. And control 3 reads text: a name written inside the drawing is not detected before the call. That is a stated limit, not an oversight, and it is why control 2 — the school's own instruction — is the strongest of the four. A school told “never” that then receives a quarantine notice has been misled, so we tell you now.

The honest position is the four controls above, the no-training term in section 8, and the fact that a child may still write identifying details on their page — which we detect and quarantine where we can.

The school's part, and it is the strongest one

The school instructs students not to write their names or other identifying details on the working out canvas, in the same way it instructs them about anything else in a test. This is control 2, it is an obligation in Terms clauses 13.5 and 22.3, and it is the cheapest and most effective of the four.

No training, no provider retention, and where the calls go

  • The model provider is contractually barred from training any model on school data, and we hold that term in writing before the first model call is made.
  • Zero retention on the provider's side is required of them — nothing kept beyond what is needed to return the response.
  • We obtain written confirmation of the region the model is served from, on the same terms and before the same first call.
  • If we ever change providers, these are the first terms we check. A provider who will not agree to them is not a provider we can use.
  • We do not use school or student data to train, tune or improve any model of our own either. We do not have one, we hold no weights, and no teacher decision, stored answer or output is ever sent back to a model.

One country, named

Model calls are made to Anthropic in the United States. That is the only place a child's work leaves Australia. A second, separate call to the same provider — made when we set up a school's own workspace — carries the school's name and nothing else: no student data, no staff data, no email addresses. It is named here because it is a real outbound call and it is not covered by the description above.

Everything else — the live application, the database, the stored images, the backups, support and administration — stays in Australia. We do not claim that nothing leaves Australian jurisdiction, because that claim would not be true (Terms clauses 11.2 and 13.6). The full component-by-component picture is in the Security Policy and the sub-processor list in Terms clause 20.1.

A change to the country the model is hosted or served from is a hosting change and carries 90 days' written notice — and in every case the notice goes out before the first call is made from the new location, whichever comes first. See section 9.

What we keep of an AI call, and for how long

The provider keeps nothing. We keep some of it, and these are the periods (Terms clause 17, Privacy Policy section 11).

Retention of AI-related records.
WhatHow long
The image of the working out12 months from the close of the assessment run
The AI call log, name-guard events and quarantine records24 months — a log deleted early is a control removed
The raw text of a quarantined response90 days
Any of it, once a school instructs deletionGone from the live service within 30 days, and expired from backups a further 100 days after that

What the call log does not hold

The call log records the question, the model and prompt versions, how long the call took, the outcome and any guard hits. It never records the image, and it never records an identifier.

Pinned versions, and notice before anything changes

A model that changes underneath a teacher without telling them is the failure mode this section exists to prevent.

  • Model and prompt versions are pinned in source, not in configuration. We never point at “latest”, so a provider updating their model cannot silently change what a teacher reads — and a change is then a diff somebody reviews rather than a setting somebody edits on a Friday (Terms clause 15.2).
  • A change is approved in writing before it ships, by the AI Accountability Owner, who does not write the code being changed (section 16).
  • A regression run comes first. Any model or prompt change is run against our evaluation set and our adversarial corpus before it ships. The run reports the unreadable rate by authoring hand, and a material change in it stops the release the same way a boundary failure does.
  • A notice states the current version, the new version, what changes about the output, and what we tested before shipping it (Terms clause 15.1).
  • A change history is published on this website, so a school can see what changed and when (Terms clause 15.3).
  • A school that does not want a change may turn AI Analysis off under section 11. We will not treat continued use as agreement to something we did not tell you about first (Terms clause 15.4).
Written notice before a change takes effect.
ChangeWritten notice first
A change to the AI model versionAt least 14 days (Terms clause 15.1)
A material change to the prompt or the output contractAt least 14 days (Terms clause 15.1)
A change to the country the model is hosted or served from90 days, and always before the first call from the new location — this is a hosting change, not a model change (Terms clauses 15.5 and 12.2)
A new sub-processor that receives school data30 days minimum, and a school may object and terminate (Terms clause 20.2)

Where a provider gives us less notice than we owe you

If a provider forces a change on us with less notice than we owe a school — a region withdrawn, a service retired — we tell every school within 2 business days of learning of it, say plainly that we could not meet our own lead time and why, and offer the same right to object. We do not have a mechanism to give a school 90 days of notice that we ourselves did not get, and we will not pretend otherwise.

How the boundary is held: tests, gates and monitoring

The boundary in section 5 is not a thing we intend. It is a thing a test asserts. Here is exactly how far that goes, and where it stops.

Eleven suites, run by hand

The payload allow-list, the output contract, the vocabulary guard, the name guard, the release gate and the adversarial corpus are covered by eleven automated test suites. They are run by hand on every change.

There is no continuous integration system, and we do not describe one. The suites pass because someone ran them, not because a pipeline blocked a merge. Wiring them into a blocking pipeline is a committed item, and until it is done we will not say “on every build” or describe a control as enforced by a pipeline that does not exist. A gap you are told about is one you can price; one you find yourself discounts everything else we say.

The adversarial corpus, and the sentence we will not write

A single corpus of hostile cases — instructions a child could write on the page addressed to the model, malicious and malformed model responses, and payload tampering — runs against the guards on every run of the suites, and will be replayed against the real endpoint before the first production model call. It carries declared gaps of its own, written into the corpus file and asserted to be there, because a corpus with no failures is a corpus written to pass.

The distinction we hold

Over-claim: “the system has been adversarially tested.”
True: we maintain an adversarial corpus that runs against every run of the test suites, and it will be replayed against the model endpoint before the first production call.

The strongest defence against an instruction written on a page is not a filter. It is that there is nothing for it to reach: the model has no tools, no retrieval, no memory, no ability to write anywhere, and a four-field reply schema with no field that asks it to act. The worst outcome of a successful attempt is a wrong description of one page, shown to one teacher, labelled as AI, sitting next to the child's actual working out.

Before a new feature ships

No new user-facing feature ships until it is checked against seven questions, and the answers are written into that feature's design note:

  1. Can this let one child reach another, or let a non-school person reach a child?
  2. Does it widen what a child can see, and is the widening scoped to their own record?
  3. Does it add any third-party request to a page a child sees?
  4. Does it create free-form content that nothing inspects?
  5. Does it compare a child to another child, or to a norm?
  6. Does it produce an evaluative statement without a source a reader can see?
  7. Does it collect a field we do not need?

A “yes” to 1, 3, 5 or 6 stops the feature. A “yes” to 4 requires a named compensating control before it ships — the drawn canvas is the thing that fails that question, and section 7 is the consequence.

What is watched once it is running

Monitoring signals, and what a move in one means.
SignalHow oftenWhat a move means
Quarantine — responses discarded by the contract or the vocabulary guard, and pages withheld by the name guard, counted by reasonWeeklyA rising rate is the earliest sign that a model or a prompt has drifted
Name-guard hits, counted separately before the call and after itWeeklyRising hits after the call mean the canvas instruction is not landing — a design problem, not a child's mistake
Unreadable rate across the whole cohortMonthlyA jump across the board is usually a problem with how pages are being captured
Unreadable rate compared across cohorts — class, run and schoolMonthlyA class materially above the others is a fairness finding and is investigated as one, not written off as a hard class (section 4)
Teacher agree-rateMonthlyAn agree-rate near 100% is deference, not accuracy — it is the signal that the design's central bet is failing
Stored AI responses read against the rules in this policyAnnually, sampled and written upReported to all three founders, whatever it finds

Turning AI Analysis off

Terms clause 13.7 — the switch is yours

A school may turn AI Analysis off for the whole school, at any time, in writing. We action the request within 2 business days, without argument and without a commercial conversation.

Everything else keeps working. Assessments, marking, results, curriculum judgements and growth history all run without AI Analysis, because marking is arithmetic. This is not a degraded mode and it is not priced differently; it is the same product with one feature switched off.

That is a deliberate design property, and it is also what makes the halt powers in section 16 usable. An owner who knows that stopping the AI breaks the product will hesitate, and hesitation is the failure mode those powers exist to prevent.

A school that wants AI Analysis stopped for a single student does that through the consent record in section 12, and it takes effect immediately.

Contesting an output, and what counts as an AI incident

AI Analysis is a description of a method, produced by a statistical model, for a teacher to weigh against the child's actual work. It can be wrong. When a school thinks an output is wrong or harmful, this is the path.

Tell usFrom launch, every AI-sourced statement a teacher sees carries a way to say the reading is wrong, unhelpful or inappropriate, attached to that response, with a box to say why. It goes to the AI Accountability Owner. You can also write to support email — to be confirmed, or use the school's usual support contact. A privacy concern about the same output can go to privacy email — to be confirmed.
We acknowledgeThe same day, and triage it the same day for whether it indicates a breach of the rules in this policy.
Who reviews itThe AI Accountability Owner, who is the Founder and CTO and writes no part of the AI (section 16).
What they can doHold a release, block or revert a model or prompt change, or stop model calls entirely. Those powers are the point of the role — an owner who cannot stop the thing is not an owner.
You get a written answerWithin 5 business days (Terms clause 13.8).
How it closesWith a change and a test, or a written record of why no change is warranted. One that produces neither has not been closed.

What counts as an AI incident

Any of these, and the list is not exhaustive:

  • a model output reaching a teacher having breached the rules in this policy;
  • an identifier reaching a model, or a name detected in a model's response;
  • an AI-sourced statement presented as a verdict, or rendered as a positive, before a teacher decided it;
  • a quarantine rate that moves materially;
  • a school reporting an AI output it considers wrong or harmful; or
  • a fairness finding — a pattern in which the model reads one group's pages materially less well than another's, or any credible report from a school that the AI serves some of their children better than others.

Containment first, and who decides

Where the boundary or the no-student-output rule is implicated, model calls are stopped before anything is investigated. The affected school is told what happened, in plain language, without waiting for the investigation to finish.

A fairness finding follows the same path with one addition: the AI Accountability Owner decides whether AI Analysis stays on for that school while it is investigated, and records the reason either way. That decision belongs to the accountability role, not to the person who wrote the code.

Terms clause 26.3 — accuracy of AI output

AI Analysis is not a mark, it is not a diagnosis, and it must never be relied on as the sole basis for a decision about a child. The teacher's judgement governs.

A complaint about privacy — rather than about an output — goes to the Privacy Officer under Privacy Policy section 16, is answered within 30 days, and always carries the right to go to the Office of the Australian Information Commissioner instead. Complaint volumes and themes go into our six-monthly privacy report and into the annual review of this policy.

What we will not publish, and why

Terms clause 4.5 — we publish no accuracy figure

We will not publish one until there is a stated evaluation set, method, date and model version behind it. Any figure quoted to you that does not carry all four is not ours.

This is the section a vendor usually fills with a percentage. We would rather be the company that has none than the company whose number nobody can reproduce. Each item below carries the condition that would change it, rather than a flat refusal.

  • No accuracy or precision figure for AI Analysis, in any form — not on this site, not in a tender response, not in a demonstration. What would change it: a named evaluation set, a described method, a date, a pinned model version, and the name of whoever ran it — published together, with the limits of a synthetic set stated beside the number.
  • No benchmark result, ours or anyone else's, presented as evidence about this product. What would change it: the benchmark named, the version we ran, the date, and who ran it.
  • No bias-audit or fairness figure. We have measured nothing, and calling our cohort comparison an audit would be the same misrepresentation with a nicer word. What exists is the design in section 4 and one monitoring signal — the unreadable rate compared across class, run and school, reported monthly. What would change it: an evaluation set with pages authored by at least three different hands, an unreadable rate reported by authoring hand, a date and a pinned model version — and even then, published with the synthetic-set limit next to it.
  • No claim of conformance to a named AI standard or framework. We map our controls against published guidance — the OWASP machine learning and AI security guides, the eSafety Safety by Design principles — and we name what those maps leave open. What would change it: a completed assessment against the named standard, by a named assessor, with a date. Where our behaviour happens to answer what a framework asks about, that is a description of the product, not a certification.
  • No status claim under Safer Technologies 4 Schools, of any kind, until a full assessment completes. Not “in progress”, not “submitted”, not “aligned”. What would change it: the completed assessment.
  • No claim of endorsement, affiliation or accreditation by any department of education, curriculum authority or assessment scheme, unless we say so in writing and name the scheme (Terms clause 29.2).
  • No description of a control as operating when it is specified but not built (Terms clause 31.2). Where something is not yet available, we tell the school in writing before it signs, say when it lands, and say how the same commitment is met in the meantime.
  • No “the AI can never see a name” (section 7), no “nothing leaves Australian jurisdiction” (section 8), and no “the system has been adversarially tested” (section 10). Three sentences that would each be easy to write and none of which is true.

What this is not

The Reason Education service is decision support for a teacher. It is not any of the following, and we will not describe it as any of them (Terms clauses 4.3 and 4.6).

  • Not a diagnostic tool. It makes no claim about any condition, disorder or disability, and there is no field in which such a claim could be recorded.
  • Not a psychological or educational assessment instrument. It is not normed, not standardised, and not validated as one.
  • Not a predictor. It does not forecast future performance, readiness or attainment.
  • Not a comparison between children. No ranking, no benchmarking, no cross-school analytics and no league table.
  • Not a monitor. No device tracking, no location, no browsing history, no activity monitoring, no keystroke capture. It records the work a child deliberately submits, and nothing else.
  • Not a substitute for professional judgement — a teacher's or a specialist's.
  • Not a conversation. The AI never talks to a child, and there is no messaging of any kind in the student app.

A school evaluating this against a departmental checklist should read the list above as the answer to the “what is the AI used for” question, section 4 as the answer to the fairness question, and sections 5 to 8 as the answer to the data questions.

Who owns this, the separation, and how it is reviewed

Officer roles are named rather than people, so a change of founder does not date this page. The signed Order Form names them.

The separation is the control

The AI Accountability Owner writes no part of the AI. The role does not write the analysis pipeline, does not own the staff dashboard and holds no decision about what the product builds next. It reviews, tests and approves the developer's work rather than authoring it. An arrangement where the person who builds the AI is also the person who decides whether it may run is not oversight — it is the same decision made twice.

Who holds what, and why they are different people.
RoleHeld byHolds
AI Accountability OwnerFounder and CTOWhether what was built may run. The boundary in section 5, the output contract, the model and prompt versions, the AI risk assessment, the weekly quarantine review, every contested output, and the halt powers below. Holds the credentials to execute a halt without anyone's help.
AI development and product ownershipFounder and COOWrites the analysis pipeline, the payload builder and the prompt handling; owns the staff dashboard; decides what is built next. Ships through review, holds no standing production access, and cannot deploy.
Privacy OfficerFounder and CEOPrivacy complaints, access, correction and deletion requests. Holds no production access of any kind, which is why this role keeps the copy of every halt direction — a record that cannot be quietly edited.

The three halts

What the AI Accountability Owner can stop, and how fast.
PowerEffectIn force
Hold a releaseA pending deployment of the staff portal or the student app does not ship.Immediately
Block or revert a model or prompt version changeThe pinned model or prompt version stays where it was, or returns to the previous pinned pair.Same business day
Stop model calls entirelyThe pipeline refuses to run. Marking, growth history and the diagnostic are unaffected, because all three are deterministic and read no AI field.Within one hour of the instruction, and in any case the same day

How a halt works, and how it ends

A halt is a written instruction to the founder who develops the AI, copied to the Privacy Officer, naming which halt and the reason in one sentence. No form, no meeting, no approval. The developer has no discretion to refuse it, to delay it, or to ship around it. It is lifted only by the AI Accountability Owner, in writing, saying what was found, what changed and what test now covers it — and “nothing was wrong” is a legitimate outcome that gets written down as one.

Where the two officers disagree, the halt stands while the disagreement is resolved. Where the disagreement concerns a change the developer authored — which is most of them — there is no casting vote: the more restrictive position stands until all three founders agree in writing. Nobody casts the deciding vote on their own release.

Staying current

Quarterly, the AI Accountability Owner writes a one-page scan of the model provider's model cards, safety documentation and usage policy changes; the OAIC's guidance on privacy and AI; the eSafety Commissioner's Safety by Design material; the Australian Framework for Generative AI in Schools and any Victorian departmental guidance. It says what changed, whether anything we do is now wrong, and what we will do about it.

Anything it finds that changes a control becomes a change to this policy or a row in our AI risk assessment within one month, or a written note of why not. A scan that never changes anything is not being read. On any model version change, the provider's own evaluation and safety notes for that version are read before the change is approved, not after.

Questions from outside, and what we do with them

We take questions about this product from outside it, and we write down what we change as a result.

  • We respond to the Victorian Department of Education, to Safer Technologies 4 Schools, and to any school's own privacy or AI review, on request.
  • Where an Australian regulator consults publicly on AI in education, we will make a submission and record it in the change history on this page.
  • We publish this policy and our AI risk assessment in full rather than summarising them, so that a parent, a principal or a journalist reads the same document an assessor does.

Stated as what it is

There is no engagement to report yet. This is a commitment with a first review date of go-live date — to be confirmed, not a description of something already happening.

Review

  • This policy is reviewed at least annually, on the review date at the top of this page.
  • It is reviewed again on any change to the model version, the prompt, the output contract, the payload, or the country the model is served from — before that change takes effect, not after.
  • The ambiguous-name exemption list in section 4 is reviewed annually with it, and so is the decision not to collect attributes about children in order to measure fairness.
  • Each review records the version, the date, who reviewed it and what changed. A review that finds no change still gets a dated entry.

Contact us, and when this policy changes

Entityentity name — to be confirmed (ACN acn — to be confirmed, ABN abn — to be confirmed), trading as Reason Education
AI Accountability OwnerFounder and CTO — contested outputs, fairness findings, model and prompt changes
Privacy OfficerFounder and CEO — privacy questions and complaints
Contesting an AI output, support and faultssupport email — to be confirmed
Privacy questions, complaints, access and deletion requestsprivacy email — to be confirmed
Security reports and responsible disclosuresupport email — to be confirmed, with SECURITY at the start of the subject line. There is no separate security mailbox: a three-person company that advertises four addresses it does not watch is worse off than one that advertises three it does.
Child safety concernschild safety email — to be confirmed
Urgent — security or child safety, including out of hoursphone number — to be confirmed
Postregistered office address — to be confirmed

When this policy changes

  • We notify every account holder at least 14 days before a change takes effect, with a plain summary of what changed and why. Schools are told at the same time as their staff.
  • A change to the model version, the prompt or the output contract is notified under section 914 days, with the test basis stated — and appears in the published change history.
  • A change to the country the model is served from is notified as a hosting change: 90 days, always before the first call from the new location, and a school that objects may terminate the affected part of the service, or the whole agreement, with a pro-rata refund.
  • A school that does not want a change may turn AI Analysis off under section 11 and keep everything else.

Issued

Responsible AI Policy v1.1, 26 August 2026. Next scheduled review 26 August 2027, or sooner on any change to the model version, the prompt, the output contract, the payload, or the country the model is served from.