Concept brief · for faculty

Noetanoeta

noh-EH-tuh· from the Greek noēta — “that which can be understood.” The knowable, made visible.

Assessment assurance for the AI era.

Students can use AI to produce the work — noeta checks whether they can still explain the thinking behind it, in conversation, at a scale one professor could never run by hand.

The one intervention with real evidence behind it is asking students to explain their own work out loud. A professor with 120 students cannot run 120 oral exams. The bottleneck was never the pedagogy — it was scale.

120
students submit
AI-assisted or not — noeta never inspects the document.
20
sampled for an oral exam
A randomized 10–20%, after everyone does the first one.
2–3
reach you
Only where understanding didn’t match the work. You decide.
Borrowed from audit practice: don’t inspect every transaction — test controls, sample the population, escalate exceptions.

Why this exists

The evidence

Detection and proctoring are a losing arms race. Detectors are unreliable and legally indefensible; surveillance inspects the room or the document, never the student’s head.

88%

of students now use generative AI on assessments

HEPI Student Generative AI Survey 2025 — up from 53% the year before

HEPI, 2025

61%

false-positive rate on non-native English writers

Stanford (Liang et al.) across seven detectors — the students least able to absorb a wrong accusation

Liang et al., arXiv

0 of N

students told to cheat who were caught by automated proctoring

one human reviewer caught one; the authors compared it to a placebo (cited in CHI 2023)

CHI 2023

Vendors withdrew their own detectors

OpenAI retired its AI Text Classifier in 2023 citing low accuracy — the company with the most reason to make detection work, saying publicly that it did not.

OpenAI, retiring the classifier

Room-scan proctoring has been ruled unconstitutional

Ogletree v. Cleveland State University (N.D. Ohio, 2022) held a remote room scan an unreasonable search under the Fourth Amendment.

Ogletree v. Cleveland State

Proctoring deters. The evidence that it detects is weak

A peer-reviewed review finds strong support for a deterrent effect, and limited evidence that remote proctoring actually catches anyone.

Higher Education Research & Development, 2023

Students name the second device themselves

A lockdown browser controls the machine running the exam and nothing else. In a peer-reviewed study of student perceptions, a fifth named a second device as the obvious way around it.

Examining the Examiners, arXiv

Oral assessment improves the learning, not just the assurance

A 2025 systematic review of oral assessment in higher education reports better retention and deeper conceptual understanding. It is the only method here that gives a student something back.

Assessment & Evaluation in Higher Education, 2025

And it has never scaled past one professor’s calendar

Which is the whole reason this exists. Everything above is already known; nobody has been able to act on it at the size of a real course.

The check

How it works

  1. 01

    Students submit as usual

    AI-assisted or not — we stop policing inputs entirely. Assignments can carry a declared AI-use level, so the rules are explicit instead of implied.

  2. 02

    A short voice conversation

    Adaptive, and about their specific work: explain this claim, define this term, what breaks if X changes. A few spoken minutes, in the student’s strongest language where policy allows.

  3. 03

    Only exceptions reach you

    Sessions where the explanation matched the work drop to the bottom of the list and need nothing from you — but noeta never records a decision on your behalf. A gap between the work and the comprehension behind it comes up top, with the transcript and the student’s own words, so you check the reasoning rather than trust a score. This replaces the flag queue; it is not a second one.

The unseen case

One exchange cannot be prepared for. After the focus items, noeta builds a single scenario live out of what the student just said, about a case the assessment never covered — it does not exist until the student has been talking. Anyone holding the original questions can rehearse those; not this one. It raises the cost of outside help by a round trip, and it is reported as exactly what it is — never presented as proof.

The alternatives

The landscape

AI detectors

Turnitin AI, GPTZero

~26% accurate on paraphrased text, 61% false positives on non-native writers, disabled by 50+ universities, and ruled insufficient evidence in federal court.

Proctoring

Honorlock, Proctorio, Respondus

An arms race already lost — Cluely defeats it invisibly. Room scans ruled unconstitutional. Verifies identity and behavior, never understanding.

Process tracking

Grammarly Authorship, Turnitin Clarity

Keystroke surveillance, not proof of comprehension. Faculty backlash, and trivially gamed by retyping.

AI oral exams

Vivaproof, VivaEdu, Sherpa

The closest competitors — they validate the category. All sell per-assignment checks to individual teachers. None offer sampling methodology, institution-level assurance, or a defensible evidence workflow.

Why now

Detection has failed publicly enough that institutions are acting on it, faculty are already improvising oral exams by hand, and no product has made that affordable. The demand is validated and the category has no winner.

How it spreads

Bottom-up: instructors free, departments licensed — the path Turnitin took. No LMS integration is needed to start; students follow a link.

The difference

What only noeta does

An assurance model, not a checker

Statistical sampling, exception-based review, control testing — the methodology provosts and accreditors already trust. Competitors verify assignments; noeta assures programs.

Evidence that survives a hearing

Detector scores now lose in court. A transcript of a student unable to define the words in the submitted paper — with rubric mapping and a review trail — is built for due process, not for a dashboard.

Accreditation reporting

Aggregated evidence that graduates actually hold the competencies you assessed — direct input to SACSCOC-style reporting. That turns an instructor tool into a provost-level purchase.

Aligned with the pedagogy

The oral exam doubles as formative feedback, and assignments can carry a declared AI-use level rather than an implied ban. Redesign over detection is where the research consensus already sits.

It puts the speaking reps back

Assessment moved to text and took the practice with it — a student can finish a degree without explaining their reasoning aloud to anyone. Employers rate graduates roughly 25 points below how those graduates rate themselves on communication, and nothing in a text-based degree measures the gap. This is the one thing a proctoring product structurally cannot claim: watching someone take a test adds nothing to their education.

NACE career-readiness perceptions gap

Getting started

Ways to introduce it

A pilot that reads as new surveillance won’t get volunteers. Each of these gives the student something in return — pick whichever fits your course.

  • Bonus points for opting in

    Lowest friction

    Extra credit for students who complete the oral. Nobody is compelled, your first cohort self-selects, and you see what the conversation looks like before it carries any weight.

  • Let it replace part of the exam

    No added workload

    Swap 10% of the exam for the oral instead of adding to it. Students trade written work for a few minutes of talking — for many that is a better deal — and it costs you no extra grading.

  • Everyone does the first one

    Removes stigma

    Run the first oral exam for the whole section, then sample after that. A universal first pass sets the baseline, calibrates you to the format, and means being selected later carries no stigma.

  • Make it the way to clear a flag

    Protects students

    When a detector or a suspicion puts a student under review, the oral is how that student answers it. That protects honest students — above all the non-native English writers detectors flag at 61%.

  • Trade it for the lockdown browser

    Popular with students

    Students who complete the oral skip the proctoring software on that assessment. You get better evidence than a room scan, and they get their privacy back.

Straight answers

Questions faculty ask

Does this work for multiple-choice or quantitative exams?
Yes — noeta samples answered questions and asks the student to walk through the reasoning: “you chose C on 7; why is B wrong?” Guessed or copied answers can’t be reconstructed. The highest value is still take-home work, where detection has collapsed entirely.
Do instructors review it? Is an AI grading students?
The AI conducts the conversation. It never assigns a grade and no penalty is ever automatic. Aligned results post as verified; only exceptions — typically two or three per class — reach you, and you read the transcript and make every call.
Couldn’t a student cheat the oral exam itself?
Honestly: a determined relay is not reliably detectable, and we tested that rather than assumed it — an early version was beaten by photographing an on-screen question and reading a chatbot’s reply back. Questions are spoken now and never displayed, which removes the thing being photographed and forces any relay through a round trip that shows up as a wait in the recording. Noeta reports that wait. It does not claim to have caught anyone. What it does claim is that a student fluent enough to survive adaptive follow-ups is demonstrating the comprehension being assessed.
What about non-native English speakers?
The oral exam runs in the student’s strongest language where course policy allows — 50+ languages — and scores substance, never accent or fluency. Their barrier is English polish, not comprehension. The status quo is worse: detectors falsely flag 61% of non-native writing.
Isn’t this hard on students with speaking anxiety?
The design answers anxiety with time, not with an exemption. Delivery is never graded — noeta assesses reasoning and refuses to score cadence, fluency, or confidence. There is no visible timer, a long silence gets “take your time” rather than a prompt to hurry, and a student can have a question asked again as often as they need — re-asks are recorded, never scored. For a student who cannot speak at all, a typed-answer accommodation exists, granted by the instructor rather than self-selected — and today it applies to a whole section, not one student, because a per-student grant needs roster integration that isn’t built yet.
What happens when an honest student gets flagged?
A flag is a signal, not an accusation, and most resolve as “knew it, explained it differently.” Evidence packets are assembled only if you choose to pursue misconduct. Recordings are encrypted, private, and reachable only through the instructor console. Training on session data would happen only under written, revocable, opt-in consent taken separately from enrollment — never by default, and never as a condition of taking the check.
What does the first semester actually look like?
A syllabus statement and two assignments. Everyone completes the first oral exam; after that a randomized 10–20% sample. No LMS integration needed to start — students get a link, and any browser with a microphone works. Canvas and LTI come later.

Noeta never issues a verdict.
The instructor decides.

It brings the instructor the student’s own words, says plainly where its evidence runs out, and leaves the judgment where it has always belonged.