Pilot proposal · next term

NoetanoetaUniversity of
West Florida

A degree should mean the student can explain the work.

AI can produce the work. It cannot stand in for a student asked, out loud, to walk through their own reasoning. Noeta holds that conversation — a few spoken minutes after an assessment — and brings the instructor what was said, in the student’s own words. It never issues a verdict, and it never runs a camera. The instructor decides.

This is a proposal to build that check with UWF rather than sell it to UWF: a design-partner pilot in two or three sections, where each side leaves with something it could not have made alone.

01Our position

A higher standard, not a softer one.

Swapping surveillance for a conversation can sound like softening the check. It is the opposite, and it is worth stating precisely: what the standard asks of students, what it produces when a finding is challenged, and what it owes the students who did their own work.

It raises the bar, it does not move it

The question stops being “does this document look original” and becomes “can this student account for the work.” A student who can explain the reasoning has met a higher standard than one whose file passed a similarity check.

Evidence that survives a hearing

A detector score has already been ruled insufficient in federal court. A recording of a student unable to explain the terms in the submission — with the transcript, timestamps and their own words quoted — is built for a misconduct process, not a dashboard.

Protecting honest students is integrity too

A wrongly accused student is an integrity failure, not a rounding error. That is why “insufficient evidence” is a real result here rather than a polite way of saying guilty, and why a session that failed to record is reported as a session to re-run.

The finding belongs to a person

Noeta never issues a verdict — not because it is neutral about misconduct, but because a misconduct finding is an academic judgment that belongs to a faculty member and a process. Noeta’s job is to put better evidence in front of that judgment, and to be explicit about what it does not know.

Departments hold this differently, and Noeta is built to sit under the stricter reading rather than the looser one: it records everything, quotes the student verbatim, keeps the audio, and states plainly where its evidence runs out.

02The workload

Fifty reviews become two or three.

UWF runs Respondus LockDown Browser inside Canvas eLearning today. Because its flags cannot be trusted, instructors watch every session by hand — the ranking reorders the work without reducing it. The published evidence below says that review is not even telling them what they think it is.

Today

  • Every session is watched — flagged or not. The ranking is not trusted enough to skip any of them.
  • So the flag saves nothing. It reorders a queue that gets worked through in full either way.
  • What ranking there is comes from a model that cannot see a second device.
  • And none of it shows whether the student understood the work.

Instead

  • No video review at all. The whole pass is retired, not replaced with a second one.
  • Two or three recordings per section, where the explanation didn’t match the work.
  • Each one arrives with the student’s own words quoted next to the finding.
  • The instructor still decides — but starts from evidence, not a ranking.

50 2 or 3

In a 50-student section that is fifty sessions watched today, and two or three tomorrow. The saving is not a better ranking — it is not having to open the other forty-seven at all.

What that review actually produced

Observed at UWF

Both of these were observed in a single sitting, working through one course’s sessions.

False positive

Ranked high for something that was allowed

A UWF instructor, working through one course’s sessions, watched the system rank a student near the top of the cheating scale for using a handheld device — on an assessment where that device was permitted. The tool had no way to know the difference.

False negative

Ranked low while clearly cheating

In the same sitting, a student who was plainly cheating was scored low on that same scale. Both were in front of a webcam the whole time. Watching every session in the course is what produced these two answers, and both of them are wrong.

Neither outcome is a misconfiguration. A browser lock can only see the machine running the exam, and a webcam cannot tell a permitted device from a forbidden one. The review was working exactly as designed.

This is documented, not anecdotal

Flags are not findings, and the vendor says so

Respondus documents flagging as a review aid rather than a determination — a flag marks a moment for an instructor to look at, not evidence of misconduct. Every flag still has to be watched and judged by a person, which is the workload this proposal is about.

Respondus Monitor product documentation

It deters. The evidence it detects is weak

A peer-reviewed review of remote proctoring finds strong support for a deterrent effect, but that the limited evidence on its ability to actually detect cheating suggests it may be ineffective.

Higher Education Research & Development, 2023

The second device is a structural blind spot

A lockdown browser can only control the machine running the exam. In a peer-reviewed study of student perceptions, 21% named a second device as the way to cheat through a proctored exam.

Examining the Examiners, arXiv

UWF already runs this stack

Respondus LockDown Browser is deployed inside UWF's Canvas eLearning, with the university publishing its own setup guides for both instructors and students.

UWF Public Knowledge Base

03The other half

What it puts back

Everything above is work removed. This is the part that is worth doing even if academic integrity were a solved problem — and the part a proctoring product structurally cannot claim, because watching a student take a test adds nothing to their education.

Assessment moved to text, and the reps went with it. A student can complete a degree without once being asked to explain their reasoning out loud to anyone — and then graduate into a job market that rates verbal communication among the things it most wants and least reliably finds.

25 pts

The gap students cannot see

Employers rate new graduates well below how those graduates rate themselves on communication — a roughly 25-point gap in NACE's career-readiness data. Nothing in a text-based degree measures it, so nobody finds out until an interview.

NACE, perceptions gap

Deeper

Orals change how students study

A 2025 systematic review of oral assessment in higher education, and course-level studies alongside it, report better retention and measured improvement in final marks — students move from surface learning toward explanation because they know they will have to say it out loud.

Assessment & Evaluation in Higher Education, 2025

2 min

Short and low-stakes is the point

Repeated low-stakes speaking reduces communication apprehension — frequency works because explaining yourself out loud stops being exceptional. Two minutes to a machine, with delivery explicitly not graded, is closer to an exposure protocol than to a presentation.

Research on speaking apprehension
A measurement UWF would make first

This is also measurable, and a pilot should measure it. A standard communication-apprehension instrument administered before and after a term costs almost nothing and would tell UWF something nobody currently knows: whether short, frequent, ungraded speaking makes students more willing to speak.

04The partnership

What UWF gets, and what noeta gets

This is not a finished product being trialed. It is an early one being shaped, and the shaping is the offer: a design partner decides what noeta asks, what it reports and what it refuses to do — while those decisions are still cheap to change. In plain terms, the trade is this.

What UWF gets

University of
West Florida
  • First say in what this becomes

    The conversation, the report and the refusals are still being decided. A design partner decides them with us, while changing them is still cheap.

  • The adjudicated dataset

    Nobody has a labeled set of oral checks where a person judged which explanations were the student’s own. Graduate assistants reviewing pilot sessions would build the first one — the piece the field is missing — on UWF’s campus.

  • A measurement nobody has made

    A standard communication-apprehension instrument, before and after a term, costs almost nothing and would tell UWF something nobody currently knows: whether short, frequent, ungraded speaking makes students more willing to speak.

  • An answer either way

    Review hours are measured, not promised. If they don't fall, the pilot has answered its question — and that is worth knowing from two sections rather than a department.

What noeta gets

Noeta
  • A real course to learn from

    Real rooms, real microphones, real sections. No demo reproduces the conditions a campus provides by default.

  • Judgment beside every session

    Both failure modes in this proposal were caught by a UWF instructor reading sessions by hand. That reading — a person saying what the system got right and wrong — is what calibrates a product. Without it, every claim is guesswork.

  • Scrutiny before scale

    A university asks about FERPA, consent, retention and review before any real student data moves. Meeting that bar for two sections is how noeta earns the right to more of them.

05Demo economics

What it costs to run today

These are the demo’s costs, at public list prices, on hosted services. Every number below is what this working system bills right now with no volume agreement, no reserved capacity and nothing self-hosted — which makes it the ceiling rather than the estimate. It is here so a pilot can be budgeted from something real instead of a projection.

Measured from the working system, not estimated. A complete check is three things: a short spoken conversation, a transcription of the recording, and one assessment pass over what was said. Here is what each one costs at list pricing.

$0.103

noeta’s voice

~1,040 characters spoken — Flash v2.5, ElevenLabs

$0.051

The conversation

Four or five short turns — Claude Opus, Anthropic

$0.109

The assessment

One pass over the transcript — Claude Opus, Anthropic

$0.003

Transcription

The whole recording — Whisper Large v3 Turbo, Groq

The voice is the newest line and the second largest. It is there because the questions are no longer displayed as text: a question on a screen can be photographed and handed to a chatbot in one tap, and in our own testing a session run exactly that way scored 93 out of 100 with nothing raised. Spoken, there is nothing to photograph. Roughly ten cents a student is what that costs.

Transcription is about 1% of the bill and does the most work of the four. The browser’s own speech recognition is free but unreliable in a real room — in testing it dropped eighty seconds of a student answering every question, and the report that came back said the student had been non-responsive. The recording is now transcribed properly before anything judges it, for a third of a cent.

One student, one oral check$0.27
A 50-student section, every student$13.30
A 50-student section, 20% sample$2.66
A 120-student section, every student$31.92

For context, checking a section of 50 students in full costs about what two coffees do, and about a fifth of that under the sampling model this proposal actually recommends. The cost that matters is not the model bill — it is the hours currently spent watching video.

Cost scales with the number of checks run, so an instructor sampling 20% of a class pays for 20% of a class. Nothing here is a per-seat license — these are the underlying model and synthesis bills at list price, before any volume rate.

What the demo is running on

Named in full, because a pilot needs to know whose infrastructure a student’s voice passes through before anyone signs anything.

Application
Next.js on Vercel. The eLearning shell is a look-alike, not an Instructure product.
Voice out
ElevenLabs Flash v2.5, streamed through noeta’s server so the key never reaches a browser.
Voice in
The browser's microphone. Each answer is its own clip; the whole session is also recorded.
Transcription
Groq-hosted Whisper Large v3 Turbo, server-side. No browser speech recognition is required.
Conversation & assessment
Claude Opus 5 via Anthropic's API — two separate passes, one to talk, one to judge.
Room analysis
ffmpeg on the recording. Levels and frequency bands only; no speaker identification.
Storage
Vercel Blob, private. Recordings are reachable only through an authenticated route.

06Security & data

Security, and what a pilot would require

A student’s recorded voice is about as sensitive as anything a course collects, so this table separates what a pilot would have to put in place from what the demo actually does today. The right-hand column is deliberately unflattering — a demo that overstates its posture is worse than one that admits it is a demo.

Access

LTI 1.3 launch from eLearning, so identity comes from UWF SSO and only the section's instructor of record can open its console.

None. The demo is reachable by URL — treat every link as public and use only volunteer recordings.

In transit

TLS 1.2+ everywhere, including to each model vendor.

Already the case — HTTPS end to end.

At rest

AES-256 on recordings and reports, with keys managed by the storage provider.

Vercel Blob, private buckets, encrypted at rest. Recordings are served through a server route rather than a public blob URL — but that route is not yet authenticated, which is the row above.

Retention

A stated window — our recommendation is the grade-appeal deadline plus 30 days — then automatic deletion, and deletion on request within 30 days.

No automatic deletion. Sessions persist until removed by hand.

Subprocessors

Anthropic, Groq, ElevenLabs and Vercel, each under a written DPA with no-training-on-customer-data confirmed in writing before any real student data moves.

Same four vendors. Their terms have not yet been papered for UWF — that is a pilot precondition, not a claim.

FERPA

Recordings and reports are education records: instructor-only, never shown to the student, never shared with other students, disclosed under the school-official exception.

Enforced in the product — student-facing responses are stripped of scores, signals and analysis at the API layer.

Independent assurance

Vendor SOC 2 reports on request; noeta itself would need a security review before handling real coursework. It has not had one.

None. Stated plainly rather than implied.

On learning from the data

The design-partner work

The most valuable thing a pilot would produce is not recordings — it is adjudicated recordings. Nobody currently has a labeled set of oral checks where a human has said which explanations were the student’s own, and without one, every claim about detecting AI-assisted answers is calibrated against guesswork. Graduate assistants reviewing and scoring sessions would build exactly that, and it is the piece the field is missing.

Training on the audio itself is necessary — how something was said is most of the signal, and a transcript throws it away. It is also allowed. What it is not is allowed by anonymization, and that distinction is where these proposals usually go wrong. A recorded voice cannot be de-identified. It is a biometric identifier; removing a name from a file does nothing, because the voice is the identifier. The route to training on it is consent, not stripping.

Concretely, that means: written, opt-in consent taken separately from course enrollment, specific to voice and to model training, revocable at any time and with no consequence for a student’s grade either way; IRB review where the work constitutes research; recordings held in UWF-approved storage with identity kept in a separate keyed table; a stated retention window; and withdrawal of consent propagating into the training set, not just the archive. Florida has no biometric-privacy statute, but a remote student may be sitting in Illinois, Texas or Washington, which do — so the program should be built to the strictest of them rather than to the campus’s own state.

There is also a technical route worth piloting alongside it. Speaker anonymization — an active research area with its own recurring benchmark, the VoicePrivacy Challenge — transforms a recording so the speaker’s identity is no longer recoverable while the timing, hesitation and prosody survive. That is a genuinely useful split for this product, because the identity is the part we do not need and the manner of speaking is the entire part we do. It would have to be validated rather than assumed: the open question is whether a transform strong enough to defeat speaker recognition also flattens the very disfluency and pacing noeta is trying to learn from. Testing that on consented pilot data is a concrete first piece of work, and if it holds, later cohorts could contribute to training without their voices being retained at all.

07The ask

What we’re asking for

A design-partner pilot in two or three sections, next term.

  • Two assessments per section

    The first oral exam runs for everyone, which sets a baseline and means being selected later carries no stigma. After that, a randomized 10–20% sample.

  • LockDown review switched off for those assessments

    This is the part that makes it a replacement. Running both would prove nothing except that faculty have more work.

  • No LMS integration required

    Students follow a link. Canvas and LTI come later, once the pilot has earned it.

  • Volunteer participation, with a stated incentive

    Extra credit, or the oral standing in for part of the exam weight — the instructor’s choice.

What we measure

  • Faculty hours spent on review, before and after.
  • How often noeta’s flag agreed with the instructor’s own judgment.
  • How often it flagged a student the instructor then cleared.
  • Student experience, collected directly rather than inferred.
  • Time each check took, end to end.

If review hours don’t fall, the pilot has answered its question and we would rather know that from two sections than from a department.

08Questions

Questions we expect

Does this add another thing to review?
No — that is the point of it. The pass through every session goes away. What arrives instead is two or three recordings per section where the explanation didn’t match the work, each with the student’s own words quoted beside the finding. If it ever becomes a second queue on top of the first, it has failed.
Is there a camera?
No. Camera-based flagging is the part that produced both failures above, room scans have been ruled unconstitutional elsewhere, and video is exactly the review burden faculty are asking to be rid of. Noeta uses voice and the open microphone only.
Then how do you know they weren’t reading off a phone?
For a determined student, we don’t — and we tested that rather than assumed it. Photographing an on-screen question, handing it to a chatbot and reading the reply back scored 93 out of 100 with nothing raised. That is why the questions are now spoken and never displayed: there is nothing to photograph. Relaying is still possible — a second device can record audio — but it costs a round trip on every question, and that wait sits in the recording between noeta finishing and the student starting. Noeta reports the wait; it does not claim to have caught anyone. No remote check can honestly promise more, and the research on camera proctoring says the same.
Isn’t this hard on students with speaking anxiety?
It is the objection worth taking seriously, and the design answers it with time rather than with an exemption. Delivery is explicitly not graded — noeta assesses reasoning and refuses to score cadence, fluency or confidence. There is no visible timer and no cost to thinking: after about forty-five seconds of quiet noeta says some version of “take your time,” never “hurry up,” and a student can ask for the question again as often as they need — re-asks are recorded on the session, but nothing scores them. It is also a machine rather than a room of peers, and repeated low-stakes speaking is the established way apprehension comes down. What anxiety does not do is turn the check into a typing exercise. Typed answers are an accommodation for a student who cannot speak, granted by the instructor through the same route as any other accommodation on a campus — and today that setting applies to a whole section, so a per-student grant would need roster integration the demo has not built.
Is student voice data used to train anything?
Not in the demo, and not in a pilot without written, revocable, opt-in consent that carries no grade consequence. A recorded voice is a biometric identifier — it cannot be de-identified the way a name in a spreadsheet can — so any training work would use derived text and instructor labels with identifiers stripped, never raw audio, and would go through UWF’s IRB if it counts as research.
What if a student’s microphone fails?
Noeta asks the question again. If two answers in a row come through silent it stops and tells the student the microphone isn’t being picked up, rather than recording blanks. A session that captured nothing is reported to the instructor as a session to re-run — never as a finding about the student.
Is this fair to non-native English speakers?
It scores substance, never accent, fluency or delivery, and the check can run in the student’s strongest language where course policy allows. The comparison worth making is against the status quo: AI detectors falsely flag non-native English writing at 61%.
Who decides whether something is misconduct?
The instructor, in every case — and that is a position on integrity, not a way of avoiding one. A misconduct finding is an academic judgment that belongs to a faculty member and a process. Noeta never assigns a grade, never issues a verdict, and posts no penalty automatically. Where it has no evidence it says so; “insufficient evidence” is a first-class result rather than a euphemism for guilt.

Noeta never issues a verdict.
The instructor decides.

It brings the instructor the student’s own words, says plainly where its evidence runs out, and leaves the judgment where it has always belonged.