Pilot proposal · next term
AI can produce the work. It cannot stand in for a student asked, out loud, to walk through their own reasoning. Noeta holds that conversation — a few spoken minutes after an assessment — and brings the instructor what was said, in the student’s own words. It never issues a verdict, and it never runs a camera. The instructor decides.
This is a proposal to build that check with UWF rather than sell it to UWF: a design-partner pilot in two or three sections, where each side leaves with something it could not have made alone.
01Our position
Swapping surveillance for a conversation can sound like softening the check. It is the opposite, and it is worth stating precisely: what the standard asks of students, what it produces when a finding is challenged, and what it owes the students who did their own work.
The question stops being “does this document look original” and becomes “can this student account for the work.” A student who can explain the reasoning has met a higher standard than one whose file passed a similarity check.
A detector score has already been ruled insufficient in federal court. A recording of a student unable to explain the terms in the submission — with the transcript, timestamps and their own words quoted — is built for a misconduct process, not a dashboard.
A wrongly accused student is an integrity failure, not a rounding error. That is why “insufficient evidence” is a real result here rather than a polite way of saying guilty, and why a session that failed to record is reported as a session to re-run.
Noeta never issues a verdict — not because it is neutral about misconduct, but because a misconduct finding is an academic judgment that belongs to a faculty member and a process. Noeta’s job is to put better evidence in front of that judgment, and to be explicit about what it does not know.
Departments hold this differently, and Noeta is built to sit under the stricter reading rather than the looser one: it records everything, quotes the student verbatim, keeps the audio, and states plainly where its evidence runs out.
02The workload
UWF runs Respondus LockDown Browser inside Canvas eLearning today. Because its flags cannot be trusted, instructors watch every session by hand — the ranking reorders the work without reducing it. The published evidence below says that review is not even telling them what they think it is.
Today
Instead
50 2 or 3
In a 50-student section that is fifty sessions watched today, and two or three tomorrow. The saving is not a better ranking — it is not having to open the other forty-seven at all.
Both of these were observed in a single sitting, working through one course’s sessions.
False positive
A UWF instructor, working through one course’s sessions, watched the system rank a student near the top of the cheating scale for using a handheld device — on an assessment where that device was permitted. The tool had no way to know the difference.
False negative
In the same sitting, a student who was plainly cheating was scored low on that same scale. Both were in front of a webcam the whole time. Watching every session in the course is what produced these two answers, and both of them are wrong.
Neither outcome is a misconfiguration. A browser lock can only see the machine running the exam, and a webcam cannot tell a permitted device from a forbidden one. The review was working exactly as designed.
Respondus documents flagging as a review aid rather than a determination — a flag marks a moment for an instructor to look at, not evidence of misconduct. Every flag still has to be watched and judged by a person, which is the workload this proposal is about.
Respondus Monitor product documentationA peer-reviewed review of remote proctoring finds strong support for a deterrent effect, but that the limited evidence on its ability to actually detect cheating suggests it may be ineffective.
Higher Education Research & Development, 2023A lockdown browser can only control the machine running the exam. In a peer-reviewed study of student perceptions, 21% named a second device as the way to cheat through a proctored exam.
Examining the Examiners, arXivRespondus LockDown Browser is deployed inside UWF's Canvas eLearning, with the university publishing its own setup guides for both instructors and students.
UWF Public Knowledge Base03The other half
Everything above is work removed. This is the part that is worth doing even if academic integrity were a solved problem — and the part a proctoring product structurally cannot claim, because watching a student take a test adds nothing to their education.
Assessment moved to text, and the reps went with it. A student can complete a degree without once being asked to explain their reasoning out loud to anyone — and then graduate into a job market that rates verbal communication among the things it most wants and least reliably finds.
25 pts
The gap students cannot see
Employers rate new graduates well below how those graduates rate themselves on communication — a roughly 25-point gap in NACE's career-readiness data. Nothing in a text-based degree measures it, so nobody finds out until an interview.
NACE, perceptions gapDeeper
Orals change how students study
A 2025 systematic review of oral assessment in higher education, and course-level studies alongside it, report better retention and measured improvement in final marks — students move from surface learning toward explanation because they know they will have to say it out loud.
Assessment & Evaluation in Higher Education, 20252 min
Short and low-stakes is the point
Repeated low-stakes speaking reduces communication apprehension — frequency works because explaining yourself out loud stops being exceptional. Two minutes to a machine, with delivery explicitly not graded, is closer to an exposure protocol than to a presentation.
Research on speaking apprehensionThis is also measurable, and a pilot should measure it. A standard communication-apprehension instrument administered before and after a term costs almost nothing and would tell UWF something nobody currently knows: whether short, frequent, ungraded speaking makes students more willing to speak.
04The partnership
This is not a finished product being trialed. It is an early one being shaped, and the shaping is the offer: a design partner decides what noeta asks, what it reports and what it refuses to do — while those decisions are still cheap to change. In plain terms, the trade is this.
First say in what this becomes
The conversation, the report and the refusals are still being decided. A design partner decides them with us, while changing them is still cheap.
The adjudicated dataset
Nobody has a labeled set of oral checks where a person judged which explanations were the student’s own. Graduate assistants reviewing pilot sessions would build the first one — the piece the field is missing — on UWF’s campus.
A measurement nobody has made
A standard communication-apprehension instrument, before and after a term, costs almost nothing and would tell UWF something nobody currently knows: whether short, frequent, ungraded speaking makes students more willing to speak.
An answer either way
Review hours are measured, not promised. If they don't fall, the pilot has answered its question — and that is worth knowing from two sections rather than a department.
A real course to learn from
Real rooms, real microphones, real sections. No demo reproduces the conditions a campus provides by default.
Judgment beside every session
Both failure modes in this proposal were caught by a UWF instructor reading sessions by hand. That reading — a person saying what the system got right and wrong — is what calibrates a product. Without it, every claim is guesswork.
Scrutiny before scale
A university asks about FERPA, consent, retention and review before any real student data moves. Meeting that bar for two sections is how noeta earns the right to more of them.
05Demo economics
These are the demo’s costs, at public list prices, on hosted services. Every number below is what this working system bills right now with no volume agreement, no reserved capacity and nothing self-hosted — which makes it the ceiling rather than the estimate. It is here so a pilot can be budgeted from something real instead of a projection.
Measured from the working system, not estimated. A complete check is three things: a short spoken conversation, a transcription of the recording, and one assessment pass over what was said. Here is what each one costs at list pricing.
$0.103
noeta’s voice
~1,040 characters spoken — Flash v2.5, ElevenLabs
$0.051
The conversation
Four or five short turns — Claude Opus, Anthropic
$0.109
The assessment
One pass over the transcript — Claude Opus, Anthropic
$0.003
Transcription
The whole recording — Whisper Large v3 Turbo, Groq
The voice is the newest line and the second largest. It is there because the questions are no longer displayed as text: a question on a screen can be photographed and handed to a chatbot in one tap, and in our own testing a session run exactly that way scored 93 out of 100 with nothing raised. Spoken, there is nothing to photograph. Roughly ten cents a student is what that costs.
Transcription is about 1% of the bill and does the most work of the four. The browser’s own speech recognition is free but unreliable in a real room — in testing it dropped eighty seconds of a student answering every question, and the report that came back said the student had been non-responsive. The recording is now transcribed properly before anything judges it, for a third of a cent.
For context, checking a section of 50 students in full costs about what two coffees do, and about a fifth of that under the sampling model this proposal actually recommends. The cost that matters is not the model bill — it is the hours currently spent watching video.
Cost scales with the number of checks run, so an instructor sampling 20% of a class pays for 20% of a class. Nothing here is a per-seat license — these are the underlying model and synthesis bills at list price, before any volume rate.
Named in full, because a pilot needs to know whose infrastructure a student’s voice passes through before anyone signs anything.
06Security & data
A student’s recorded voice is about as sensitive as anything a course collects, so this table separates what a pilot would have to put in place from what the demo actually does today. The right-hand column is deliberately unflattering — a demo that overstates its posture is worse than one that admits it is a demo.
Area
What a pilot requires
What the demo does today
Access
LTI 1.3 launch from eLearning, so identity comes from UWF SSO and only the section's instructor of record can open its console.
None. The demo is reachable by URL — treat every link as public and use only volunteer recordings.
In transit
TLS 1.2+ everywhere, including to each model vendor.
Already the case — HTTPS end to end.
At rest
AES-256 on recordings and reports, with keys managed by the storage provider.
Vercel Blob, private buckets, encrypted at rest. Recordings are served through a server route rather than a public blob URL — but that route is not yet authenticated, which is the row above.
Retention
A stated window — our recommendation is the grade-appeal deadline plus 30 days — then automatic deletion, and deletion on request within 30 days.
No automatic deletion. Sessions persist until removed by hand.
Subprocessors
Anthropic, Groq, ElevenLabs and Vercel, each under a written DPA with no-training-on-customer-data confirmed in writing before any real student data moves.
Same four vendors. Their terms have not yet been papered for UWF — that is a pilot precondition, not a claim.
FERPA
Recordings and reports are education records: instructor-only, never shown to the student, never shared with other students, disclosed under the school-official exception.
Enforced in the product — student-facing responses are stripped of scores, signals and analysis at the API layer.
Independent assurance
Vendor SOC 2 reports on request; noeta itself would need a security review before handling real coursework. It has not had one.
None. Stated plainly rather than implied.
The most valuable thing a pilot would produce is not recordings — it is adjudicated recordings. Nobody currently has a labeled set of oral checks where a human has said which explanations were the student’s own, and without one, every claim about detecting AI-assisted answers is calibrated against guesswork. Graduate assistants reviewing and scoring sessions would build exactly that, and it is the piece the field is missing.
Training on the audio itself is necessary — how something was said is most of the signal, and a transcript throws it away. It is also allowed. What it is not is allowed by anonymization, and that distinction is where these proposals usually go wrong. A recorded voice cannot be de-identified. It is a biometric identifier; removing a name from a file does nothing, because the voice is the identifier. The route to training on it is consent, not stripping.
Concretely, that means: written, opt-in consent taken separately from course enrollment, specific to voice and to model training, revocable at any time and with no consequence for a student’s grade either way; IRB review where the work constitutes research; recordings held in UWF-approved storage with identity kept in a separate keyed table; a stated retention window; and withdrawal of consent propagating into the training set, not just the archive. Florida has no biometric-privacy statute, but a remote student may be sitting in Illinois, Texas or Washington, which do — so the program should be built to the strictest of them rather than to the campus’s own state.
There is also a technical route worth piloting alongside it. Speaker anonymization — an active research area with its own recurring benchmark, the VoicePrivacy Challenge — transforms a recording so the speaker’s identity is no longer recoverable while the timing, hesitation and prosody survive. That is a genuinely useful split for this product, because the identity is the part we do not need and the manner of speaking is the entire part we do. It would have to be validated rather than assumed: the open question is whether a transform strong enough to defeat speaker recognition also flattens the very disfluency and pacing noeta is trying to learn from. Testing that on consented pilot data is a concrete first piece of work, and if it holds, later cohorts could contribute to training without their voices being retained at all.
07The ask
A design-partner pilot in two or three sections, next term.
Two assessments per section
The first oral exam runs for everyone, which sets a baseline and means being selected later carries no stigma. After that, a randomized 10–20% sample.
LockDown review switched off for those assessments
This is the part that makes it a replacement. Running both would prove nothing except that faculty have more work.
No LMS integration required
Students follow a link. Canvas and LTI come later, once the pilot has earned it.
Volunteer participation, with a stated incentive
Extra credit, or the oral standing in for part of the exam weight — the instructor’s choice.
What we measure
If review hours don’t fall, the pilot has answered its question and we would rather know that from two sections than from a department.
08Questions
It brings the instructor the student’s own words, says plainly where its evidence runs out, and leaves the judgment where it has always belonged.