← Noor Akbari

Cost is not a goal

September 2026

Every product I've worked on has had a version of the same argument: someone wants one number that says whether we're winning. I understand the wish. I've stopped granting it.

At Rosalyn the sentence we wanted to be true was this: every exam completes, every violation is caught, reviewed, and defensibly actioned, at a cost both sides accept, without students revolting. Four clauses. The temptation is to fold them into one score. Don't.

A composite hides which term moved. When the score drops you have to decompose it before you can act, which means the decomposition was the real instrument all along. Composites also need weights, and whoever picks the weights is setting product strategy without anyone noticing.

The worse problem is cost. If cost sits inside the north star, the team wins by cutting it: a cheaper model, one fewer camera, thinner human coverage. Every one of those raises the metric and degrades the product. Cost and quality are the same knob. So cost lives outside the north star as a ceiling, a constraint the team has to respect and can never be rewarded for lowering.

That leaves one north star, which for us was the share of sessions that ended in a state the institution could defend. Not "flagged," not "clean," but defensible: the recording is complete, every event has a proctor decision, every escalation carries a specific observable description, and the institution actually recorded a verdict. It's a conjunction, not a judgment call. If any input fails, the session doesn't count.

Then guardrails, because a north star is structurally blind to some failures. Student experience is the hard stop; a product can produce immaculate records while students hate it, and the north star won't flinch. Detector coverage guards against the obvious attack on the metric, which is to flag less and close everything. Review integrity guards against rubber-stamping. Cost is a ceiling. Validity is checked annually, through appeal overturns and the integrity incidents that arrive from outside.

And above all of it, a falsifier: a red team of consenting confederates running known cheating methods, a few hundred sessions a quarter. If recall on known behavior falls while the north star rises, the north star is suspended as a steering metric until recall recovers. A falsifier with veto power beats a second north star. Two co-equal top metrics get traded against each other and the organization quotes whichever is up. A veto can't be averaged away.

None of this is specific to proctoring. It's what I'd do anywhere a product can look better on paper while getting worse.