What a benchmark is really a contract for
I've been the person who had to make a vendor's system hold up to an audit. That's why the industry's move to computer-use benchmarks reads, to me, as a procurement problem first.
Richard Rios — Operator & Writer
I spent fourteen years inside regulated financial services — operational risk and collateral controls at a global bank, then client advisory at Morgan Stanley. Now I build AI systems for regulated workflows, and write here about where the hard parts actually are.
What I
work on
Running the workflows where mistakes are expensive and "move fast" is not an option — reconciliations, controls, exceptions, the unglamorous machinery that has to be right.
Designing for the failure case first. Where the real edges are, who owns them, and how a control keeps working after the person who built it has moved on.
Putting language models to work inside processes that already have rules, auditors, and consequences — usefully, and without pretending the hard parts disappeared.
The thinking lives in the Field Notes. Each one is a small attempt to be precise about something I had to figure out in practice.
A studio with no client case studies owes its visitors a different kind of proof. So I ran my own methodology on my own tool — and the tool failed its first review.
I've been the person who had to make a vendor's system hold up to an audit. That's why the industry's move to computer-use benchmarks reads, to me, as a procurement problem first.
Fourteen years in regulated finance taught me that "same input, same output" is sacred. Then language models broke that contract — and I could see where the trouble would land.
Upstream / downstream
This site is where ideas get worked out in public. When they turn into systems built for a specific business, that happens through my practice, Rios Applied AI — the same judgment, applied under contract.