writingsstate-of-industry / Oct 7, 2026

Who Pilots? -- OpenAI's Decisions API

Back in April I wrote about a library I'd built that records every prompt before it hits the API (/llm-usage-in-a-multi-surface-application), so I could find the calls a plain rule could've handled. In my mind I kept hearing the phrase, "the model call buys us latency, cost, and randomness with nothing to show for it". I still believe it. Yet, what I missed was that it was never really an engineering problem; we put a model in front of a decision because the decision was uncomfortable, & a paragraph of confident-sounding English is a wonderful place for discomfort to go & die. Recently OpenAI dropped their Decisions API (hello jev!). This really is interesting to me, and it isn't the speed (an apparent 150 ms against 1.6 s for a regular Luna call) or the cost point; but it's that the answer stops being prose. You write a question & the "allowed" answers, hand it some context, & get back a pick & (if the launch-day reporting holds, since OpenAI hasn't published the response shape) a confidence score next to it. Because we've seen what "more structured" responses can get us with these decision-focused matrices, I thought this was a great paradigm shift in the industry. However, I then realized we lost the part of the process where we can see when the model decided instead of just getting back a number. Now, if it comes back at 0.94, fine, go. If it comes back at 0.71, somebody has to say whether 0.71 is good enough, & that somebody is never the model. You can bury that number three directories deep in a yaml file, but I think it's a statement about how wrong you're willing to be & who eats it when you are. It's not really a tuning knob but rather it's a policy with your name stamped at the bottom as the responsible party. The prose used to absorb that for us So who's actually deciding? You are. You always were :0 Funny enough, my April piece already had a slot for this. I said some tasks should "defer, abstain, or return a constrained response instead of improvising", & a threshold is just that line drawn in code: above it > act, below it > defer to a human / a rule / nothing. In this, changing it should have to pass evals the same way a prompt change does (most people think a float in a config is the least dramatic diff in the repo, and honestly that's the problem). The part I can't get past is that nobody has published whether any of it is calibrated. When it says 0.8, does the thing happen eight times out of ten? From the coverage I've read, the advice is to treat the scores as indicative, not true probabilities, UNTIL calibration data shows up. Jev (the "pioneer" in this "new field of model decision making") doesn't escape it either; someone on Hacker News noticed its probabilities shift when you just reorder the answer list. Checking isn't even hard which is the annoying part: log the score, log what actually happened, & see if the 0.8s came true about 80% of the time (that's the April library with one more column). I suspect most people won't check. I suspect the number gets used anyway, because a number feels like it has already done the thinking for you; that was ALWAYS the appeal, it just used to be harder to see. The prose is gone now, & for the first time in a while it's obvious who's actually piloting the decision matrix -- hint, it's not the Decisions API.