Answers You Can Act On — Ep 1: Not Every Call Needs Words
Many LLM calls that return a paragraph were really a decision: approve, escalate, decline. This episode shows how to spot those calls, what the new decision models from TypeSafe, OpenAI, Cloudflare, AWS and Perplexity do, and why Red Hat's test keeps a small classifier in the running.
Transcript
A customer writes: "My order arrived damaged. Can I get my money back?" Suppose a large language model reads that, writes a paragraph about it, and your code digs out one word: approve, escalate, or decline. How sure was the model? You don't know. But the only possible answers fit on a card: approve, escalate, decline. So a small model can simply score each one, and pick: say, escalate, at eighty-one percent. No paragraph to wait for, nothing to dig out, and a confidence you can see. So here's the question most of us skip: does this call need a model that writes at all? By the end of this episode, you'll be able to spot the calls that don't, in your own systems, and know what to try first.
Answers You Can Act On: six fixes for today's language models. Part one of six. For the people who build with them, and the people who sign off on what they cost. Six problems you already have with those models, one per episode, each with the fix the industry is moving to: answers that are almost right, so spot the decisions; output your code can't trust, so get answers your code can use; confidence you can't audit, so check the confidence; facts it made up, so decide where a person steps in; inputs that talk it into anything, and a model that gets retired, so outlive your model; and a bill that doesn't add up, so control the bill. Every episode: the problem, how the fix works, what the industry runs today, the evidence and its limits, and the budget, and who signs.
Today, episode one. That ticket travels with us through all six episodes. One word on scope. This course won't teach you to make a language model write better. That's its own discipline. This course is about the calls that never needed to write.
First, the frustration. In Stack Overflow's twenty twenty-five survey, the most-cited frustration with AI tools, named by two out of three developers, was answers that are almost right, but not quite. In the 2026 survey, nearly half said they trust AI output when they can easily verify it. Most of that is about generated code, which this course won't fix. But the same shape of problem, almost right, and slow to check, is what any call gives you when it returns a paragraph and your code wanted a decision. That part, this course fixes. Because an answer picked from a card you wrote is far easier to check, and to measure across thousands of cases, than a paragraph.
Now, how the fix works. Before the call, could you write every acceptable answer on a card? If you could, it's a decision: approve, escalate or decline; yes or no; a score from one to five. The model's only job is to pick from your card. If you couldn't, because it's a reply, a contract summary, or a draft of code, it's generation. The words are the product. Mix both into one prompt, and the decision waits for the writing, on a model big enough to write.
Now apply the test to our ticket. Is this a refund request at all? A decision. Does it match an order we actually shipped? That isn't a job for a model. It's a database lookup. Does the damage claim fit the returns policy? A decision. Approve, escalate or decline? A decision. Then one piece of writing: the reply. Three decisions, one lookup, one paragraph. Send it all to one large model in one prompt, and the decisions wait for the paragraph, and come back as text you have to parse.
Decisions also hide inside the agent patterns you already use. Anthropic's guide to building effective agents describes routing: a step that "classifies an input and directs it to a specialized followup task". Its own examples include refund requests, and sending easy questions to a smaller, cheaper model, Claude Haiku. The same guide describes guardrails, where one model screens what another processes, and evaluator loops, where one call writes and another evaluates. A router, a guardrail, a judge. Each one picks from a short list, even when it's built on a model that writes.
For the decision calls, there's now a new kind of tool: one that doesn't write at all. You give it the situation and your card of answers, and it returns a probability for each answer. TypeSafe's Jev launched in September. Others followed, including four that arrived within about a week of each other: OpenAI's Decisions API, Cloudflare's Clef, the experimental Strands Decider from Amazon Web Services, and Perplexity's decision model. Three of those four publish open weights you can run yourself. Episode two takes the mechanics apart.
So should every decision go to one of these? On the second of October, Red Hat's researchers tested several of them as safety filters, against older tools. On prompt injection, a large language model acting as judge scored highest, at 89.3 percent. A small classifier, running on a laptop, was a third of a point behind, and answered in under a tenth of a second; the large model took about a third of a second, network included. Jev scored eighty-six percent. On content safety, Jev scored highest of everything tested.
Red Hat's conclusion: pre-trained classifiers "remain extremely competitive". And a hope worth sharing: that "the excitement around Jev signifies a shift toward greater pragmatism in model selection". Two things to weigh: the tests were in English only, so for Arabic traffic, run your own test; and Red Hat is making those small classifiers the defaults in its own product. The newest model is one option, not the default.
Your options for a decision call come in the same three columns every episode. Open-weight, which you run yourself: a fine-tuned classifier like the ones Red Hat tested, or an open decision model such as Cloudflare's Clef, the Strands Decider, or Perplexity's. Commercial and hosted: OpenAI's Decisions API, TypeSafe's Jev, or the hosted versions of Clef and of Perplexity's model. And the third column: the model you already have. Route the decision to a smaller model, or constrain its answer to your card. Episode two shows how. Choose on volume, on where your data may go, and on whether you can check the answers.
Last, the budget, and who signs. One number, and one question. The number: how many of your model calls return a paragraph, when your code keeps one word. The question for the board: which of our calls are really decisions, and what does each one cost? You're not alone in asking. Researchers at Nvidia argue that small language models are "sufficiently powerful, inherently more suitable, and necessarily more economical" for many calls inside agents. The benefit: answers that can arrive sooner, cost less, and only ever be one of the options you allowed. Give that list one named owner.
The first pocket rule, one of six you'll collect. Could you write every acceptable answer on a card before the call? If yes, it's a decision, so start simple. A lookup, or a plain rule in your code. Then a small classifier, if one fits the job. If neither does, test a decision model and a large model held to your card, side by side, on your own cases, and keep whichever passes your audit at the lower cost. If no, it's generation. Let the large model write, and check it.
Episode one, in three lines. Decisions pick from a card you could write in advance; generation writes. The new decision models are real options, and so are a lookup, a small classifier, and the model you already run. And the pocket rule: could you write every acceptable answer on a card? Next time, problem two: output your code can't trust, and the fix: the answer arrives as a type. What constrained decoding really guarantees, and where decision models differ. See you there.