Imagine hiring an assistant who gives you the wrong answer, thanks you for catching it, and bills you for the thank-you. The correction also costs money. So does the explanation of how the error occurred, which may itself require correction. Somewhere, a management consultant is watching this business model with the wounded expression of a man whose family recipe has been stolen.
We used to approach language models with the elaborate precautions of people summoning something in a basement. “You are a world-class expert,” we began, before supplying a role, a task, six constraints, three examples, and an emotional backstory. Asking for a packing list could involve more preparation than the trip. Prompting became a minor clerical profession, complete with people selling the correct incantations.
Now Grok 4.7, GPT-6 Astra, and Claude Fable 5.1 arrive with promises of greater competence at sustained, complicated work. The hopeful implication is that we can retire some of the ceremony. Clear instructions still matter. But a system sold as an intellectual colleague ought to survive a request written by someone who has other things to do that morning.
This makes its mistakes more irritating. A chatbot that can barely follow a recipe inspires indulgence. One presented as capable of helping with scientific research creates different expectations when it invents a source. The embarrassment has acquired a premium tier.
As the prose improves, the remaining mistakes can also become harder to spot. A visibly confused assistant invites scrutiny. A lucid one may borrow the credibility of its nine correct answers for the tenth. The customer supplies the quality control and pays retail.
The explanation for the bill is straightforward. Metered APIs generally charge for tokens processed and generated: units of text, rather than units of truth. OpenAI also bills internal reasoning tokens as output, including reasoning the customer cannot see. Its documentation explains this arrangement with admirable composure. Subscription plans package usage differently, but the underlying distinction remains: access to computation comes without a guarantee that every answer will be correct.
There is a respectable case for this. An unsuccessful calculation still uses hardware and electricity. Human professionals charge for investigations that end inconclusively. Yet a meter explains how a bill was calculated; it cannot, by itself, establish that the service deserved the money. A taxi consumes fuel while taking you to the wrong airport. You may nevertheless wish to discuss the fare.
The discussion becomes difficult at “wrong.” Some errors are satisfyingly definite. A quotation was fabricated. A sum does not add up. The instruction said “leave this file alone,” and the file has enjoyed a thorough renovation. These failures can be checked against evidence or an explicit requirement. They do not become the customer’s fault because the customer neglected to append “and please observe reality.”
Other disappointments resist inspection. “Make this warmer” may produce a business letter that sounds as though it wants to borrow your cottage. “Make it concise” may remove the one qualification that kept the sentence honest. Here the system has made a judgment about an underspecified preference. Sometimes the judgment is poor. Sometimes we discover our preference only upon seeing its opposite. This is familiar territory for anyone who has selected paint with another person.
Then there is the result that satisfies the instruction and appalls its author. Suppose an assistant with access to your calendar clears your Friday by cancelling everything. You meant the tiresome meetings. It included lunch with your daughter. Asked to find the cheapest flight, it selects an itinerary requiring a night on an airport floor. The price is correct. Your spine has objections.
Such examples reveal how much ordinary language leaves to judgment. We routinely expect people to infer what we value, notice consequences, and ask before doing something difficult to undo. More capable systems may need less coaching through individual steps while requiring clearer limits on what they may decide. “Handle it” contains a remarkable amount of unnegotiated authority.
Calling these problems “stupid questions” is convenient for the seller. A person often asks because they lack the expertise to formulate a better question. A capable assistant should challenge a false premise and identify a consequential ambiguity. The human, meanwhile, cannot delegate judgment entirely and then insist that every unwelcome consequence was unforeseeable. Responsibility belongs partly in the request, partly in the model, and partly in the software deciding which buttons it can press.
Refunds for mistakes sound attractive until someone must adjudicate whether a poem was insufficiently haunting. Still, difficulty at the edges should not erase the center. Fabricated evidence and violations of explicit instructions are reasonable candidates for credits. Providers could distinguish those failures from changes of mind, make spending limits easier to set, and require confirmation before consequential actions. None of this requires a universal theory of disappointment.
For buyers, the useful price is the cost of reaching a result they can trust, including retries and the afternoon spent checking it. A cheaper answer that requires an unpaid auditor may be expensive. A dearer model that gets there reliably may earn its fee. The comparison should include the human still sitting at the desk.
We can reasonably hope to spend less time teaching machines how to understand us. We will still have to decide what we want, how to recognize success, and who pays when success fails to appear. Until that last question is settled, the most commercially sophisticated sentence in artificial intelligence may remain: “You’re absolutely right. Let me try again.”




No comments yet