When Your Software's User Stops Being Human
Tool-calling and open protocols are turning applications into machine-readable surfaces. For a small software company whose entire product is an interface, that relocates where the value sits.
Commercial software was built on an assumption so basic it is rarely stated: on the other side of the screen there is a person, and that person has to be convinced, guided and retained. The entire discipline of product design descends from it — onboarding, learning curve, engagement, habit-driven retention.
Language models with the ability to call tools break that assumption at one specific point: the operator of the software may no longer be the party that has the goal. The person still has the goal. The operation is performed by an agent. And an agent is not persuaded by a good interface. It is bounded by what your application exposes in machine-readable form, and by the cost of discovering it.
Open protocols for exposing tools to models are the formalisation of that layer: a standard way for a model to discover and use external capabilities without bespoke point-to-point integration. If that approach consolidates — and the signals are strong, though the outcome is not settled — the consequence for small software companies is larger than any argument about which model is best.
What actually changes
The interface stops being the product and becomes one of the product's outputs.
That does not mean interfaces die. It means they stop being the only path, and value shifts toward layers previously treated as plumbing:
- The data model. When an agent operates your system, schema quality stops being an internal matter. Ambiguous field names, implicit states and rules that exist only in the head of whoever wrote the code become execution failures visible to your customer.
- The authorization rules. An agent executes faster and with less hesitation than a person. A permissive rule that never caused a problem because "nobody would actually do that" becomes a probable cause of incident.
- The ability to reverse. Software designed for human use assumes the person reads before confirming. Software operated by an agent needs idempotency, logging and undo — because the reading step may simply not occur.
The predictable strategic error
The natural reflex of a small company facing this is to add a chat box to the product. It is the wrong move and the cheapest one to make, which explains its popularity.
A chat inside your product competes with the assistant the user already runs outside it — and loses, because the outside assistant has context on everything the person does, and yours has context only on your product. The exceptions are real but narrow: cases where the value comes from proprietary data or a regulated workflow the general assistant cannot reach.
The defensible move is the opposite one: make your system the best place for an external agent to act.
In practice that means exposing operations with names that require no interpretation, declared side effects, and explicit rather than silent failure. It is API design work, not conversation design work — and it is unglamorous enough that most competitors will skip it.
What this does not solve
The enthusiastic version of this thesis is everywhere and the sceptical version almost nowhere, so it is worth stating the counter-arguments plainly.
Reliability across long chains. A sequence of operations executed by an agent accumulates error probability at every step. For processes with financial or legal consequence, the required success rate is high enough that human supervision remains the bottleneck — and the bottleneck is the real cost. Any specific reliability figure should come from a dated, published benchmark rather than intuition.
Accountability. When an agent takes a wrong action inside your system using the customer's credential, the question of whose fault it is has no settled answer. This is a contractual and regulatory problem before it is a technical one, and it is moving.
Cost. Agent-driven operation is orders of magnitude more expensive per transaction than a click. That confines the model to processes where the human time saved is worth more than the inference — which is true in fewer cases than the current discourse suggests.
The practical question
For anyone running a small product today, the decision is not whether to adopt agents. It is this: if a competent agent tried to execute your product's core process tomorrow, could it — and would you be able to say what it did?
If the answer is no for the first reason, the work is schema and API design. If it is no for the second, the work is logging and audit. Neither requires adopting new technology, and both leave the product better even if the thesis of this article turns out to be wrong. That is the definition of an asymmetric bet.
References & notes
- Model Context Protocol — open specification and reference documentation.
- No specific multi-step agent reliability figure is quoted here: published benchmarks vary widely by task definition and date, and a single number would misrepresent the spread.
- Liability for actions taken autonomously inside a customer's account remains unsettled in most jurisdictions at the time of writing.
Corrections are published inline and dated. Write to us if something here is wrong.
// weekly dispatch
One email. Every Tuesday.
The week's analysis, one tool we actually tested, and one behavioural pattern worth practising. Unsubscribe in one click.