Agents Need an Index They Can Act On

· updated · Lewis Tham

Imagine asking an agent to check an order in a supplier portal. It finds the portal's homepage.

Fine. We have located the building.

The agent still needs the operation that accepts an order reference and returns its current status. It needs the right account. It needs to know whether an empty answer means “no such order” or “you have been signed out.”

A page index helped with the address. An index of callable operations could help with the work that follows.

That is the case for a capability index. Its useful unit is an operation the caller can understand and verify.

A directory is the easy part

Assume the portal offers search, order details and a shipment view. A directory could list a route for each. The agent would still have to decide which route answers the user's question.

“Status” might mean whether the supplier accepted the order, whether it shipped or whether a carrier delivered it. Sending the same words to the nearest matching tool can produce a plausible answer to the wrong question.

An entry therefore needs to describe the operation precisely enough to choose it. Its inputs must distinguish an order reference from a shipment reference. Its output description should say what kind of status is being returned. If a route only covers open orders, that limit belongs beside the route, before the caller selects it.

This is the index I think agents need. A semantic search over endpoint names is a starting point, but selection depends on the behavior behind the name.

Evidence belongs next to the description

A useful entry should carry a reason to believe its description.

For the order lookup, that could include evidence that the requested reference appears in the result, that the correct account was used and that a different reference produced the corresponding order. Those checks connect the route's promise to its execution.

They also expose a common trap in learning from traffic. A successful request may contain a fixed identifier that should have become an input. Replay can return a valid order forever while answering every question incorrectly.

An index that records only “last request succeeded” misses the point. The caller needs to know what success meant and what assumptions were tested.

Here is a proposed entry for the order lookup, not a schema Unbrowse claims to implement in full:

Entry detail What the caller needs to know
Input Supplier order number, not the buyer's purchase-order reference
Meaning Supplier acceptance; shipment and carrier delivery are separate operations
Access Caller's supplier account
Evidence Which inputs were checked, when, and how the returned order was matched
Failure No such order differs from an expired session

The entry makes a claim that execution can contradict. That is more useful than a confident description attached to an endpoint name.

Without that evidence, the index can make an agent more confident without making it more capable.

Access belongs to the caller

An order lookup also makes the boundary around shared knowledge easy to see.

Knowing how the portal accepts an order reference does not give another person access to your orders. The route can describe an operation; the session supplies authority to act within an account. Combining those concepts into a single “learned skill” risks hiding where the private data went.

Unbrowse permits eligible read-only, secretless learned routes to be scrubbed and shared, with an owner opt-out. Private sessions, logins and writes are excluded from public tool definitions. Some account-specific knowledge must stay private altogether.

A collective index should spare the next caller from repeating discovery. It should never lend them the previous caller's identity.

Freshness changes the decision

Suppose an order route was tested successfully and the supplier later changes its status response. Perhaps the field still exists but now describes fulfillment rather than acceptance.

A transport check can remain green while the meaning drifts.

This is why I would evaluate freshness against the operation's claim. Does the request still produce the promised kind of result? Under which access conditions? If those checks fail, the entry should stop presenting itself as a dependable answer to that task.

There may be a supported way to repair it. The agent may need renewed access. The operation may no longer be available. Discovery should help distinguish these cases so the caller does not keep trying stale instructions with increasing enthusiasm.

Unbrowse exposes capability discovery and version information, and callers are instructed to inspect returned results. Indexing can learn and verify supported operations. Those are useful building blocks; a recent version label alone is no guarantee about today's answer.

Judge the index by the choice it improves

For an infrastructure builder, I would test a capability index with an intentionally ambiguous request. Give it the supplier “status” job and see whether it identifies what the user needs, selects an appropriate operation and returns evidence tied to the right order.

Then present a missing order. Finally, make the required account unavailable. A useful system should give different answers to those situations. Treating all of them as “no result” leaves the next agent, or the person, to repeat the investigation.

That is a proposed evaluation of selection quality. It does not require a huge catalog to be revealing.

The Agentic Web Will Not Wait explains why discovery should not depend entirely on vendor roadmaps. The Second Run Is the Test of Your Browser Agent explains why discoveries should survive the browser session that produced them. The Unbrowse documentation provides a concrete starting point for trying discovery and inspecting what a returned capability actually supports.

An index earns its place when the agent chooses a better next action because the index exists.

Finding the building was never the whole errand.

All posts