Imagine hiring an assistant who forgets where the filing cabinet is every time they leave the room.
On each visit, they examine the furniture, reason about which drawer might contain the records and carefully work the handle. The demonstration is impressive. You start wondering whether the amnesia is included in the subscription.
A browser agent deserves the same expectation as a colleague: take notes. If it has learned a reliable operation, the next run should benefit. The browser can teach the system how a website works. That lesson should outlive the tab.
Keep the structure when you can
Imagine a weekly export task. Suppose the dashboard receives structured records from its server and renders them into a table. An agent could inspect the table, work through its pages and reconstruct the records. It might also be able to reuse the request that supplied them.
The first route can involve this round trip:
Structured response → rendered interface → extracted records
A verified request can preserve the original structure:
Request with the intended inputs → checked response
Browser automation has more options than looking at screenshots. DOM access, accessibility information and in-page requests can all be useful. The decision belongs at the operation: what does this task need from the browser, and what work is the system repeating merely because it repeated it last time?
If the task is to inspect how a chart renders, the interface is the result. If the task is to retrieve the records behind the chart, rendering may be incidental. An agent builder should know which job they are testing.
A recording is a hypothesis
“Just call the API” skips the discovery work.
A network trace contains requests unrelated to the task. A useful request can contain a date that should be variable, a session value that should stay private or a token that expires. The response may cover only the visible page of results.
A reusable operation has to account for those details. Copying a request that returned the right answer once cannot establish that it will answer another question.
For the export, a new reporting period must reach the server and produce the corresponding records. Repeating the original response would demonstrate memory of an answer, not knowledge of an operation.
Unbrowse's supported learning flow can compare observed sessions and compile reusable requests. Its indexing process verifies browserless replay before adding supported read-only tools. Successful browsing can still finish without a reusable capability; the result must distinguish those outcomes.
The learned route earns its place by working again.
Make the benchmark remember its costs
I would test a browser agent in two stages. This is a proposed evaluation, not a set of benchmark results.
| Stage | What to test | What to count |
|---|---|---|
| Cold discovery | Unfamiliar export; output checked against the source | Time spent discovering the route, including failed learning attempts |
| Repeated execution | Changed reporting period, a period with no records, renewed or unavailable access, and a broken route | Correctness, execution time, recovery work and human checking |
A prepared route and an unfamiliar website start with different information. Calling their timing difference a pure execution improvement hides who paid for the knowledge. Keep both stages visible. A system that learns only the easiest operation should not inherit the result of the harder task it avoided.
Check the result before counting a run. Missing pages, stale dates and records from the wrong account matter more than an elegant request. An empty period should remain distinguishable from an expired session. Otherwise a faster route merely returns the wrong explanation sooner.
Include the person's work. If someone must compare every exported row against the dashboard, that audit belongs in the task's cost. Saving navigation while retaining a full manual check may help, but it is a smaller accomplishment than completing the job unattended.
Failure is part of what was learned
A request may require browser state even after its parameters are understood. That is useful knowledge too. Keep the browser dependency explicit instead of pretending the route became independent.
A later failure needs a similarly precise response. Renewed access, a changed operation and a temporary service error are different situations. The system should follow the supported recovery path or return an understandable blocker. It should not interpret every failure as an invitation to rediscover the whole site, or every empty response as success.
Save the lesson, including its limits
A capability index gives learned operations somewhere to live. The stored knowledge should include the route's inputs and enough context to use it correctly. Private account knowledge stays within its access boundary; eligible public routes can be shared without sharing a user's session.
For a task that depends on the presentation itself, retaining the browser may be the right result. The chart-rendering task should not lose its visual check merely to improve a browserless-execution score. What matters is retaining the useful lesson, including where interaction remains necessary.
The Unbrowse documentation offers a way to try this with a recurring read-only task. Judge the repeated run after changing its input, and keep the failed attempts in view. Do not infer universal support from a successful example.
Give the agent a browser when the work needs one.
Then let it keep its notes.