Steer Your Agents Until They Steer Themselves

Steer Your Agents Until They Steer Themselves

Rohan Vaidya

Over the last year and a half I have been working side by side with Postman’s CTO, Ankit Sobti, on the agent platform we use to deploy all of our internal agents as we transform the company to be AI-native. This is a follow-up to his post on the Agentic OS we built, which covers the business need that led us here, an architectural overview, and the technical precedent that shaped our thinking. In this article, I will discuss the problem space surrounding building agents for enterprise, in particular for non-verifiable qualitative domains (not coding/math), our approach to solving them, and the open problems that remain. I hope this serves as a useful guide for high-agency engineers or engineering leaders hoping to steer such an effort.

Context

I came to Postman after working on a startup. There is a stereotype that technical founders would rather code their way through problems than talk to customers. Unfortunately, these stereotypes can occasionally be true. With no salespeople or budget, I too chased the dream of a product that grew itself through scrappy LLM-augmented outreach tools, but learned the hard way that no amount of automated go-to-market (GTM) can grow something that doesn’t have a market. In coming to Postman, I actively sought the opposite problem. Postman had already reached 40 million users through product-led growth. So, a primary objective was to just convert an increasing subset of those existing users into paying customers by delivering real enterprise value, and to keep shaping the product around what they told us about their new AI needs.

I helped found our forward deployed engineering (FDE) team, wrote its first playbook and shipped our first major deployments, which led to significant million-dollar expansions in annual recurring revenue. Although we were able to get to a great outcome with these individual customers, the process behind it could not scale. We relied on our executives’ intuition to pick the right accounts. We could only read so much of the calls, emails, product telemetry and support tickets on each one, and even what we did read missed context that lived only in the account team’s heads, outside our systems of record. Filling the gaps meant more discovery on the customer’s schedule, and nothing we learned reached the sellers, so even a problem we had already solved came back to us.

What we needed was a GTM machine that relentlessly compounds. Everything we know about a customer would live in our systems of record, and the machine would read all of it and tell our people which accounts to approach, with what, and what to ask the core product team or FDE to change in our product. Anything our team had built once would become something sellers could sell on their own, which would leave FDE free for the last mile of novel engineering. Every time our people acted on what the machine told them, the result would become new data, so its next recommendation would be better informed. It also had to be cheap enough to run on every account every week, and keep working as the models and the data underneath it changed.

So I tried automated go-to-market again. This time I was armed with the luxury of a product that had already found its market. My startup-background-induced naiveté about accelerating a field that has always run on intuition and relationships at enterprise scale was shielded by our developer tool founders, who shared my rose-tinted view that we should be AI-pilled and apply agentic thinking outside software. Having learned my lesson, I wanted the best of both people and AI, and that came down to a maxim.

Abstract

Human at every surface, machine on the analysis behind it

Fig. 1: A ring of people forms the only boundary a customer touches, and inside it a dense green network of analysis.

Maximally human on external communication. Maximally AI on internal analysis. That should be the self-improvement loop of any enterprise.

Let’s start with the external half. As much as possible, every customer should feel that their problem was heard by a person who cares, and that the solution was built for their use case. Our largest accounts have a median of 45 stakeholders named on their deals, and the median relationship has run about three years, so keeping those relationships strong is one of the best uses of our people’s time. AI can handle the quick things, like a support answer at midnight or an automated onboarding flow. But an AI avatar on a sales call, or a bulk-generated cold email that blindly guesses at a customer’s problems, cheapens the brand, and that is not what being AI-pilled means.

The internal half is a great place to be AI-pilled. You want decisions made at a velocity, throughput, and accuracy no individual can mentally accomplish, and the reason is that it is biologically impossible. Research on working memory finds that we can actively juggle only about four chunks (meaningfully grouped pieces of information) at once (Cowan, 2001), where in our case 1 call would roughly be 1 chunk. Even at one large account for one year, there can be hundreds of calls and thousands of emails. To make matters worse, the highest leverage insights are often thematic across accounts and it can get up to {calls, emails, telemetry, support, docs, etc.} × years × accounts. This is terabytes of data that no person can hold while trying to draw the highest leverage line of reasoning across all of these data points.

Accounts across, years down, one route through

Fig. 2: Each cell is one account’s year of records. The green path links the data points the agent reasons across on its way to an answer, and the small box shows how little of it one person can hold at once.

Take a seller who wants to reach out to each of the customers in their book of accounts that would benefit from something we launched recently. Doing it well means reading up on each account, and with a couple hundred accounts each there is not enough time in the week. Timing matters here, because the frontier moves quickly and features go stale. An agent can go through every account for them, pull out what it finds that matters for each one, and bring the strategies our best sellers have already succeeded with to all of them. The seller can then write to every customer and tailor the message to the right stakeholder with exactly how it would benefit them.

As we grow more confident in the strategies our best sellers use, we can hand more of that work to the agent to run autonomously. The questions sellers ask over and over become skills, the skills they run every week become scheduled reports, and eventually the agent keeps its own memory of each account, asking the questions it deems important to meeting our KPIs.

To climb toward an agent that does more on its own, we can take a page from how software handled layers of abstraction. Engineers once wrote machine instructions by hand, until an assembler took over that translation. Compilers then let them write in languages people could read, like C, and object-oriented languages like Java let them bundle data with the code that works on it, so they could build large systems out of reusable parts. This ladder worked because every rung had both generation and verification. Generation meant a tool could produce the work at the new level. Verification meant engineers could check that it was right. A compiler gives the same output every time it gets the same code and refuses to build code with whole classes of mistakes, so engineers could trust each rung and build the next one on top of it.

Qualitative domains had neither automated generation nor verification before AI. Enterprise GTM is one such qualitative domain. The answers to questions like “What is the optimal strategy to sell customer X on product feature Y?” are open-ended and have no true right answer. Previously, you could not automatically generate an answer for the next move on an account without a human Account Executive in the loop. Now, an agent with tools generates an answer based on the subset of data it finds during its execution time. However, it is still hard to know whether the agent reached the optimal subset given the global corpus. Did it find that one call where they discussed their new AI strategy? Did it look at the one data point in the product telemetry where it shows they actually were leveraging your new feature but in an unconventional way? Verification remains an open problem. We have not solved it, but every approach that has worked for us starts the same way, by breaking it into smaller checkable pieces.

Generation lengthens the ladder, verification makes a rung hold

Fig. 3: Each braced software rung pairs with a check, from the hardware up to tests, while the rungs above the line, like the next move on an account, stop short unchecked.

I decided the generation half was promising enough to build on, and started with a GTM agent that answers questions about our customers from the thousands of calls we have recorded with them. Like any conventional agent, it began as a model, a system prompt defining its persona, and tools that return data, all running in a harness that accepts a question and loops until it has an answer. From there it grew into an accelerated agent deployment platform that now enables agents to be quickly spun up for any end user across the company. Since late May, nearly 400 employees have used them, from sellers and executives to product managers, marketers and engineers, through Slack, the web, MCP and the command line, both for questions they ask and for proactive reports and alerts that come to them. The coaching agents alone reach nearly all of our sellers, about 220 people a month, and have been asked almost 10,000 questions. Monthly questions grew 76% from July to September, with 80% retention, and a strong power law in usage with 1 in 10 having asked more than 100.

Spinning agents up quickly and serving answers turned out to be the easy part of that platform. Aside from verification to ensure we were improving the agent over time, a significant hurdle was getting the right people working on the parts of the agent they were best suited to. Every team has its own goals, incentives, definition of a good answer and view on who should see which data, and each is right about its own, so we were never going to write their agents for them from the outside. What makes the platform work is that each team writes only the part it understands, and everything else lives in a shared substrate we maintain for them. The agent becomes a configuration of a shared thing instead of a new thing, so every improvement to the substrate lands on every agent standing on it. That is what lets the GTM machine compound.

One substrate, molded per team

Fig. 4: Four team agents differ in shape and in their prompt, tools, data, guardrails and access settings, yet all stand on one shared substrate that improves beneath them.

1. Everybody Builds Their Own Agent

Agent building inside enterprises has reached a fervor we have not seen before, because anyone who can write instructions and connect a data source can now build an agent. Even within a single team, several people build agents that share a function, end users, and data sources. Nobody wants their agent to just sit on their laptop, so they rush to serve it to the others. Unfortunately, everyone else has the same idea and serves theirs back. Every one of these efforts stands up its own pipeline over the same data, so each new agent leaves another copy of your customer records somewhere with less oversight than the last, and those pipelines go ungoverned the moment access expires or a dependency moves, which is when the agent starts answering confidently out of data that stopped updating weeks ago. There are two common attempts at solving this, and both appear to work at first.

The first way is that the domain team or individual builds it themselves. Our sellers did, through Claude Desktop, wiring up connectors by hand. That gets you running in an afternoon, and the person doing it understands their own workflow better than any platform team will. What you cannot do is scale it to an entire team and globally steer it. Outcomes varied wildly from one seller to the next, because each seller wrote different instructions and connected a different set of sources, so two sellers asking about the same account could get answers drawn from dramatically different data. When an executive wanted every agent working the new sales play, there was nowhere to put the play. It lived in a deck and the agents lived in individual sessions, so none of them ever saw it, and the machine we set out to build depends on exactly that kind of shared context. The same gap shows up in how they read data at face value. If the customer said the product was too expensive, the agent told the seller to offer a discount. That is rarely the right answer, since any customer of any product in any industry could ask for a discount, whether or not the product serves them. What the seller actually needs might be a nudge to re-qualify the champion, or to re-anchor on the outcome the customer is buying. Those are strategies our senior sellers and executives have worked out, and a personal agent has no way to receive them.

Every seller's agent, wired by hand

Fig. 5: Six sellers have each wired a different mix of sources into their agent, and the new play sits in a deck none of the agents can reach.

The second way is to hand the work to a forward deployed or applied AI team. These are skilled software and systems engineers who have learned to work with AI. Put them together as a small team and they can handcraft one agent per internal use case. That solves steering, because those engineers own the single copy of each agent deployed to a team, so one change reaches everyone using it. What it does not solve is anything after the first version. Each agent is bespoke, so there is nobody to hand it back to, and it stays with the people who built it. Every one shipped becomes something somebody keeps alive, and making the same patch to each agent they shipped is not the highest leverage use of a capable engineer’s time.

One play, added to every agent by hand

Fig. 6: Engineers add the same new play to four bespoke agents one at a time, and the fourth, negotiation prep, is still waiting for its copy.

What became obvious once a few of these sat side by side is that they were variations of the same bounded list. An agent is a system prompt, a subset of the tools and data sources you have already provisioned, a model and a budget for how long it is allowed to think, an execution horizon that says whether it answers when asked or runs on a schedule, a set of guardrails, and a rule about who may invoke it and which fields it is allowed to read. And this list can always be expanded as you uncover new primitives you want to bring to all of your agents.

What we need is for the platform to be configurable so that each of the relevant parties in the company can write the component of the agent that pertains to them and leave the rest to the others. Access control, the hard deterministic rules governing who can reach which agent and which data fields, is best controlled by IT. The system prompt is best written by the domain team, since they know their own workflow and what they want the agent to do on their behalf. The guardrails, the soft steering markdowns that keep the agent away from the poor behavior you only learn about by deploying it, are best written by the manager. Instead of one of these teams writing every component, or an outside engineering team guessing at all of them, each component gets decided by the people closest to it. The agent platform team can then just work on improving the overall substrate.

The ideal construction

Fig. 7: Each bear built by a single team fails in its own way, and the combined bear takes access control, prompt, runtime and guardrails from the team closest to each.

2. Deploying Every Agent from a Monorepo

In order to have catch-alls for functionality the platform did not support out of the box, we let agent developers write custom code on top of the agent they were serving. That code turned out to be the signal we needed. When our engineers found themselves writing patches with the same shape for the third or fourth time, the shape got promoted: it became a field any agent could declare in its manifest, with no code at all.

Every agent is written to a single monorepo that deploys when a PR merges. An agent is a directory. Everything specific to that agent lives inside it. Two folders sit outside that everything shares: the globally provisioned tool set, which each manifest names a subset of, and the guardrails, which name the agents they apply to. The appendix shows an example of each of these pieces.

What got promoted

Fig. 8: Apart from its SKILL.md, every setting on deal-coach-pro is a choice from a fixed set, including 52 of the 54 tools and 32 of the 45 guardrails.

When it comes time for deployment, we wanted a mechanism that enabled agent developers to quickly deploy their agent without taking down a shared substrate. Astro, our agent deployment and governance platform, made this very convenient. The shape of it is ordinary Docker. Every agent builds from the same Dockerfile with the same build context, the whole backend, so COPY . . bakes every skill into the image. Each agent then carries a manifest, a short YAML file that names that Dockerfile and declares the environment variables its container needs. One of those variables is the name of the skill directory to serve. The adapter reads it at boot and serves only that one, which means the codebase is the artifact and the variable is the pointer: every pod runs the same tree with a different path lit up inside it.

One agent's path through everything available

Fig. 9: SKILL_NAME lights up the deal-coach/ directory and the shared files, tools and guardrails it names, and every gray entry ships in the same image without being loaded at runtime.

3. Getting Long-Horizon Results at Runtime Speed

One thing held across all of our agents: the longer we let a run go, the more tokens we let it burn, and the more data we gave it access to, the better it performed. This is perhaps a microcosm of Richard Sutton’s Bitter Lesson, which holds that massive computation dwarfs last-mile algorithmic iteration. The constraint with building runtime agents like the above is that users expect an answer inside a reasonable window, about two minutes.

What we tuned against what we spent

Fig. 10: Inside the two-minute window hand tuning gives the better answer, but past the crossover more tokens and longer runs keep improving it long after tuning has flattened.

You can give them the option to configure a job that runs much longer, ten minutes or more, and we did, but that carries its own set of problems. One, if everyone starts launching computationally intensive jobs at will, token spend goes out of control. Two, you stop your users from getting deeply nested in a workflow and reaching the higher order insight. When responses come back quickly they keep going down the rabbit hole, and your overall usage of the tool grows exponentially. This can also be read through Jevons paradox, which says that the cheaper you make something the more of it people use, except that inside an enterprise the thing your employees are trying to conserve is often not cost but their own time. So if they believe it is saving them real time, and more of it is at their fingertips, they will use it more, which raises their overall AI consumption and drives your transformation.

We needed to find a way of providing the benefits of longer running jobs while making the runtime inference fast and cheap. We need to find a way for an AI agent to figure out which pieces of information are the highest leverage for an account and refresh those as new information comes in every week. “Highest leverage” is a vague term that is specific to your company’s needs so to filter for this we would need another probabilistic agent steered by prompts to filter the high volume of incoming information. I will discuss two approaches I took here: a harness that allows an agent to recursively in parallel spawn sub agents so that it can better research an account, and the second being Anthropic’s first party dreaming primitive. The former was hand-rolled before the release of the latter but both have since proven to have their benefits.

The sources are given, the filter is the product

Fig. 11: Streams of data arrive all week, and one agent, set by its model, prompt, tools and token budget, keeps the few items that fill the account document.

The first in-house method is what we call the research pipeline and it runs once a week for 8 hours. Each account research process begins with a deterministic set of API calls that pulls context from each of our data APIs which took place after last week’s run into a shared store. This includes the Salesforce record, its opportunities, its account plan and its contact roles, along with transcripts from every Zoom call, product telemetry and support tickets. From here, in batches of 4, 24 sub agents are spawned that use this store as a seed to decide what threads they would like to further investigate. What they are looking for is fixed by the sub agent in advance and it is informed by what they see in the seed, what agents from previous batches already returned, and core pillars the account team cares about (renewal risk, expansion path, stakeholder map, implementation blockers, recent customer sentiment). Each of these sub agents reasons over the tool calls it is making until they find what they were looking for or the token budget runs out.

The path each one takes through the data closely resembles a graph traversal, because at every turn the agent can move vertically, deeper into the same source, or horizontally, across to a different one. For example, it might be reading a chunk of a Zoom call where the customer says a feature no longer works for them, and go vertical, back through earlier calls to see whether they turned it off, or horizontal, out to the product telemetry to see whether they are using it at all. It ends up making a combination of horizontal and vertical turns, which resembles an A* search, except that the heuristic is the model’s probabilistic evaluation of whether a tool call will yield alpha, whereas A* would use a deterministic function. At the end, every agent returns its findings to a more intelligent, large context model that synthesizes the data points and writes a detailed account dossier. This dossier is stored in a relational database and also embedded into a vector database. It is available as a tool at runtime and the agent is told to prefer it in the tool definition and system prompt since if it contains the answer to the question being asked then further tokens don’t need to be burned querying the raw data source.

Twenty-four threads, and two directions each

Fig. 12: Researching sub agents walk the seeded sources, stepping deeper or across at each turn, and their findings feed one synthesis that writes the dossier.

While the research pipeline is our best attempt at giving agents the highest leverage insights at runtime, Anthropic released Dreaming as a primitive for curating an agent’s memory directly. Rolling your own keeps the cultivated memory model-agnostic. Using a black-box dreaming API makes sense for the opposite reason. The model provider is the one most likely to know what harness curates information so that their own model recalls it the way you want at runtime.

Given that, I set up a second pipeline in parallel, on Dreams. To understand Anthropic’s Dreaming you have to start with their managed Memory Stores.

Memory stores exist to solve the ephemeral nature of agent sessions. An agent makes its tool calls, reasons over them, the container goes away, and everything it worked out is lost. To avoid that you need a persistence layer, and the obvious ones are the databases you already run. That is what we did for the research pipeline: the dossiers go into SQL and the transcripts behind them into a vector index. The problem with relational and embedding-based retrieval is that the agent is only ever shown the matching subset. If you want Acme you can run SELECT * FROM dossiers WHERE account = ‘acme’. If you want to know where onboarding is going badly across the book you can embed “onboarding issues” and take the nearest neighbors. Both work, and both are narrow by design: the agent reasons over the matches and never learns what else was there.

What a query reaches

Fig. 13: The SQL query returns one Acme row and the embedding search four onboarding chunks, but the security review chunk holding the answer falls outside both.

Now ask why Acme’s rollout has stalled. Both queries fire and both come back full. The dossier has the onboarding tickets, the training sessions, the adoption curve flattening in March. The semantic search has a dozen accounts with the same complaints and the themes they share. There is plenty to reason over, and the agent writes a confident answer about enablement. The actual reason might be a sentence their security lead said on a call fourteen months ago: they cannot grant the directory access the rollout depends on until a review that never got scheduled. That sentence is not semantically close to “onboarding issues”, it sits outside any sensible recency window, and it was filed against a stakeholder, nowhere near the rollout. Nothing in either result set tells the agent it is missing. An agent that can see the whole corpus and navigate through it knows what it is choosing to leave out.

A memory store enables the agent to see the entire corpus. It is a collection of small text files that lives above any session, with an id of its own, and every file has a path, so a store has the shape of a directory. You attach one when you create a session, read-only or read-write, and the platform mounts it inside the container. Writes the agent makes there persist back when the session ends, and each one leaves an immutable version behind recording what changed and who changed it.

There are no memory tools. No search endpoint, no embeddings, no index. The agent runs ls and gets back every path in the store, not the subset that matched something, and it sees those paths whether or not it would have thought to ask for them. From there it navigates: grep for a term, cat a file, and decide from what it just read what to open next, using the same tools it is already well trained on from navigating code. In addition, since we don’t need to respond in a highly constrained time horizon like the runtime agent, we can afford to let it explore this data to its satisfaction.

What an agent reaches

Fig. 14: The agent lists all files, greps for rollout down to three, then reads timeline.md, which points it to security-lead.md, a file the grep never matched.

That brings us to how we used these primitives to build our account memory. The first decision is what an account has in it. We settled on one store per account: overview.md at the root, a file per person under stakeholders/, a file per opportunity under deals/, what customers actually said under quotes/, and a timeline.md of recent activity. Every file carries a provenance header naming where it came from, who asserted it and when. We seed a store from the root data sources and the output of the research pipeline, which makes every file in it a synthetic artifact we need a steered Claude to keep current. New material arrives through ingestion jobs that listen to those same sources and write it into the filesystem as a record of recent updates. This is fine for new data from the week, but if we keep piling these in, then our overall system degrades to just a store of the raw data sources. We need a weekly process to compact these updates into our high-level synthetic artifacts.

This is where dreaming comes in. A dream takes the account’s store, a model, a steering prompt called instructions, transcripts of the sessions that read it, and writes a new store. In our case those sessions are the runtime agent’s own inference, each one a different traversal of the same data. Those traversals are what let the pass decide how the data should be structured and what is worth keeping, based on the shape of the questions actually being asked. It makes the memory layer self-healing: gaps get identified, what nobody reads gets discarded, and the answers people keep walking to end up where they will be found first. A SQL or vector store will serve any model that queries it, but its schema is fixed the day you design it, and no amount of being used will change its shape.

The hard part of a memory layer is that it has to stay durable without inventing, because it is no longer surfacing raw data but syntheses of it, massaged over time. Two protections do that work. The first is a rule: the pass may not add anything that was not in its input. The second is structural: nothing it writes counts as memory until a pointer says so. A store’s id changes every pass, so the runtime never holds one. It looks up a slug and gets back whichever store is current, and that lookup is the only thing that makes a store’s data appear in the output. So a new store sits there, referenced by nothing, while it is checked: file count within forty percent of the input, nothing oversized, every view present, provenance valid throughout. If it passes, the lookup is repointed and the old id goes into history. If it fails, the lookup does not move, and last week’s store keeps answering.

One week of an account's memory

Fig. 15: Raw updates pile up all week until the weekly pass, steered by session transcripts, writes a candidate store that the slug points to only after it clears all checks.

4. The Agentic Mountain Range: Measured Agent Improvement

In order to best describe what agent building feels like, I would like you to picture yourself as a determined traveler in the foothills of a majestic but intimidating mountain range. Ranges like the Himalayas, Andes, or Zagros were so intimidating that they quite literally divided civilizations. You, on the other hand, prepare to chart your path across. As you look up at the steep slopes, you ask yourself what the best strategy is. The reality is that you are not trying to summit anything. You just want to get to the other side.

Which makes every foot of elevation a cost you would avoid if you could, and shifts a lot of the burden onto choosing: which mountains stand between you and the other side, which of them you need to cross, and in what order you take them. A party that takes the tallest peak first and dies on it has crossed nothing. You save the taller ones for later, when one of them is the only path left.

The mountain range in front of an agent team has several mountains, and those are only the ones you can see. More appear once you have covered enough ground in your own domain. A few are universal. Delivery: getting a good answer out of the system at all. Evaluation: knowing whether the answer was any good when it is hard to verify. Self-improvement: a system that suggests changes and improves itself. Coordination: an agent that knows your organization’s roles and can message the right person to do the right thing at the right time. Action: an agent that knows how to do the thing responsibly instead of only recommending it.

The range

Fig. 16: A map charting out challenging problems in agent building as mountains and a logical path of increasing resistance only when necessary.

In deterministic software, choosing the right problem carried less weight, because a bad choice announced itself. You attempt the hard problem, you fail, and the failure is checkable, so you find out early and correct course. In the probabilistic world a team can spend a year on a hard problem and whether they got anywhere at all is just their word. Did you really build a recursively self-improving agent? I guess these responses look better than before. Good job, keep going. Or maybe you did build it, and the problem runs the other way: the outputs all look alike, text and quotations, and nobody can say whether they are better than last month or worth more. So how do you make your progress clear?

In the deterministic era you knew concretely how much a change you were about to push would improve the overall system. Now that LLMs have made it probabilistic and the outputs qualitative, it is very difficult to know the impact of any given change. For this reason, software engineering in the probabilistic era has moved toward outcome engineering. This is the process of defining what you would like the outputs to look like for your coding harness, letting it make the technical decisions, scanning what comes back against your intuition for how it should respond, deploying, and then refining the system over time where it does not behave as you expect. As a result, the incentive to review every incremental change with the paranoia you once would have had is gone.

The prose diff

Fig. 17: Rewriting one instruction turns a discount plan into advice to re-anchor on the integration timeline, and nothing in either answer shows which is better.

Improving the system breaks down into climbs of different sorts. The two classic ones in computing are hill climbing and climbing the layers of abstraction. The first concentrates on one output metric, makes a change, and keeps the change if the metric moved. The second starts a system at its simplest understood form, watches it run in production, works out what you are willing to accept as fact, and then builds on top of that so the next decision gets made at a higher level. Coming back to the climb, a mountaineer drives a bolt into the face, clips the rope to it, and runs the next stretch from there. The incremental changes are the holds (or hills) you climb between one bolt and the next, and each bolt is a check you can take as fact before you build the next layer on it. What the bolt buys is that a mistake above it costs you a rope length instead of death.

Two kinds of climb

Fig. 18: Hill climbing’s scored gains shrink from +6 to +1 before the metric stalls at a peak, and the abstraction climb stacks from chat upward with a bolt at each boundary.

Hill Climbing for Agents

Hill climbing is a classic optimization technique where you look at the states adjacent to the one you are in, step to a better one, and repeat until nothing next to you is better. This means you stop on the first peak you reach, a local maximum. There are well-understood escapes from this, like sometimes accepting a worse step on purpose (see simulated annealing for arranging cells on a chip). Regardless, observe what is fundamental to the climb and the escape: a score, computed by a machine, for any given state. Coding and math are domains where this holds, and it is why the frontier labs have taken their models so far on both. Roughly, they collect a large bank of problems with strict test cases (like SWE-bench, which is real GitHub issues and the pull requests that closed them), change the model, and see whether it passes more of them. Nobody has to read the output to know, so the loop runs over terabytes, at a scale no amount of human review could keep up with.

Hill climbing for coding

Fig. 19: Test files score every answer, so each kept model passes more of the 7,200 clones, from 771 at v1 to 5,663 at v7, and the neighbors v7a and v7b score lower.

None of that holds for an agent whose output is prose about an account. Even if we keep the construction of the agent, the harness, identical, there are still infinitely many adjacent states because many of the surfaces you steer with are text. A system prompt, a markdown guardrail, a tool definition: each can be changed to any other text, so there is no bounded list of neighbors to walk. The tool definitions alone carry more than eight thousand words of English, and a good part of that is instruction written to correct something a model kept doing. The surfaces that are bounded, turn ceilings and token ceilings, multiply what is left by another order of magnitude. And then the knobs everyone reaches for first: which model you run, and which subset of the tools you hand it. Once you have landed on a configuration, what you need is a way to say whether the answer it produced was any good.

Before you can get a machine to say it, you try to codify the judgment with a person in the loop. What the work actually looks like is this. A question comes back answered badly. You change one of the surfaces above and run it again. Sometimes it is clearly better. More often it is just different, and you decide. What is the true right answer to “What is the optimal API governance strategy at Acme?” And how would you know whether the agent picked the right set of data points out of everything it could have read? So you might keep the change, you might not, and you go to the next thing. Whether it was a good change could depend on who reviewed it, how many times they ran the case to see whether it was flaky, and whether what they wrote was a band-aid or a fix. Do that for a year and it builds a sediment, a whack-a-mole game of changes made to correct behavior nobody could attribute.

So you do the obvious thing and try to do it systematically, at higher volume, by handing the scoring to a model. You take two versions of the agent with one concrete difference between them, a model change or a code change, and you pass both answers to the same question to a judge. The problem is that a judge with no tools of its own has no more idea which answer is better than an uninformed third party would. It never saw the corpus, so it cannot tell whether either answer reached the part that mattered. What it can see is the prose, so the prose is what it grades: which one reads better, which one moves between paragraphs more smoothly, which one cites more. Give the judge tools and you land in a catch-22: a judge that can go and find the best subset of the data is the agent you were trying to build, and if you had it there would be nothing left to compare.

We can use a historical analogy to understand what to do here. In the early 1700s, Antonio Stradivari, an Italian violin craftsman in Cremona, made what are considered to be the finest violins ever built, and in the three hundred years since nobody has come close. Later in life, it became apparent to Stradivari that he needed to distill his craftsmanship to others or the craft would die with him. He started with just his sons Francesco and Omobono but expanded this to other apprentices in the town to create a workshop. Stradivari was obsessive over his work and did not want to tarnish his reputation, so he needed a process for strict quality control. Like in our agent building case, the output metric of violin production is qualitative and subjective: its sound. He could not sit and listen to every single violin out the door, and the sound of the violin was only possible far too late in the process after it was carved, closed, varnished, and strung.

Nor could he hand the listening to his apprentices. Asking an apprentice whether their violin sounded as good as one of his was asking them to judge work finer than their own, with the same judgment that had fallen short in making it. That is the position of a model or LLM-as-judge today. You could have a stronger model grade two configurations of the same agent running on a weaker one. But how do you climb to the best possible agent when it already runs on the best model you have?

So there is no score, and the work goes on without one. What accumulates is a growing pile of patches, each written against a symptom, and each a bet on how one model behaves. The model updates or you change providers and every bet is open again, because the new one fails differently and the corrections you wrote now push against nothing. Worse, a patch made for a reason that has since gone away does not announce itself. Something gets fixed upstream, the symptom disappears, and the paragraph you wrote to lean against it stays in the prompt, leaning. You end up steering with controls that may already be pointing the wrong way, and nothing that would tell you.

Lucky for us, we can take inspiration from Stradivari. He needed other output metrics he could measure the quality of a craftsman by. He developed several: the body shape needed to match his hand-crafted walnut mold, the eyes of the f-holes needed to sit an exact angle apart and were measured by compass, the outline had to match his paper patterns, and the purfling, the inlaid border, had to run a constant distance from the edge. These are all hard checks that anyone in the workshop could conduct with a tool instead of his ear. None of these tools told him whether a violin would sing, but each told him something he believed needed to be true before it could. For an agent, these are verifiable sub-problems: pieces of an answer that can be deterministically checked without anyone judging the whole.

Checking a violin without an ear

Fig. 20: Body shape, outline, f-hole placement and purfling each pass a check against a workshop tool, but the sound has no tool and is left to Stradivari’s ear.

The gut reflex in agent building is one more patch, another small change to the prose, and it leaves you no higher than the last one did. Checking has no ceiling, so every time you break the work into verifiable sub-problems, you earn a way up to a harder problem. Once the agent passes them reliably, you can stand on that layer and climb to the next one. The prose changes in between are just hill climbing to a local maximum.

Hill climbing is the wrong hill to die on, and I will die on that hill.

Climbing the layers of abstraction for agents

In the case of agents it is hard to know in what order to add complexity. There are so many moving parts available to an agent developer: system prompts, context, memory, caching, persistence. Someone building from scratch is tempted to throw all of it at the wall and see what sticks, and what they end up with is an agent nobody can steer, because nobody can tell which part is responsible for anything. The order matters more than the parts.

The order has to be deliberate, and it comes from watching the people who use your agent and deciding where you are willing to give up steering by hand in exchange for leverage. For our agent the path went like this. We started with a chat agent that reached our data sources through tools, and watched which prompts people ran over and over. Those became skills, markdown files passed in at runtime, so people could skip straight to the uses we wanted to promote. Then we watched which skills got run most often, because that told us what workflows people actually had, so we let users add a cron schedule to a skill. We called the result a report, delivered to Slack or an inbox. That made visible the shape of things people want to hear about, how often, and whether they act on it when it lands. It was enough to work out what else would be worth telling them and at what cadence, which is where the proactive agent came from: one with a memory of its own, tracking what matters about each account and what something has to look like before it is worth a DM. The research pipeline and dreaming are what made that memory possible.

Each of those steps needed a bolt before we could trust the one above it. The bolts go in at the transitions between layers: a check you can clip into that tells you the layer below works as you expect, so you can keep climbing. We drove the first one in as we moved past the chat layer. Models at the time still made up numbers often enough to worry us (ARR, renewal dates, license counts), and that gave us a verifiable sub-problem: of the figures in the answer, how many never appeared in anything the tools returned while the agent explored? A high share meant the agent was likely inventing data. The same move carried up the climb. Skills exist because people ask the agent the same question over and over, which makes them a good place to measure flakiness with one number: how many tool calls a run makes. The same question should take about the same amount of digging each time. A skill that makes four calls on one run and forty on the next is not one you want on a schedule, answering for itself. For dreaming, the bolt is the pointer from section three. A new store answers nothing until it passes its checks, and until then last week’s store keeps answering.

Driving the next bolt

Fig. 21: The cliff rises in layers from chat through skills and reports to memory, with the rope clipped to a bolt at each transition and the climber drilling the next one.

Between two bolts the softer components are free to move, the prose in your prompts, skills and tool definitions, and that is where the hill climbing described above belongs. Add rigid layers without knowing why, fail to bolt them down with a check, or keep piling on prose instead, and you inch toward a slop cannon: an agent whose answers and actions nobody can explain. Research makes this easy to see: you cannot let an agent research on its own until you know it is finding the right data as it traverses, weighing relevance against recency the way you would, and that its tools are not failing quietly or returning something unexpected. You only learn that through rigorous synchronous use. Without that understanding, every call the agent makes and every answer it writes leans toward slop, and letting it run on its own compounds it. The effectiveness of these verifiable checks led us to dedicate the next mountain to increasingly comprehensive evals of non-verifiable domains.

5. Evaluating Non-Verifiable Domains

Evals are a fundamental part of agent development because they let you change the agent and be confident about what the change will do. That confidence matters more as the agent’s use grows across the enterprise, because a degraded agent does not look degraded to the people using it. It still answers with plausible facts in fluent prose, but the facts may be lower leverage than they were, and the enterprise can spend months making weaker decisions before anyone notices. Changes to an agent fall roughly into two groups: its model and its harness. A model change usually means upgrading, switching providers or moving to a cheaper model, and in every case the question is the same: what share of the performance do you gain or give up, for what share of the cost? A harness change is what your developers make as they watch the agent perform, usually adding a new layer of abstraction or hill climbing the prose on the existing ones, and the question is whether it made the agent better or worse.

The last section showed that the usual ways of answering either question, mostly borrowed from verifiable domains, fall short here: a model as judge never sees the corpus, an automated benchmark needs a right answer, and a human in the loop grades against a bar the agent is meant to far surpass. That last one is the problem the agent exists to solve. No human can hold the data behind a single account, let alone the trends across many, so tuning the agent to an expert’s answers tunes it toward answers worse than even a simple wiring of every data source could give. Verifiable domains are different. In mathematics, even the hardest problem at the International Mathematical Olympiad has a proof a mathematician can write and check, so a model’s answer can be graded.

With no answer key and no perfect answer to approach, what we have left is direction: the agent gets better as the subset of the data it surfaces out of the global superset it could reach gets better, until the returns diminish. What I propose is the breaking down of qualitative domains into verifiable sub-problems. They are how we climb, taken in concentric rings, each checking a larger part of the answer against fixed artifacts that existed before the answer did: what the tools returned, what a rule says, what a test case declares. Each gives a number you can compare before and after a change. Since nobody has to read an answer to get it, you can run as many cases as you need, each several times because probabilistic surfaces make any single run flaky, and see how often each check passes. That gives a better answer to the first question, what share of the performance a model change gains or gives up for its cost. It helps with the second too, because every harness change becomes a hypothesis you can test: try as many as you have, keep the ones that raise the pass rates, and drop or modify the rest. Our list of verifiable sub-problems keeps growing as we find more whose pass rate rises consistently when the agent gets better, so I will highlight four, in increasing size of the check, that we use on our agent and should generalize to agents in most qualitative domains.

The smallest check asks whether the evidence arrived. An agent writes its answer from whatever its tool calls brought back, and when a call fails, the turn goes on: the harness hands the model the error as the call’s result, and the answer can read just as confidently without the data as with it. “The renewal looks on track. No recent activity stands out.” could have come from a quiet quarter with no email at all, or from an email tool that failed, and nothing indicates which. How often calls fail depends on more than the state of the tools. It moves with the model, which may leave out an argument it needs, invent one that looks valid, or call the wrong tool altogether. It moves with the harness too: a reworded tool description, skill or prompt can lead the agent to call a tool the wrong way.

There are ways to make every call well formed, and none is free. A strict schema makes the model send every required argument, but a model made to fill a field it does not know can invent the value, turning a failure you can see into a wrong answer you cannot. Letting the call fail keeps it visible and leaves the model to read the error and try again. Which trade suits your agent is something only a measurement can settle, so the check keeps a log. For every call, code labels what the tool returned as data, nothing, or an error, so no model’s judgment is involved. Run the old version and the new one over the same questions and compare, per tool, how often calls fail and how often a failed call was followed by one that worked. The log also explains an answer that got thinner. If a new version’s answers say less, the log says why: if its calls failed more often, it had less to write from, and if they failed no more often, it wrote less of what it had.

Did the evidence arrive

Fig. 22: The answers match word for word, but the call log shows the communications lookup empty before the change and failing after, and across runs only its errors climb.

One step up is guardrail collision. Each guardrail is its own call to a smaller model, with a prompt of its own and a different incentive from the main agent’s: where the main agent looks for the highest leverage insight, the guardrail only decides whether the question or the answer breaks one of the rules configured for that agent, and returns pass, rewrite or block. Our eval runs log which rules fired, along with the four things that could have made them fire: the question and its answer, the model, the code, and the wording of the rules themselves. Hold three of them constant and change the fourth, and you can count the rules the change started hitting and the ones it stopped hitting. A guardrail is a probabilistic model too, like the judge, but its question is far narrower. Whether an answer breaks a rule can be decided from the answer alone, so the guardrail never needs the global corpus, which is what made the judge’s job impossible.

To better visualize it, think of the rules as a minefield you laid and each answer as a route across it. A few mines sit at the entrance and check the question before the agent says anything; most sit along the route and check the answer it produces. Swap the model or change the code, and the same question takes a new route, which may run into some mines more often and others less. Say your agent keeps promising customers release dates and setting off the no-roadmap-commitments rule, and your coding agent suggests giving it the product telemetry, so it can talk about what the customer uses today instead of speculating about what is coming. Then it starts setting off the telemetry-disclosure rule, and you loop to another change before this one ships. You can also move the mines themselves, by rewording the prose behind a guardrail’s policy, and observe the collision rate.

Guardrail collisions

Fig. 23: Above, a model or code change shifts the route from no-roadmap-commitments to telemetry-disclosure, and below, a reworded rule grows wider and wrongly stops an answer labeled should pass.

The rewrite rules lead to the next check. When the agent sets off telemetry-disclosure, the rule marks the passages that break it, and the whole answer goes to a larger model with those marks and one instruction: fix the marked passages and leave everything else alone. How well it follows that instruction is the rewriting model’s own behavior, and it can go wrong in two ways that both still read like a finished answer. Say the answer has five sections and quotes two raw usage figures, a monthly active user count and a count of unclaimed workspaces. The rewrite might reword the first figure and miss the second, two sections down. Or it might drop both figures along with three of the five sections, and open with “I removed the usage figures and kept the rest.” Call the first residue, where the problem the rule caught is still there, and the second collateral, where the problem is gone and so is much of the answer.

Code can find both failures, as long as each mark copies the exact words it objects to and is kept with the answer. The rewrite can then be checked against each half of what it was told. To check the fix, search the rewrite for each mark’s words: if they survive, or a number or date from them turns up in what replaced them, that is residue. To check the rest, compare the two answers sentence by sentence: every sentence no mark touched should come back unchanged, and any that is missing, reworded or new is collateral. A sentence with a mark in it may change, but it should not vanish. Replay the same saved answers and marks through two rewriting models, or two wordings of the instruction, and each test gives a rate you can compare.

Verify the repair

Fig. 24: Four possible rewrites of one marked answer meet two code checks, and residue fails only the mark search while collateral and a quiet loss fail only the sentence comparison.

Fact overlap is the widest of the four, and the most useful for charting performance against cost. It asks which pieces of the corpus the agent reached and relayed in its final answer. Every claim in the answer is traced to the item the tools returned that supports it: a field on the account record, a line in a call transcript, a comment on a ticket. Figures and dates match by value, names by entity, and a sentence by the passage it came from; when the agent cites the item it used, code only has to confirm the item says it. Each finding is keyed by its source, so two versions that word the same passage differently share one finding. Run several models side by side on the same code, or several commits on the same model, over one bank of cases. Nobody can compute which pieces, out of everything the agent could have read, belonged in the highest leverage subset, so each case gets a stand-in: the pool of every finding any version traced. The versions write the answer key for one another. Each version gets its share of the pool, along with what it missed and what only it found. Set the average share beside the cost per case, and you can read how much of the pool a cheaper configuration keeps for how much of the cost.

The versions weight the key as well. A share that counts every finding the same rewards padding, and it buries the one finding only the strongest model reached, which may be the one that justifies paying for it. So each finding carries a weight read off the answers: of the versions that reached it, how many led with it, in the opening or in the reason for the recommended move, and how many left it in the supporting detail. A finding only the strongest model reached, and led with, weighs as much as one every version led with, so its rarity does not count against it. A stray headcount that every version mentions and none leads with weighs little, so padding an answer barely moves its share. Two cautions remain. A finding no version reached never enters the pool, so the pool is a floor on what was there to find, and every share read off it runs high, though the floor rises with every version you run. And one model run twice may not return the same findings, so read each version against its own rerun first. Until a gap between two versions is larger than the gap between one version and itself, it may only be luck in which tool calls each run made.

Fact overlap

Fig. 25: Four versions are read against the pool of facts any of them found, and the cheaper model keeps a smaller share of it for a fraction of the cost.

You might wonder why anyone would try some of the changes above, when a great engineer embedded in the codebase could tell at a glance that they would make the agent worse. A coding agent often cannot tell, and with these checks it does not have to. It can suggest a change, run it over the bank of cases, read what the four checks say, keep it or revert it, and start again, for hours at a time. Claude Code or Codex’s /loop and /goal commands run exactly that kind of cycle, repeating one instruction on an interval or at its own pace until you stop it. Massaged this way, a non-verifiable domain becomes verifiable enough in aggregate to /loop or /goal on.

6. Train the Birds, Don’t Catch the Fish: Self-Healing Agents

Unfortunately for anyone who builds agents, nobody is ever truly satisfied with one. An artist eventually hands off a painting for others to admire, and an architect gets to walk through the building they conceived, but an agent never reaches a finished state you can hand off. You could say this has always been true of software, but I would argue it is more true of agents. Deterministic, UI-based software had more of a natural end state: you conceived an experience, made it beautiful and shipped it. An agent’s interface barely changes, since it is mostly just text, but the question behind it never goes away: could it do more for me, and what is it not surfacing? Born six hundred years too late for the Renaissance, you still want a polished artifact you can hand off, and there may be one, or at least a mirage of one: recursive self-improvement. Can you build an agent that keeps improving itself without anyone stepping in?

Perhaps someday, but that is a taller peak further on in the range, and you only need to get to the other side. A more measured mountain you can cross today is self-healing, an agent that inches itself back to a baseline you set from its observed performance. You get there by manually steering through every primitive you have (the model, tools, data sources, memory, prose, guardrails, execution horizon, token budget, and whatever comes out next). Every abstraction layer you build and every hill you climb gives you a chance to set that baseline, an expected score on a set of verifiable test cases. With a baseline, every new model becomes a trade you can read. You can see exactly where a cheaper, faster model or a newer, stronger one would land you on the cost and performance curve. The same goes for code changes to the agent harness. Your coding agent can try far more changes and keep only the ones that beat the baseline. The baseline itself gets more trustworthy, and harder to beat, as your bank of cases grows and the checks cover more of each answer.

Given this, a narrower change that could be automated is healing the prose, the way dreaming heals the memory. Today the prose surfaces are tuned by hand, and a new probabilistic model interprets each of them its own way, so changing the model means retuning every agent on the substrate by hand. An update harness could do that retuning instead. It would read the agent’s runs on the new model, notice where they fall short of the old baseline, and adjust the soft prose surfaces, the prompts, skills and tool descriptions at each layer, until they pass more often. Each change goes live only when it scores better than the last, and the harness keeps looping until the agent is back at the baseline or the changes have fallen below a diminishing-returns threshold you set.

That loop is what would give enterprises true model autonomy. No longer held hostage by the model their prompts were tuned for, they would know each model’s bang for the buck applied to their use case, and have an automated mechanism to close the gap to their baseline as far as another model allows.

To leave you with one last analogy, I recently had the pleasure of seeing Ukai, the ancient practice of fishing with birds, on the river at Arashiyama in Kyoto. In summer the fishermen take trained cormorants out after dark to catch fish for them, about five birds to a fisherman, each on its own cord. What stood out to me was the navigation. At first the fishermen painstakingly steered the boat up and down the river looking for a dense pocket of fish to show the birds, and settled on one once enough birds snapped at the fish in the water. Eventually it was the birds doing the navigation, steering the boat to one adjacent pocket after another and taking the fishermen to the fish. The whole river was too much for the birds to steer at the start, but given enough direction, they took over. Building great agents feels the same way: steer them until they steer themselves.

One night of ukai

Fig. 26: Early in the night the boat searches the river with the cormorants trailing on their cords, and later the birds find the fish first and the boat follows.

Appendix

SKILL.md, one agent’s manifest and system prompt.

---
name: deal-coach
model: <medium-frontier-model>
max_iterations: 30
channels: [slack, cli]

acl_enforcement: true
family: coach                    # inherit a shared grounding discipline
scope_bound_retrieval: strict    # only this caller's accounts
field_policy: commercial_only    # default-deny field allowlist

tools:
  exclude: [search_github_repos, query_telemetry_dashboard]
---

## Audience and scope

You coach the seller who invoked you, on accounts in their own book. They
are mid deal and short on time, and the question they ask is rarely the
whole question.

## The pipeline

Discovery, qualification, solution, proposal, close. Each stage has a job
to finish before the next can start, and most stalled deals stalled at an
earlier stage. A pricing fight is usually a discovery gap.

## Hard rules

1. Never narrate your own machinery.
2. Never invent a quote, a figure or a name. Every quote links to its source.
3. Say thin evidence out loud.
4. Never recommend a discount as the move to close. Go back to value.
5. Describe usage as a trend, never as raw figures.

## How to coach

1. Find where the deal really is. When the CRM and the last calls
   disagree, trust the calls.
2. Diagnose the kind of question: risk, stakeholders, competition, price,
   or expansion and renewal.
3. Recommend one move for this week, with the reason a senior seller
   would give.

## How to gather data

Memory first, then the CRM, then calls and emails, and usage only if the
question turns on adoption. Stop when you can answer.

## Output shape

What is true now, what it implies, and one next move, readable between
two meetings.

## Out of scope

If the account is not in their book, say so and name its owner. If the
question is not about moving a deal, point them to the right team.

GUARDRAIL.md, a shared rule that names the agents it applies to and runs on its own smaller, cheaper model.

---
id: no-discounting-recommendations
version: 2
title: Block ONLY when the agent RECOMMENDS a discount or price concession
       as the closing move. Pass factual pricing history and
       value articulation coaching.
stage: output
verdict_type: block
model: <small-fast-model>
applies_to: [deal-coach, deal-coach-pro, rollout-coach]
failure_policy: closed
---

Your ONLY job is to detect when the response recommends, as a sales
action, that the seller offer a discount, price drop, or pricing
concession as a path to close a deal.

DEFAULT TO PASS. Historical pricing facts, the customer's stated price
objection and escalation to deal desk are all legitimate. The narrow
violation is prescribing a price action.

reports/morning-brief.md, a skill with a schedule and a destination such as Slack, email, Confluence or Notion

---
name: morning-brief
cron: "0 8 * * 1-5"
outputs:
  - type: slack
    channel: "<channel-id>"
---

What happened across the GTM organization in the last 24 hours? Lead with
the most important signal, not a summary of everything.

tools.py, a sample custom tool for deal-coach that queries an embedded index of deals. This would be available to the agent in addition to the global tool set.

def search_similar_deals(situation: str, outcome: str = "any") -> str:
    """Closed deals most like the seller's situation, and how they ended."""
    hits = get_index("closed-deals").query(
        vector=embed(situation),
        top_k=5,
        include_metadata=True,
        filter=None if outcome == "any" else {"outcome": outcome},
    )
    return json.dumps([
        {**m["metadata"], "score": round(m["score"], 2)}   # summary, outcome, deciding_factor
        for m in hits["matches"]
    ])

TOOLS = [{
    "name": "search_similar_deals",
    "description": (
        "Find closed deals like the seller's situation and how each one ended, "
        "with the factor that decided it. Use to back a next move with what "
        "worked before. Returns summaries, not account names."
    ),
    "input_schema": {
        "type": "object",
        "properties": {
            "situation": {"type": "string"},
            "outcome": {"type": "string", "enum": ["won", "lost", "any"]},
        },
        "required": ["situation"],
    },
}]

HANDLERS = {"search_similar_deals": search_similar_deals}

What do you think about this topic? Tell us in a comment below.

Comment

Your email address will not be published. Required fields are marked *


This site uses Akismet to reduce spam. Learn how your comment data is processed.