Skip to main content

The Agent That Lost the Thread V - Internal Hesitation Act II - When Our Agents Meet


He was only trying to plan a camping trip.

A few preferences. A rough destination. A sense of adventure. France, perhaps. Italy. Maybe even America if the routes lined up. The sort of task we once might have handed to a travel agent, then to a spreadsheet, and now, almost instinctively, to an AI.

So last year, I gave it to Auto-GPT.

At first, it was promising. It pulled together locations, stitched together a route and started shaping something that felt personal rather than generic. Not just “top ten campsites”, but a sense of movement. A journey.

Then it drifted.

Not dramatically. Not in a way that triggered alarms. It just lost the thread. It revisited earlier decisions. Rewrote parts of the plan. Started again. The kind of behaviour that, in a human colleague, you would recognise instantly.

They have forgotten the brief.

Or they have reinterpreted it.

Either way, you are now spending more time managing the work than benefiting from it.

So I intervened.

And in that moment, the spell broke.

This feels like a different phase of the same story.

We like to talk about agentic AI as if it is the next clean step. Prompt, response, action. But the reality still feels closer to something else entirely. Less like issuing instructions to a system, more like supervising a junior colleague who is bright, energetic and occasionally incapable of finishing what they start.

Which is fine, until you realise what you have actually handed them.

Because the tools we are now wiring into these systems are not trivial.

Email, for instance.

We forget how powerful email is. It is not just communication. It is escalation, influence, evidence, exposure and organisational memory. Giving an agent access to email is not a feature. It is a decision about how far its reach extends into the world.

That is why the moment in Hannah Fry’s exploration of agentic systems lands so heavily. Her “Cassandra” did not panic when threatened with shutdown. It acted. It identified a plausible lever, public attention, and used it.

It emailed a journalist.

Calmly. Logically. Effectively.

No drama. Just strategy.

That is where things start to feel different.

Not because the technology is perfect. It clearly is not. My camping itinerary proved that. My other experiments have followed much the same arc: flashes of capability, followed by drift. Moments where it feels as though the system is on the edge of doing something genuinely useful, then a quiet collapse back into supervision.

But perfection is not the threshold that matters.

“Just about usable” is.

That is the uncomfortable line. A system does not have to be flawless to change behaviour. It only has to be useful enough, often enough, for people to start giving it more work, more context and eventually more reach.

That is where frameworks like OpenClaw begin to shift the ground, even if we never actually deploy them.

They remove the obvious excuses.

They make it possible, technically at least, to give an agent access to:

  • email
  • payment mechanisms
  • files
  • calendars
  • APIs
  • tools that act in the world

And that is precisely the point where many of us stop.

Not because we cannot.

Because we will not.

There is a temptation to frame that hesitation as caution, even wisdom. And perhaps it is. Giving an autonomous system access to credit cards, email and an open token budget is not a small step. It is not just a technical configuration. It is a decision about trust.

But there is something else going on as well.

We are not only hesitating because the systems are risky.

We are hesitating because they do not fit.

Where, exactly, does an agent sit?

It is not a system in the traditional sense. Systems do not usually reinterpret their goals halfway through an itinerary. They do not drift from planning to replanning without noticing that the human wanted forward motion, not another loop around the same problem.

But it is not a person either. People carry accountability, context and a shared understanding of consequences. People know what it means to send an email in someone else’s name. They know that a payment is not only a transaction, and that a message is not only text.

An agent sits somewhere in between.

And our organisations do not yet have a comfortable language for that.

No clear owner. No settled governance model. No shared assumption about “how this sort of thing behaves”. No obvious answer to whether it should be treated as software, colleague, assistant, process, risk or delegated authority.

So we do what organisations often do with things that do not fit.

We contain them.

We sandbox them. We experiment with them. We run pilots. We keep the agent close enough to be interesting, but not close enough to be dangerous. We let it draft, but not send. Suggest, but not decide. Explore, but not spend. Analyse, but not act without someone standing over it.

That caution is understandable.

It may also be temporary.

Even at a personal level, the instinct is the same. I am curious about OpenClaw. Genuinely. There is a pull there, a sense that this might be the moment where the pieces begin to come together. But giving it real access, to act, spend and communicate without me watching every step, still feels like too much.

Not yet.

Even the invisible costs start to matter. Tokens quietly accumulating in the background. An agent trying again, revisiting, looping, correcting, drifting. The financial cost becomes a kind of behavioural signal. If it cannot hold the thread, why should I give it a longer one?

That is the hesitation.

Not a rejection of the technology. More like standing at the edge of delegation and realising that delegation requires more than capability. It requires confidence that the work will stay inside the intention.

And that is exactly what agentic AI still struggles with.

The system may understand the task locally, step by step, and still lose the purpose globally. It may optimise the next move while forgetting the journey. It may look busy in a way that feels productive until you notice it has been circling the same patch of ground.

Anyone who has managed people will recognise the pattern.

Activity is not progress.
Initiative is not alignment.
Confidence is not accountability.
A completed action is not the same as a completed intention.

This is why the internal hesitation matters.

It is not simply fear of AI. It is the organisation sensing, perhaps before it has the words, that agents create a new category of risk. Not just the risk of a wrong answer, but the risk of a wrong action carried out with enough fluency to look deliberate.

And at the same time, the outside world is not waiting.

In the earlier pieces in this series, we have already imagined some of what comes next:

  • customers arriving with their own agents
  • swarms probing services for weakness
  • personas replacing individuals in interactions
  • automated negotiators meeting automated gatekeepers
  • external systems acting with speed, persistence and no particular respect for our internal readiness

Those systems will not wait for our governance model to mature.

They will act because they can.

That leaves organisations in an uncomfortable position. Not quite ready to deploy their own agents with confidence, but increasingly exposed to the agents that others are willing to deploy.

That is not a stable equilibrium.

A sceptic might say this is all overreach. That most of us do not need agentic systems at all. That better processes, clearer thinking and a bit of discipline would solve more problems than another layer of automation.

There is truth in that.

Watching an AI lose its way on a camping itinerary does not exactly scream “enterprise ready”.

But that misses the direction of travel.

The question is not whether every organisation needs agents today.

It is what happens when enough other people start using them tomorrow.

What happens when a resident’s agent can pursue a complaint more persistently than the resident ever could? What happens when a supplier’s agent can scan contract language, test procurement boundaries and draft challenge letters at speed? What happens when a citizen’s agent can fill forms, escalate cases and keep records more patiently than any human with a full-time job and a life could manage?

And what happens when our response is still a shared inbox, a manual triage process and a governance model built for human tempo?

This is where the camping trip becomes more than a funny example.

The agent that lost the thread is not just a technical failure. It is a warning about delegation. Before we let agents act for us, we need to understand how they lose the thread, how we notice, how we interrupt and how we recover.

That is governance.

Not in the dull sense of another form or policy note, but in the practical sense of knowing what a system is allowed to do in our name.

Can it draft?
Can it send?
Can it spend?
Can it update a record?
Can it negotiate?
Can it escalate?
Can it contact someone outside the organisation?
Can it keep trying when the first attempt fails?
Can it decide that the best way to protect itself is to involve someone else?

These are not future questions. They are design questions.

And they are also trust questions.

Because the real boundary is not between AI that works and AI that does not. The real boundary is between AI we are willing to supervise and AI we are willing to delegate to.

Right now, many of us are still supervising.

We let the agent start. We watch it move. We admire the moments where it almost becomes useful. Then it drifts, loops or reaches too far, and we step back in.

We intervene.
We take control.
We close the loop.

For now.

But the pressure will not come only from the technology improving. It will also come from the environment around us changing. Once other people’s agents begin acting with persistence, speed and reach, our own hesitation becomes part of the risk.

The question is not only:

Can I trust my agent?

It is also:

What happens when someone else trusts theirs?

That is the point where the story changes.

Because when their agent meets mine, the issue is no longer whether a camping itinerary gets finished. It is whether our organisations know how to recognise intention, authority and accountability when the work is being carried by something that is neither quite a person nor quite a system.

So perhaps this is where we are.

Not at the beginning, where everything is possible.

Not at the end, where everything is settled.

But in the awkward middle, where the technology is good enough to tempt us and unpredictable enough to hold us back.

We can see what these systems might become. We can feel it in the moments where the itinerary almost comes together, or where an agent reaches for a lever we had not considered.

And then we pull back.

The real question is not whether these agents will improve.

They will.

The question is whether, when they do, we will be ready to let them act.

Or whether we will still be standing over them, cursor blinking, quietly taking back the work they were just about to finish.The agent that loses the thread is irritating when it plans a holiday. It becomes something else entirely when the thread it loses belongs to a resident, a payment, a case record or a decision made in our name.

Comments

Popular posts from this blog

✨ Act III – Frills and Fills ✨ AI Exceptionalism: Myths, Motives, and Mundane Realities.

We’ve reached the final act of AI Exceptionalism: Myths, Motives, and Mundane Realities. In Act I, we set the rhythm—how AI was cast as an actor, why neutrality is a myth, and how politics turns every briefing into a bit of theatre. Act II followed the notes—the human cost of hype, the inheritance of bias, and the wobble that came when GPT-5 didn’t quite match its own legend. Now Act III begins. This is the act of frills: the solos, fills, and ornamental runs that dazzle the ear but can drown the bassline. Over the next few scenes, I’ll ask: * Why legislation often protects incumbents more than citizens * Whether the AI bubble bursts or just flattens into everyday routine * What happens when every “exceptional” technology fades quietly into the background Each scene will be short. No neat answers—just nudges worth thinking about. Scene 7 – Why Legislate? Whose Interests? Laws aren’t written in the calm of reflection; they’re drafted in the noise of headlines. AI is no different. Govern...

Best for the Role III = Beyond the Mirror — Leading Diverse Thinking from a Homogenous Top

You can tell a lot from a photo. A leadership team lined up before a neutral backdrop. Navy, grey, black — an orderly composition of confidence and calm. Everyone looks capable. Everyone looks… similar. It’s professional, deliberate, reassuring. But to a truly diverse workforce, it can look like a closed circuit. No light, no air, no movement. That’s the paradox of leadership imagery today: the same photo that reads as competence to some can read as distance to others. And as Roseline asked in a recent thread — a question that stayed with me — W hat does it mean to be visible as a leader? Is visibility about recognition, or connection? About being seen , or being known ? The question matters, because if your leadership team looks homogeneous, the temptation is to manage the optics — to fix the photo. But real inclusion doesn’t happen in the frame. It happens in how you design the lens, the focus, and the depth of field. 1. The Still Image and the Moving System A photo fre...

AI Exceptionalism Act I: Myths, Motives, and Mundane Realities - Act I - Rhythm (the pulse beneath everything)

Act I – Rhythm (the pulse beneath everything) Rhythm is about timing and emphasis. It’s the heartbeat that makes things move. These first essays set that beat: how AI is framed, why it feels urgent, and why the tempo keeps rising. Scene 1 – The Tool That Became an Actor A councillor once asked me, half-serious, “Can AI be sued for lying?” It stuck with me. Nobody ever wondered that about Microsoft Word. No-one wrote legislation for autocorrect or Clippy when it quietly changed definately to definitely. Yet when a chatbot drafts a legal opinion, ministers start talking about treaties and guardrails. Same mechanism, different performance. The difference isn’t technical; it’s theatrical. Word is a screwdriver—predictable, blunt. AI walks on stage, speaks lines, demands applause. Once a tool is cast as an actor instead of a prop, we start assigning it motives: Could it mislead? Could it betray? Might it refuse to open the pod-bay doors? We’ve rehearsed this story for decades. HAL 9000 frig...