How would you know if your AI starting giving you less?

·

Old-school automation broke the same way every time, so you always caught it. AI doesn’t.

Dave Heywood discusses why your AI could start doing less for you right now, and what to do about it.

The lesson from the first automation wave that still applies now

A decade ago, the first wave of automation projects broke in predictable, static ways. A bad rule failed the same way on the same input every time, which meant someone eventually (hopefully!) noticed and fixed it.

AI doesn’t offer that safety net. Same prompt, same process, different day, and you might get a different result.

Forrester’s research into RPA failure found the technology itself was rarely the problem. The real cause was automating processes with too much variability, documents in different formats, rules that shifted every quarter, edge cases a rigid script couldn’t handle. The projects that worked were the ones where the manual process was mapped in full before anyone touched an automation tool. The ones that failed skipped that step and simply automated the mess, faster.

Why AI breaks the old model

Rigid, rules-based automation fails identically every time, which is exactly what makes it catchable. AI doesn’t offer that same courtesy. Gartner found that roughly 60% of programmes ended up as “zombie bots,” automated processes nobody’s actively monitoring or maintaining anymore.

If that’s the story for something rigid and rule-based, what happens when the thing quietly drifting off course is making judgment calls instead of following a fixed script?

Dave’s hunch, and what’s actually proven

Dave puts a theory on the table: that AI providers may be quietly rationing model output to manage the eye-watering cost of running these systems at scale. No study confirms this. But the underlying economics are driving this way. The cost of running an AI firm scales with every query, and it’s the single biggest cost line for every provider in the market. That’s a standing incentive to do less per call, whether or not any single provider is acting on it.

What is documented is that vendors already make quiet, undisclosed changes to live models. OpenAI’s rollback of an overly sycophantic model update is a real precedent, and the company admitted the change had been shaped too heavily by short-term feedback signals. Whether or not the rationing theory holds, the underlying risk does: if a provider changed what you’re getting tomorrow, most businesses have no way of knowing, and no way of knowing what it already cost them.

The problem with baking a cake blind

Without checkpoints, diagnosing an AI failure is like pulling a ruined cake out of the oven with no idea what went wrong. Wrong ingredients, wrong order, wrong temperature, wrong tin, any of it? all of it?

A checkpoint isn’t friction for its own sake, it’s the only instrument available for isolating where a process actually went wrong.

Two things to do to get a better grip on your AI workflows

Name an owner. Every AI workflow needs one specific person, not a committee, whose job is to review outputs and flag when something feels off.

Compare and contrast your outputs. Take a query you run regularly, rerun it a week later, and compare the two results. It takes ten minutes, and it’s the only real way to catch drift, whether that’s a model quietly doing less, a provider swapping what’s running underneath, or anything else happening upstream with zero visibility.

Neither approach needs a new platform or a transformation budget. It just needs someone actually looking, on a schedule, instead of assuming the tool’s working because it was fine last time.

Full episode transcript

How would you actually know if your AI was doing less for you today than it did yesterday? This is Scale. An OX7 Partners Podcast. and I’m Dave Heywood

A lot of firms went through this exact same thing about a decade or so ago with the first wave of automation. So when we think about things like robotic process automation,

And that had the problem of breaking

in a predictable manner based on a static set of inputs. So somebody eventually caught on and fixed it.

Now with this new wave of AI automation, that doesn’t really happen.

It fails in different ways, some louder, some quieter. half the time it’s really hard to notice and pick up on

And something I genuinely can’t prove, but I’ve felt it all the same.

Is the notion that some of these providers starting now to ration what models actually give us in order to manage their own costs. I I couldn’t find any particular study which proves this one way or the other.

But by the end of this conversation I’ll I’ll show you what is proven and a couple of things that you can practically do that’ll

help you catch

where a well thought through AI process might be falling short.

So let’s travel back a little bit first, because this isn’t really a new problem for us.

So if we look at that first wave of automation, and robotic process automation is a particularly pertinent subset of that, Forrester did some research into why those projects particularly failed. And it wasn’t the technology itself that let people down.

It was down to the fact, but a majority of failures came from automating processes that had too much variability in them. So think of documents in different formats, rules that changed every quarter, messy, unstructured inputs, that a rigid script just couldn’t cope with. And that’s why there’s been so much excitement around AI driven processes.

and I saw some of this myself earlier in my career, running early automation projects. And the ones I found that worked were the ones where we sat down and got a blank piece of paper and a pen and manually mapped exactly what was happening where.

Every exception, every weird little edge case where a machine couldn’t perform to the quality that we expected,

And designing routes for that to go, so we got the expected outcomes.

The ones that didn’t work were where we skipped that where we bowed to this pressure to move quickly and we ended up just automating a bit of a mess. It was a faster, more efficient mess, but a mess all of the same.

Old school automation is based on this premise of same input, same output every single time.

Which gives us the advantage of when something breaks or doesn’t work, it breaks in a very similar, predictable way, someone notices the pattern and we can go and fix that.

But AI is a bit of a different beast here. Same prompt, same process, different day, and you might get a different output.

And this matters perhaps more than we might think, because Gartner found that something like sixty percent of classic automation, end up with what they called zombie bots. So processes that got automated, and then people just set it and forget it. No one’s monitoring it, no one’s maintaining them, no one really understands.

how they’re designed and what they do anymore. Now, if that’s happening with something that’s very rigid and rules based that fails identically each time, what happens when the thing drifting off course is actually making judgment calls instead of just following a fixed script?

It demands even greater vigilance.

and to add to the mix, a little hunch that I’ve got, and it’s just a hunch, something that I’ve been feeling but can’t really prove one way or other, I’m pretty convinced that AI providers are to some extent rationing what these models actually do.

In a quest of finding that path to profitability, as we know they’re all hemorrhaging money like it’s going out of fashion, we’re seeing moves towards price increases and I’m pretty sure some token rationing going on quietly behind the scenes as well.

we know running these models are unbelievably expensive and that cost isn’t fixed at all. It scales with every single query and as we become more confident we perhaps use more complex queries.

So the mechanics of just running these models on a global scale is the single biggest cost line for every single provider out there. And so it creates this incentive How can we achieve an output that satisfies the user

while also using less resources per call

and we’ve seen slightly different versions of this play out a little bit. So OpenAI had to roll back an update to one of their models because it had become extremely sycophantic.

You know, I could have given it some of the

worst business ideas ever, like opening a car wash at the Tour de France. And it probably would have

lauded that as a visionary idea but only I could have dreamed of. Aren’t you really clever with a little virtual pat on the head?

and as they rolled it back, they even admitted that the change had been shaped too much by some of that short term feedback without really thinking about how people’s use of it evolves over time. Now that’s a slightly different example, but it’s proof that these firms are pushing these quiet little changes into live models.

For their own reasons, and you only find out about it after the fact, if perhaps you find out at all. So I may or may not be wrong about rationing specifically, but ask yourself this question: if your provider decided to give you less tomorrow, how would you actually know? And what would it have already cost you by the time you noticed?

Now that’s all well and good, but we’re in the business of doing something about this.

So what can you do?

I found from experience that slowing things down a little bit and building in proper human checkpoints has been perhaps one of the single biggest things that’s actually made AI work better for me and with me. And it it sounds a bit counterintuitive because everyone’s selling the dream of speed and scale here. But think about it like baking a cake.

If you throw all your ingredients together, shove it in the oven and it comes out a complete mess.

You’ve got no idea what went wrong. Was it wrong ingredients? Did we mix it in the wrong order? Was the oven too hot? Did we use the wrong kind of tin? Could it be one of those, some of those, all of those? We’ve got zero way of really isolating the the failure.

And that’s exactly what running AI processes without some of those checkpoints is like.

So a checkpoint isn’t necessarily friction for friction’s sake. It’s one of the few instruments that we’ve got in a process.

that allows you to identify something and say, right, that’s where we drifted off rather than just staring at a bad output, getting really grumpy and not understanding why.

I’ve got a couple of other practical things that we can do and put in place pretty much straight away.

First of all,

have a named individual against every AI process or workflow that you’re running. Not a committee, not a team, a single individual whose job it actually is to sit down at a set point and look at what’s coming out the other end and flag when something feels off.

Super obvious.

Yet not that many people are actually doing it.

The second one, again which also costs you very, very little, is take a query or process but you run regularly.

And run that at set intervals and check the outputs against each other.

it might take you ten minutes tops but it’s a really, really good way of benchmarking those outputs and getting a sense of if something’s drifted. whether that’s a model that’s slightly less robust and thorough than it was a couple of weeks ago.

Providers swapping out what’s actually happening under the hood or anything else happening that you’ve got zero visibility into. And neither of those needs huge spend, a new platform, or a transformation project wrapped around it. It’s just about getting into this idea of somebody actually looking and monitoring on a schedule instead of assuming the thing’s fine, because it was when we set it up.

I hope you found that useful. I’m Dave Heywood and this is Scale, an OX7 Partners podcast, and I’ll see you next time.



Leave a Reply

Your email address will not be published. Required fields are marked *