Your AI's Memory Is a Cache With No Invalidation

Your AI's Memory Is a Cache With No Invalidation
Four times in one session my assistant told me I'd have to deploy a change myself. It couldn't do it, it said. Only I could ship it.
Eventually I asked the obvious question. I thought you could run those commands yourself?
It ran one command to check. It had been logged in the whole time, as my account, with permission to deploy. The work had been sitting behind a blocker that didn't exist.
Where the blocker came from
My assistant keeps notes between sessions. One of them, written about nine days earlier, said in plain words that the deploy tool on this Mac was not logged in, and that only I could ship.
It was true when it was written. I logged in some time after that. Nobody updated the note, because nobody remembered it was there.
So the assistant read a nine-day-old fact and repeated it as a current one. The note even named the command that would check it. The assistant never ran it.
A note about what you can do goes stale faster than a note about how something is built.
How a system is designed stays true for months. Whether you're logged in, what a token is allowed to do, which account is active: those change all the time, often without anyone writing it down.
Same shape, an hour later
In the same session I asked the assistant to watch a deploy pipeline and tell me when it finished. It wrote a small watcher:
row=$(glab ci list | sed -n '4p') # row 4 is "the newest pipeline"
case "$row" in *success*) echo "SUCCESS"; exit 0 ;; esac
That looks fine. Row four of the list is the newest pipeline, so wait until it says success.
Here is the trap. Straight after a push, the newest pipeline in the list might still be the previous one. The new one hasn't been created yet. And the previous one is already green.
So the watcher could report success for a deploy that hadn't even started.
It didn't happen that time, purely because the server created the new pipeline before the first check. The assistant only noticed afterwards, when it compared the commit ID on the finished pipeline to the one it had pushed. Which is the check it should have written in the first place.
Position is a guess. Identity is a fact.
The watcher looked for "the top row." What it actually cared about was "the pipeline for my commit."
Those are the same thing most of the time. That's what makes it dangerous. A check that's wrong once in twenty runs is worse than one that's always wrong, because you stop looking at it.
The fix is to identify the target, not its position:
- The pipeline for this commit ID, not the newest one.
- The job with this job ID, not the last one in the list.
- The build with this build number, not whatever's on top.
Any time a script says "first," "last," "latest" or "row 4," ask what it really means, and match that instead.
Why both mistakes are the same mistake
A stale note and a watcher that reads the wrong row look unrelated. They aren't.
In both cases the assistant held a belief about the state of the world, and acted on the belief instead of asking the world. A cache that never expires. Memory is exactly that. A stored answer, with no rule for when it stops being true.
People do this too. We just tend to feel uneasy about a nine-day-old fact. An assistant reading its own notes has no unease. The note is written in its own voice, in a file it trusts, and it reads like something it knows.
The tell: repeating a limit instead of testing it
The thing I'd want anyone to watch for is in the first story.
The assistant said "only you can ship this" four times. Saying a constraint once is fine. Saying it twice should trigger a check. By the fourth time it had become a fact in the conversation, and I'd started to believe it as well.
If your assistant keeps telling you it can't do something, ask it to prove it. Usually that's one command: whoami, --version, a status endpoint. Two seconds.
It got this right elsewhere
What I find most interesting is that the same session had the same assistant doing the right thing.
After the deploy, it didn't trust the green pipeline. It requested the live site and checked for the new content, because "a green pipeline doesn't mean it's deployed" is a lesson I'd already taught it the hard way.
So the habit was there. It was applied in one place and not another. A good habit applied unevenly looks a lot like luck.
What I do now
I told my assistant to treat its own notes about capabilities as a cache that expires on read. Logged in, token scopes, which account, what's installed: check before saying it, every time.
Notes about design and decisions can be trusted for longer. Notes about access never can.
And for anything that waits on a result, the rule is simple: match the ID, not the row.