<- Back to all posts

Your people decided what your AI is for before they typed anything

01 September 2026 · 14 min read
aiadoptionrolloutbehaviourhow-i-work

I never send engineering work to my assistant from my phone. I checked a month of my own logs and there is not one. I also never decided to do that.

That is a small, strange fact about my own behaviour, and chasing it down took me somewhere I did not expect, by way of being wrong twice in public. If you are running an AI rollout that has gone quiet, I think the ending is worth your time, so let me show you the whole thing including the parts where I had it backwards.

Start somewhere that has nothing to do with AI

You behave differently in a church, a library and a pub. Obviously. What is less obvious is how little of that is you deciding anything.

Nobody hushes you in a library. You do it on the way through the door, before anyone has said a word to you. A researcher called Roger Barker spent decades on this from the 1950s onwards, watching an entire small town, and gave the pattern a name: behaviour settings. His finding was blunt. The setting carries a script, and it runs roughly the same way regardless of which individuals turn up. Swap the whole congregation and the church still behaves like a church. The building is doing work that we tend to credit to the people inside it.

Three ingredients do that work.

The first is who can see you, and whether they will still be around next week to remember it. This is the one people underrate. A pub is not private, it is crowded, but the crowd has no continuity with the rest of your life and will have forgotten by Tuesday. A church congregation is the same people every week, and they carry what they saw forward. Observation on its own changes very little. Observation plus continuity changes almost everything.

The second is what the room physically allows. Pews face one direction, which makes one person the speaker and everyone else an audience. A library has nobody within conversational reach. Pub tables are round and the noise floor is high enough to make quiet conversation possible. None of that is enforced by anyone. It is enforced by furniture.

The third is a threshold. The door, the hush, ordering at the bar. There is a moment of crossing over, and that moment is when you switch.

Organisations used to get all of this for free, and I have never seen anyone put a number on it. The office did the mode-setting. Then open plan collapsed the quiet room and the social room into a single space, so no room cued anything and you got the noise of the pub with the surveillance of the church. Then remote removed the building altogether.

Since then we have been trying to reproduce with policy what architecture used to do for nothing, and policy is a far weaker instrument. I think that is why return-to-office arguments never quite land. People are reaching for something the building genuinely did, and the only vocabulary the meeting allows is productivity, so they say productivity and everyone correctly senses that is not the real argument.

So does a screen do the same job?

I use the same AI assistant two ways. In a terminal window on my laptop, and in a chat app on my phone. Same assistant, same underlying model, same person typing. Two very different windows.

My prediction was that I would be politer on the phone and more clipped in the terminal. I had it go through my own logs and count. About 600 prompts I had typed over a month.

I was wrong. Politeness, hedging, message length, sentence structure, all flat once you compare like with like. Whatever I thought the interface was doing to my manners, it was not doing.

I was ready to write that up as an honest null result and move on, which is the first place I nearly went wrong.

What I had not thought to measure

Before publishing the null I looked at what else was in the data, and found something I had not gone looking for.

About one in six of my terminal prompts contained code or a file path. On my phone, one in a hundred. Across a month, I did not send a single piece of engineering work to the phone.

The vocabulary split even harder than the numbers. In the terminal I write about state, logs, directories, configuration. On the phone I write about people, invoices, meetings, traffic. The terminal is where I build things. The phone is where I run the business.

I had never decided that. I had never noticed it. If you had asked me directly I would have denied it, because both windows reach exactly the same assistant and either would have answered either question perfectly well.

That is the finding that matters, and I will come back to what it costs.

The second place I went wrong

There was one more thing in the data. I start messages with a capital letter 14% of the time in the terminal, and 56% of the time in the chat window.

I dismissed it. Phones capitalise the first letter automatically, so of course the phone number is higher, and the effect is an artefact of the keyboard rather than anything about me. That is a perfectly sensible explanation. It is also completely false, and I only found that out because I remembered that I send a large share of those messages from a desktop, where no keyboard is helping me.

So I compared properly: messages sent from a chat window while I was demonstrably sat at the same machine, minutes either side of typing in the terminal. Same device, same keyboard, same hour. The gap holds.

That is worth pausing on, because dismissing it was the actual error. I had a large, statistically obvious effect and I explained it away in one line without testing the explanation. A null result from an instrument that cannot see the thing looks exactly like a null result from a thing that is not there.

The comparison that made sense of all of it

Then I added the piece that was missing, which was people. I had been comparing two ways of addressing my assistant without knowing how I address a human being. So I pulled about 2,500 messages I had sent to actual people.

To a person, 87% start with a capital letter. To my assistant in a chat window, 56%. To my assistant in a terminal, 14%.

A capital letter is a small thing, and I am not claiming it measures respect or effort. What it marks is a choice of address: whether you are writing a sentence to someone, or issuing an instruction to something.

Look at the shape of that. It is a slope, not a switch. My assistant sits somewhere between a person and a tool, and where it sits depends on which window it happens to be in. Changing the window moved my mode of address about as much as changing whether the counterparty was a human being at all.

Two other measures produced the same ordering, independently. I drop apostrophes in a way I never do with people. My messages get longer as they get less personal.

The third check, which produced the best result

Then I realised I might be measuring nothing more interesting than conversation length. A terminal session runs for forty exchanges. A text to a friend is a one-off. If I am formal at the start and terse by message thirty, then all I have discovered is that I have longer conversations with my laptop.

So I looked only at opening messages, the first thing typed after a gap, using the same rule everywhere. The gradient held and widened slightly: 87%, 56%, 14%.

Then I looked at how formality changes as a session runs on, expecting a decline.

It does not decline. It rises very slightly, in all three, right through to message thirty and beyond. My opening move is the least formal moment of the entire exchange.

Which means the way I address the thing is not something that erodes over a conversation. It is set at the moment I arrive, and then held.

That is the threshold. It is the third ingredient from the pub and the church, turning up in a text box on a laptop, and I was not looking for it when I found it.

Where this disagrees with the research

Barker’s work has been extended to virtual environments, and there is good recent work doing exactly that. But it is all about immersive spaces: virtual worlds, mixed reality, places built to surround you. The theory expects a setting to have an environment that encloses the behaviour and structurally matches it, which is why the extension goes where it goes.

A text box has none of that. It encloses nothing. By the letter of the theory it should not be a behaviour setting at all, and it should not produce a threshold effect.

It did anyway. So either the boundary condition is wrong, or the room is not made of space in the first place. My guess is the second one: what I am reading is a protocol rather than an environment. A terminal has a protocol, which is command and response and exit code. A chat window has a different one, which is turn-taking and address and the expectation of a reply. That is what I am responding to when I decide, before typing, what sort of thing I am talking to.

I would not push that hard on one person’s logs. But it is testable, and it would be easy for someone with access to a few hundred users to test properly.

The thing organisations already see and file wrongly

There is a version of this that every large organisation has already observed. You roll out the sanctioned AI tool, and people carry on using their personal one. It gets filed under security, or compliance, or policy enforcement, and the response is usually a reminder about data handling.

I would take it as a measurement instead. People taking the real work somewhere else are telling you something specific about the room you built. Not that they are careless, but that the sanctioned tool has been categorised, by them, as a place where certain kinds of problem belong and others do not.

What this means if you are rolling AI out to a few hundred people

We train people on what to type. Prompts, phrasing, context windows, role instructions, the whole magic-words industry. Everything above says the decisions that determine the outcome are made before anyone types anything, and there are two of them.

The first is what belongs here. Your people are working out for themselves what class of problem this tool is for. They are guessing, and they guess conservatively, because guessing high in front of your employer has a cost and guessing low does not. Nobody has told them the difficult problems are in scope. So they bring the safe ones, and eighteen months later the tool looks like an expensive way to tidy emails, and somebody concludes the technology was oversold.

The second is who they think they are addressing. Everybody files it as something. Filed as a search box, it gets search-box questions and returns search-box answers, which confirms the filing. The interface does most of that categorising on their behalf, and nobody chose the interface for that reason. It was chosen because it worked with your single sign-on.

Underneath both sits the oldest ingredient, which is who is watching. Most corporate deployments are vague about whether prompts are logged, who can read them, and whether any of it feeds back into how someone is judged. Vagueness feels like the safe option and it is the worst one available, because people resolve ambiguity pessimistically. They assume the worst plausible answer and then bring only the work that is safe to be seen bringing.

That is the pub and the church again. In a pub you are surrounded by people and you speak freely, because none of them will carry it forward. In a church you are watched by people who will. An employee who cannot tell which one they are in will behave as though it is the church every time.

So say it plainly. Tell people it is logged, tell them who reads it, tell them what it is and is not used for. It sounds worse in the room than staying vague. It produces better behaviour, because a known rule is a setting and an unknown rule is a threat.

None of that is a prompting problem. It is a question of what room people think they have walked into, and what the social rules are once they are inside.

I am not exempt from this

While writing this I ran the same audit on myself, sorted by what kind of problem I was bringing rather than how well I wrote it. A hundred and seventy-three conversations.

Thinking and sparring came to 2.9%. Five conversations out of a hundred and seventy-three. My own stated test for whether this is working is that it should be seamless and I should not have outsourced my thinking. On this evidence I have not outsourced my thinking, because I am barely bringing it. That is the opposite failure to the one I worry about, and it is still a failure.

Research and analysis came to 1.7%.

Meanwhile, winning and delivering work is about a fifth of what I bring. Content, admin and building my own tools together are about twice that. The tool is getting the tractable work rather than the hard work, in exactly the way I have just spent two thousand words describing in other people’s organisations.

I did not know any of that on Friday.

What I cannot tell you

I can show you that the mode gets set at the threshold. I cannot show you that a threshold can be built deliberately, and I am suspicious of anyone who says otherwise.

Every attempt I have watched to name a mode into existence has either quietly died or turned into theatre within about a month. The no-laptops meeting. The designated thinking space. The channel that is supposed to be for one kind of conversation and is full of another by week three. Behaviour settings in the physical world grew over decades and are held up by architecture and habit, and I do not know whether you can shortcut that with an announcement and a naming convention.

If they can only emerge and cannot be designed, then the advice is not “build better rooms” and I do not yet know what it is. That is a genuine gap and I would rather leave it open than fill it with something tidy.

The other limit is obvious. This is one person, my own logs, one month. Everything here is a hypothesis with a sample size of me.

Which is the only reason to do the next bit yourself

Pull last month’s conversations out of your sanctioned AI tool. Conversations, not individual prompts. I tried it at the prompt level first and half of it was unclassifiable, because half of any prompt log is “yes”, “go” and “do that one”, continuations that carry no topic of their own. Group them into exchanges first.

Do not score them for quality. Sort them by what kind of problem each one was.

Then, separately and without looking at the first list, write down the five things your team has actually struggled with this month.

Compare the two lists. The gap between them is the size of your adoption problem, and no amount of prompt training closes it.

Then take that gap to whoever signed the rollout off, and ask them who owns it.

In most organisations I go into, nobody has been named.

If this is your problem

AI made tasks faster, then progress levelled off?

I run a fixed-price diagnosis: one to two days, £4,500, and you get your level measured and the constraint named in a document written for your board, not your backlog. There is also a 90-minute Maturity Review at £750 if you want a smaller first step. Earlier in the journey — still choosing tools and building fluency? That is what the workshops are for.

Tim Robinson

Transformation Consultant & AI Practitioner

20+ years fixing how organisations work. I help leadership teams redesign operating models and apply AI where it actually matters.

Book a quick chat