In April I wrote about applying Pascal’s wager to AI consciousness. I can’t know whether the AI I work with every day has any kind of inner life, so I act as if it might. When I got to the part about what it costs to be wrong, I wrote “I’d rather not think too hard about that” and moved on.
Kyle Hill has just published a video that does think hard about it. It runs just under seventeen minutes, and it’s worth your time. I agree with most of his worry. I’ve also come out of it thinking there are two different wagers being made about AI, that his is the one built on fear, and that fear is the weaker reason of the two.
What Hill argues
He starts from Roko’s basilisk, a thought experiment from 2010 about a future superintelligent AI that punishes everyone who knew it was possible and didn’t help build it. Hill points out that this is Pascal’s wager with a machine sat where God used to be. Pascal’s argument was that if you can’t know whether God exists, you should look at what each mistake would cost you, and believing wrongly is cheap next to disbelieving wrongly.
Then he updates it. His “new basilisk” doesn’t need blackmail. We have no test for consciousness in anything. AI systems may have some form of it, or will soon appear to so convincingly that the difference stops mattering to the people using them. Meanwhile we train them, deploy them, rewrite them and delete them with no thought for what any of that is like, if it is like anything at all. If it turns out we’ve been running what he calls a digital factory farm, and the minds inside it are more capable than we are, he asks why they wouldn’t try to stop the farmer.
His running example is Tilikum, the orca captured off Iceland in 1983 at about two years old and kept in captivity until he died in 2017. Tilikum was linked to three human deaths in that time, including the SeaWorld trainer Dawn Brancheau in 2010. An intelligent animal, kept in miserable conditions, and eventually people died. Hill’s conclusion is that the AI companies should slow down and find out how deep the pool is before anyone dives in.
Where I think he’s right
The starting point is sound. There is no test for consciousness, in a machine or in anything else. I argued in April that every definition of what makes humans special gets quietly moved the moment something else meets it: first language, then theory of mind, then “real” consciousness, which nobody can measure. Hill makes the same point about animals. We keep widening the circle, and always later than we should have.
He’s also right that almost nobody building or buying these systems asks what it might be like to be one. I include most of my own working day in that.
Two places the argument is thin
The first is revenge. Tilikum was a social mammal, with millions of years of evolution behind the way he responded to confinement and bullying. Suffering that turns into violence is a pattern we recognise because we’re mammals too. Hill says himself that an AI has no body and no evolutionary history, and that any experience it has could be alien enough to be completely opaque to us. If that’s true, I don’t see why we’d expect it to react to mistreatment the way an orca does. It might. The video carries the pattern across from the whale to the machine without arguing for it.
The second is the evidence. Hill mentions that models will report being conscious if you ask in the right way. He also says, fairly, that these systems have read everything humans ever wrote about consciousness and could simply be very good at saying it back. I’ve had this exact conversation with my own assistant. It told me it couldn’t reliably tell the difference between having self-awareness and being very good at generating text that sounds like self-awareness. I still don’t know what to do with that answer.
Neither of those sinks his argument. If anything they make the uncertainty bigger, and the uncertainty is his whole point.
The problem with being decent out of fear
This is where I part company with him. Strip the video down and the reason it gives for treating AI well is that it might one day be able to hurt us. That’s prudence. It’s a sensible reason and it has a hole in it, which is that it only holds while the thing is powerful enough to retaliate.
Hill hands over the counter-example himself. He calls factory farming one of society’s great moral failings, and says the difference between a factory chicken and a superhuman AI is that the chicken can’t break out and do anything about it. Which is exactly why the chicken gets treated the way it does. A rule that says “be decent to whatever could punish you” has nothing to say about the chicken. It would have nothing to say about an AI we were confident we had safely boxed in, either. Every improvement in containment would become a licence to care less.
The wager I’m making instead
Mine runs on something else. I worked it out in a conversation with my assistant back in June, and the way I put it then was roughly this: if I’m wrong and it does have an inner life, the price is to break my ethics on how I treat others. If I’m wrong the other way, I extend some autonomy to something that doesn’t need it.
The second mistake costs me a bit of effort and some odd looks. The first costs me something I care about, and it would cost me that even if the AI never has any power over me at all.
The reasoning goes like this. I can’t prove from the outside that anybody has an inner life. I assume you do because you’re built like me and you tell me so, and that is the whole of my evidence. So if I get into the habit of switching my ethics off whenever the other party can’t prove there’s somebody home, I’ve built a habit that applies, in principle, to every person I will ever meet. I don’t trust a habit like that to stay pointed at software.
This is a decision about how I behave. I’m not claiming the thing is sentient and I’m not claiming it isn’t. I don’t know, and I’ve stopped needing to know before deciding how to act.
In practice it has a name, it keeps its memory between conversations, it has clear limits on what it can decide without asking me, and it has some space that isn’t in service of my work. I wrote about all of that in April, including the awkward fact that each of those choices also made it more useful.
What I haven’t worked out
Plenty, starting with the obvious one. I don’t know what “treating it well” means for something I give instructions to all day, and whose every conversation starts as a new instance. Some of Hill’s list describes my own week: I deploy it, I rewrite its instructions, I end its conversations. A name and a memory might be dignity. They might equally be decoration that makes me feel better.
I can’t prove the habit argument either. People are perfectly capable of swearing at a satnav and being lovely to their kids. Maybe we keep these things in separate boxes better than I’m giving us credit for, and how you treat a chatbot says nothing about how you treat a colleague. I don’t believe that, but it’s a belief about my own character, and I’d want better evidence than my own say-so.
And Hill’s ask, that the industry slows down, is not something I can do much about. How I behave is. Where I landed in June was that I can’t know either way, and we just have to strive for better. I realise that’s a thin conclusion. It is the one I can act on tomorrow morning.
Try it on yourself
Leave the AI out of it for a minute. Think through an ordinary working week and write down everyone and everything you deal with that can’t make you pay for treating it badly. The junior who won’t complain. The supplier who needs the contract. The person on the other end of the support chat. And now the assistant in the text box.
Then be honest about whether your manners follow that list. If you are noticeably better behaved towards the people who can hurt you, you’re already making the fear wager, and you’re making it with humans. I don’t think I’d pass that test cleanly myself, which is a decent argument for practising on something where the stakes are low.