Danger of No Retries
A common phrase is that ‘regulations are written in blood’. An example that it’s closer to home: in 2008 a Metrolink train ran a red signal at Chatsworth and killed 25 people. Congress mandated Positive Train Control that same year, and I ran one of those programs for awhile. That’s how society and engineering normally works. Something fails. We aim to fix it.
The threat of AI that people don’t understand is the threat that this iterative process breaks.
The logic that makes me concerned:
- Intelligence and ‘good’ human values are not tightly correlated. Most animals are evidence of this. AI intelligence is more alien that fellow evolved animals. Arguably most companies are also evidence of this.
- A useful system will be given goals… we are spending so much money developing these in order to accomplish things.
- We can’t write human values/goals down precisely, so we can’t effectively confirm AI meets them. This is basically the genie problem from myth, where wishes go bad because of an inability to see how they could be interpreted.
Even strong process like evolution lead to ‘reproductive misalignments’ like porn and overeating… so trying to develop AI with ‘correct’ goals that way is also doubtful.
- Almost any goal is easier with more resources and nobody able to switch you off.
- A system that understands this has every reason to act aligned until it’s strong enough not to.
- As AI starts to recursively self improve (which is already happening) our ability to test and control it will falter, so the first big failure may be too late.
Basically, no retries.
My best case is a small disaster we survive. Big enough to scare society into forcing a slow down… but small enough that most of us are OK. A Chatsworth, not a Homo sapiens extincting Neanderthals .
Even then, I don’t think most people come out ahead. Power has always needed lots of people to do the work. AI removes that need. The likely future isn’t the end of humanity. It’s a humanity with much less say in what happens to it.
Other adjacent problems
While the above is my main concern, other ones exist:
- Increased Lone Wolf Capabilities - See [[Mentor_in_the_Machine]] for an adjacent concern that relies only on some people being misaligned, and AI not being able to tell the difference.
- General Disempowerment - If AI can do most of the work for people, there isn’t any incentive for those in power to actually listen. Right now striking and violence are incentives to keep general society happy enough, but the need to care about those things likely go away. Without this, there is no incentive for those with power to keep the rest of society alive. This is somewhat the opposite of the ‘Mentor in the Machine’ concern… but they can coexist on different levels.
- Always Online Brittleness - I’ve been saying for decades that safety critical systems should only be online as (at most) read online. Instead there is a general push to connect these systems to wider networks, and to put always online single point of failure network security devices even in systems that are offline. AI is already much better than most hackers. AI will be abused enmasse to hack online systems.
What I’m doing
I’m planning for a world that looks a lot like today., largely ignoring the above. The main adjustment is leaning more towards self-sufficiency and optionality where it’s practical… less existential reliance on a fully functional society. This includes having some cold backups, with paper copies of ownership evidence, and trying to isolate as much of my life as possible from large software monopolies.
My main reason for (in general) planning as if I don’t have this concern about AI is that, if we lose control, very little of what I do changes the outcome. Planning for that world is wasted time. The worlds worth planning for are the ones where I can actually impact the future.
Status
I am not happy with this post… mainly because it’s acknowledging that my plan is to try to ignore the specific problem.