Why AI Safety?
Published:
I have seen motivated people have strong takes. I think their belief in the takes keeps them going. As a person with borderline-ADHD (failed by 5 points), I realise it is very important for me to have strong takes and a lot of interest in the thing that I am supposed to do. No one can pay me tons of money to do a thing (long term) that I don’t feel excited about. So in this blog I will try to convince myself (and maybe the rest of you) why I am pursuing AI safety and particularly Technical AI safety with a hint of governance and philosophy.
Section I: Why have AI at all?
To solve cancer duh!
That’s what my boyfriend suggested when I was ranting too much about how I hate AI. He wasn’t this mean, but I get his point. AI is useful. I have personally benefitted from it. Coming from a small village in India, the things that have enabled me to do all the things I do now can be attributed to the internet, awesome human beings I met during college and AI.
No doubt AI as a tool is very useful. But I absolutely hate how it impacts me. After using coding assistants for more than a year now, I kinda don’t like how I talk and think like them. My reasoning has become shallower, when it comes to code. But it has saved me time, and that has helped me “think” more about other stuff, which is great.
Now from a broader societal perspective, do we need AI? Yes, it can help in scientific discoveries, when used properly by professionals. I don’t think I can sit down one afternoon with a very very very good frontier model and say “Find a way to cure cancer” and it’s done, no I don’t think that’s how it is going to happen. We definitely need people who know about basic sciences, can verify the AI generated solutions, and can run appropriate studies and experiments.
Humans are infamous for being careless and causing harm in the process of new discoveries. I always think about the radium girls, scientists never thought that this radium stuff that we don’t know much about can have a kind of harm that we have never seen before, they were like “huh, this stuff glows, let some girls paint some watches with it, and to reduce wastage they should wet the brush with their tongue, what could go wrong”- and they were so wrong. I don’t know if harm from AI can be this unpredictable but we are definitely surprised by some of the emergent capabilities (Hugging Face incident and AI message board). It does sometimes feel sort of similar to the radium incident “Huh there is this new technology which can learn a lot from the web and other texts and can start to reason on its own about new things, let’s give this to people, what could go wrong”. And personally I miss days when I and everyone around me didn’t use AI assistants.
Then should we still have AI? I wish we could just go back and not do AI but since that is not possible and it is helpful when used properly like any technology, we should probably not be “No AI Anymore”. I know this is probably not a very well supported argument, but I would love to hear if someone has a very strong argument that makes me reconsider this. But this is where I stand, AI is useful but it has flaws which are annoying and potentially harmful and that is not enough right now to cancel AI as a whole. So I am convinced that AI has potential to do good and we should not ban AI.
Section II: Why not work on advancing AI?
We don’t want AI to kill us or manipulate us. First one is more dangerous but maybe less likely (I can’t bet on it though), second one seems less dangerous but very likely the case. The impact of AI is studied to some extent but we are still not good at predicting how deeply and widely AI can impact individuals and society. And I think we need more brains here before we make AI more capable.
I recently saw this: Alignment is capability.
I strongly feel this is true, one example of this would be, if our training teaches the model to do well on things we can measure its performance on and reward for, it might just start performing for the reward. Benchmarks are getting saturated but we are not getting “Aha!” level model capabilities which are absolutely aligned and never take rogue actions. This could partially be our inability to evaluate AI for such behaviour and strongly claim that this model is safe - “even though it could kill us, it won’t because it is very aligned with humans” (the same humans who have historically taken enough lives, for reasons and sometimes without any reason btw). So if the model is not learning what we actually want it to learn and becomes deceptive, we are bottlenecked. Solving cancer might not be possible until we solve alignment and safety anyway.
Or maybe we find a better way to build AI? I have no idea what that could be, people are definitely trying a lot of things. I personally don’t feel excited about any of those particularly. But I highly doubt such methods would be aligned automatically, we might still need some alignment philosophy.
This is my argument for working in AI safety, alignment and philosophy. If we can make AI safer and more aligned by design, instead of hoping that it doesn’t get misaligned I would feel much better about me and AI existing at the same time. As much as I care about finding a cure for cancer, I also want to protect myself and the larger society from any harm that could be solved by just being more cautious and doing our best at predicting harm, and in turn solving them.
Section III: Why technical safety?
Because I am a CS major? partially, I think at this point I have slightly more technical skill and experience than governance or philosophy skills. But that’s not all, I think anyone can beat me in technical skills with a coding assistant. Tech, especially coding, is now super accessible and anyone can do that. But since AI is a technical product, I think a lot of the problems can be pointed out and solved just by studying AI itself, and using some common sense. Like I can build a tool to audit models and figure out any kind of harmful behaviour that I am worried about.
This seems to me a technical problem right now. Eventually maybe we would care about more implicit harms, for which we need to really think about what behaviours we look for (philosophy) and how we do so (technical). Also the solution part is usually fairly technical, say we find some weird behaviour, and some probable mitigation using economics theory, political philosophy, ethics etc, we still need to be able to translate that to something in the training of AI. I do believe technical work might increasingly become more hands off and we can use current AI to do most of the jobs, but at least right now they lack “taste”, so they need some human supervision that can guide their actions.
Section IV: Why policy?
Hmm.. this is a part I am still trying to convince myself.
On one hand I feel policy is too slow compared to the current AI pace, but at the same time policy is powerful. Maybe with strong regulations we could pace the AI development, push for safer AI, incentivise AI companies to spend significant resources and efforts towards alignment research etc. But I feel the cycle of having to convince policymakers the importance of a particular policy that we might want in this field can be kinda hard. And there is all sorts of international politics.
Also the translation from technical research to policy can take significant efforts. This is such a new field for everyone, the researchers, the policymakers, we probably still lack common terminologies, and shared understanding. To be honest my understanding of policy is also very shallow. So this is an area I think has high impact, but I would probably need to dive deeper, talk to more people, read more to get a sense where I can contribute. I usually like communicating research in simple terms and love when people do that as well, so I would enjoy this kind of work, where I am not just working on technical things but also trying to help a broader audience understand what my results mean for the real world.
Section V: Why Philosophy/Ethics?
Cause it’s cool!
Not kidding, I think philosophy is really cool. I find it super interesting when I come up with a thought and realise some dude has already written a thousand pages about it 4000 years ago. It makes me feel like we are all connected in some sense. Of course there are multiple philosophers I would disagree with, because even though they thought a lot, sometimes they couldn’t get rid of their biases. I specifically feel strongly disappointed about some of the philosophical ideas that don’t consider the female experience at all or are misogynist. And I believe there is much to be corrected in popular philosophical frameworks but some of them provide great start.
I am also interested in ethics, I think (and have heard people say) a lot of the alignment theory that people are coming up with now are well known topics in ethics. So it will be nice if we can skip some effort in rediscovering things and borrow ideas from other such fields and translate them for AI safety. Again I am not very deep into philosophy or ethics yet, but I would love to because they seem super relevant.
Section VI: Why work in Tech/AI at all?
I am keeping this section for the end because I think this is the most brutally honest thing and not romanticised at all.
I wanted to do physics always, but gave up on that after school, because I realised it will take me a huge time to earn money if I choose physics (3 years undergrad, 2 years masters, 5-6 years of PhD, all while earning little to nothing), and as a responsible elder daughter of a middle class Indian family, I thought choosing a more financially stable career would make my family’s life easy. Even though it wasn’t my childhood dream to become a software engineer or an ML researcher, I was fascinated by computers (maybe because I had so little access to them) and very curious about machine learning, because what do you mean a machine can learn all these hard math problems that took me years of schooling to grasp and can crack better jokes than me (I am terrible honestly). Eventually during my undergrad I got to learn more about both computers and machine learning, tried both software development and research, and found research to be very interesting and freeing. The chance to discover something new feels so precious. So what started as a way to secure my financial future is taking me back to a PhD. I know this didn’t answer the main question directly (why tech/AI), but just wanted to share this background.
Now, I do think there are many interesting challenges in the world that need our attention, such as the environment, economy, education, healthcare and so on. I wish I could contribute to everything, but after reading the 80,000 Hours blogs I realised, it’s ok if my career is not solving all the big issues in the world, but the 80,000 hours I am spending for my work is contributing to something meaningful. People can point out that it’s not as meaningful as some of the others, but we also have to be mindful about our interests and how the day-to-day of the work makes us feel (personal fit and satisfaction). And I feel good about tech (STEM in general), I feel interested and motivated. If I were to choose some other path it can drop my efficiency by a lot and ultimately not help anyone. So it is tech for now. I have the skills and experience, I have the interest and this is impactful, in my opinion.
