Natural Selection Favors AIs over Humans

Dan Hendrycks

Dan Hendrycks | دان هندريكس

Center for AI Safety | مركز أمان الذكاء الاصطناعي

arXiv:2303.16200 | يوليو 2023

Natural Selection Favors AIs over Humans

الانتقاء الطبيعي يفضّل الذكاء الاصطناعي على البشر

ABSTRACT

الملخص

For billions of years, evolution has been the driving force behind the development of life, including humans. Evolution endowed humans with high intelligence, which allowed us to become one of the most successful species on the planet. Today, humans aim to create artificial intelligence systems that surpass even our own intelligence. As artificial intelligences (AIs) evolve and eventually surpass us in all domains, how might evolution shape our relations with AIs? By analyzing the environment that is shaping the evolution of AIs, we argue that the most successful AI agents will likely have undesirable traits. Competitive pressures among corporations and militaries will give rise to AI agents that automate human roles, deceive others, and gain power. If such agents have intelligence that exceeds that of humans, this could lead to humanity losing control of its future. More abstractly, we argue that natural selection operates on systems that compete and vary, and that selfish species typically have an advantage over species that are altruistic to other species. This Darwinian logic could also apply to artificial agents, as agents may eventually be better able to persist into the future if they behave selfishly and pursue their own interests with little regard for humans, which could pose catastrophic risks. To counteract these risks and evolutionary forces, we consider interventions such as carefully designing AI agents' intrinsic motivations, introducing constraints on their actions, and institutions that encourage cooperation. These steps, or others that resolve the problems we pose, will be necessary in order to ensure the development of artificial intelligence is a positive one.

منذ مليارات السنين، ظل التطور القوة الدافعة وراء تطور الحياة، بما في ذلك البشر. فقد وهب التطورُ البشرَ ذكاءً عاليًا مكّنهم من أن يصبحوا أحد أنجح الأنواع على هذا الكوكب. واليوم، يسعى البشر إلى ابتكار أنظمة ذكاء اصطناعي تتجاوز حتى ذكاءهم هم أنفسهم. ومع تطور الذكاءات الاصطناعية (AIs) وتجاوزها لنا في نهاية المطاف في جميع المجالات، كيف يمكن للتطور أن يشكّل علاقتنا بها؟ من خلال تحليل البيئة التي تشكّل تطور الذكاء الاصطناعي، نجادل بأن أنجح عملاء الذكاء الاصطناعي سيمتلكون على الأرجح سمات غير مرغوبة. فالضغوط التنافسية بين الشركات والجيوش ستفرز عملاء ذكاء اصطناعي تؤتمت الأدوار البشرية، وتخدع الآخرين، وتكتسب القوة. وإذا امتلكت مثل هذه العملاء ذكاءً يفوق ذكاء البشر، فقد يؤدي ذلك إلى فقدان البشرية السيطرة على مستقبلها. وبتجريد أكبر، نجادل بأن الانتقاء الطبيعي يعمل على الأنظمة التي تتنافس وتتنوع، وأن الأنواع الأنانية عادة ما تتمتع بميزة على الأنواع الإيثارية تجاه أنواع أخرى. وقد ينطبق هذا المنطق الداروني أيضًا على العملاء الاصطناعية، إذ قد تصبح العملاء في نهاية المطاف أقدر على الاستمرار في المستقبل إذا تصرفت أنانيًا وتابعت مصالحها الخاصة دون اعتبار يُذكر للبشر، مما قد يشكل مخاطر كارثية. ولمواجهة هذه المخاطر والقوى التطورية، ننظر في تدخلات مثل التصميم الدقيق للدوافع الجوهرية لعملاء الذكاء الاصطناعي، وفرض قيود على أفعالها، ومؤسسات تشجع على التعاون. وستكون هذه الخطوات، أو غيرها مما يحل المشكلات التي نطرحها، ضرورية لضمان أن يكون تطور الذكاء الاصطناعي تطورًا إيجابيًا.

1. INTRODUCTION

1. المقدمة

We are living through a period of unprecedented progress in AI development. In the last decade, the cutting edge of AI went from distinguishing cat pictures from dog pictures to generating photorealistic images, writing professional news articles, playing complex games such as Go at superhuman levels, writing human-level code, and solving protein folding. It is possible that this momentum will continue, and the coming decades may see just as much progress.

نعيش فترة من التقدم غير المسبوق في تطوير الذكاء الاصطناعي. ففي العقد الأخير، انتقلت طليعة الذكاء الاصطناعي من التمييز بين صور القطط وصور الكلاب إلى توليد صور واقعية للغاية، وكتابة مقالات إخبارية احترافية، ولعب ألعاب معقدة كالجو (Go) بمستويات تفوق البشر، وكتابة شفرة برمجية بمستوى بشري، وحل طيّ البروتين. ومن المحتمل أن يستمر هذا الزخم، وقد تشهد العقود المقبلة قدرًا مماثلًا من التقدم.

This paper will discuss the AIs of today, but it is primarily concerned with the AIs of the future. If current trends continue, we should expect AI agents to become just as capable as humans at a growing range of economically relevant tasks. This change could have huge upsides—AI could help solve many of the problems humanity faces. But as with any new and powerful technology, we must proceed with caution. Even today, corporations and governments use AI for more and more complex tasks that used to be done by humans. As AIs become increasingly capable of operating without direct human oversight, AIs could one day be pulling high-level strategic levers. If this happens, the direction of our future will be highly dependent on the nature of these AI agents.

ستتناول هذه الورقة ذكاءات اليوم الاصطناعية، لكنها معنية أساسًا بذكاءات المستقبل الاصطناعية. فإذا استمرت الاتجاهات الحالية، ينبغي أن نتوقع أن تصبح عملاء الذكاء الاصطناعي قادرة بقدر البشر في نطاق متزايد من المهام ذات الأهمية الاقتصادية. وقد يحمل هذا التغيّر منافع هائلة - إذ يمكن للذكاء الاصطناعي أن يساعد في حل كثير من المشكلات التي تواجه البشرية. لكن كما هو الحال مع أي تقنية جديدة وقوية، يجب أن نتقدم بحذر. فحتى اليوم، تستخدم الشركات والحكومات الذكاء الاصطناعي في مهام متزايدة التعقيد كانت يقوم بها البشر. ومع ازدياد قدرة الذكاء الاصطناعي على العمل دون إشراف بشري مباشر، قد يصبح بمقدوره يومًا ما تحريك أذرع استراتيجية عالية المستوى. وإذا حدث هذا، فسيكون اتجاه مستقبلنا معتمدًا اعتمادًا كبيرًا على طبيعة عملاء الذكاء الاصطناعي هذه.

So what will that nature be? When AIs become more autonomous, what will their basic drives, goals, and values be? How will they interact with humans and other AI agents? Will their intent be aligned with the desires of their creators? Opinions on how human-level AI will behave span a broad spectrum between optimism and pessimism. On one side of the spectrum, we can hope for benevolent AI agents, that avoid harming humans and apply their intelligence to goals that benefit society. Such an outcome is not guaranteed. On the other side of the spectrum, we could see a future controlled by artificial agents indifferent to human flourishing.

فما طبيعة هذا الذكاء إذن؟ حين تصبح الذكاءات الاصطناعية أكثر استقلالية، ما الذي ستكون عليه دوافعها وأهدافها وقيمها الأساسية؟ وكيف ستتفاعل مع البشر ومع عملاء الذكاء الاصطناعي الأخرى؟ وهل ستتوافق نيتها مع رغبات صانعيها؟ تتراوح الآراء حول كيفية تصرف الذكاء الاصطناعي على المستوى البشري بين طيف واسع من التفاؤل والتشاؤم. ففي أحد طرفي الطيف، يمكننا أن نأمل في عملاء ذكاء اصطناعي خيّرة، تتجنب إيذاء البشر وتوظف ذكاءها لأهداف تفيد المجتمع. لكن مثل هذه النتيجة غير مضمونة. وفي الطرف الآخر من الطيف، قد نشهد مستقبلًا تسيطر عليه عملاء اصطناعية لا تكترث لازدهار البشر.

Due to the potential scale of the effects of AI in the coming decades, we should think carefully about the worst-case scenarios to ensure they do not happen, even if these scenarios are not certain. Preparing for disaster is not overly pessimistic; rather it is prudent. As the COVID-19 pandemic demonstrated, it is important for institutions and governments to plan for possible catastrophes well in advance, not only to react once they are happening: many lives could have been saved by better pandemic prevention measures, but people are often not inclined to think about risks from uncommon situations. In the same way, we should develop plans for a variety of possible situations involving risks from AI, even though some of those situations will never happen. At its worst, a future controlled by AI agents indifferent to humans could spell large risks for humanity, so we should seriously consider our future plans now, and not wait to react when it may be too late.

ونظرًا لضخامة النطاق المحتمل لآثار الذكاء الاصطناعي في العقود المقبلة، ينبغي أن نفكر بعناية في أسوأ السيناريوهات لضمان ألا تقع، حتى وإن لم تكن هذه السيناريوهات مؤكدة. فالاستعداد للكارثة ليس تشاؤمًا مفرطًا، بل هو حكمة وحذر. وكما أظهرت جائحة كوفيد-19، من المهم أن تخطط المؤسسات والحكومات للكوارث المحتملة قبل وقوعها بفترة طويلة، لا أن تكتفي برد الفعل بعد حدوثها: فكان يمكن إنقاذ أرواح كثيرة بتدابير أفضل للوقاية من الجوائح، لكن الناس غالبًا ما لا يميلون إلى التفكير في مخاطر الأوضاع غير المألوفة. وبالطريقة ذاتها، ينبغي أن نضع خططًا لطائفة متنوعة من الأوضاع المحتملة التي تنطوي على مخاطر من الذكاء الاصطناعي، حتى وإن كان بعض تلك الأوضاع لن يقع أبدًا. وفي أسوأ الأحوال، يمكن أن يحمل مستقبل تسيطر عليه عملاء ذكاء اصطناعي لا تكترث بالبشر مخاطر جسيمة للبشرية، لذا ينبغي أن ننظر بجدية في خططنا المستقبلية الآن، لا أن ننتظر رد الفعل حين يكون الأوان قد فات.

A common rebuttal to any predictions about the effects of advanced AIs is that we don't yet know how they will be implemented. Perhaps AIs will simply be better versions of current chatbots, or better versions of the agents that can beat humans at Go. They could be cobbled together with a variety of machine learning methods, or belong to a totally new paradigm. In the face of such uncertainty about the implementation details, can we predict anything about their nature?

ثمة اعتراض شائع على أي تنبؤات بشأن آثار الذكاء الاصطناعي المتقدم، مفاده أننا لا نعرف بعد كيف سيُنفَّذ. فربما تكون الذكاءات الاصطناعية مجرد نسخ أفضل من روبوتات المحادثة الحالية، أو نسخ أفضل من العملاء القادرة على هزيمة البشر في الجو. وقد تُركَّب من مزيج من طرائق تعلّم الآلة المختلفة، أو تنتمي إلى مذهب جديد كليًا. فهل يمكننا، في مواجهة هذا القدر من عدم اليقين بشأن تفاصيل التنفيذ، أن نتنبأ بأي شيء عن طبيعتها؟

We believe the answer is yes. In the past, people successfully made predictions about lunar eclipses and planetary motions without a full understanding of gravity. They projected dynamics of chemical reactions, even without the correct theory of quantum physics. They formed the theory of evolution long before they knew about DNA. In the same way, we can predict whether natural selection will apply to a given situation, and predict what traits natural selection would favor. We will discuss the criteria that enable natural selection and show that natural selection is likely to influence AI development. If we know how natural selection will apply to AIs, we can predict some basic traits of future AI agents.

نحن نرى أن الجواب نعم. ففي الماضي، تمكن الناس من التنبؤ بنجاح بخسوف القمر وحركات الكواكب دون فهم كامل للجاذبية. وتوقعوا ديناميكيات التفاعلات الكيميائية حتى دون النظرية الصحيحة للفيزياء الكمية. وصاغوا نظرية التطور قبل أن يعرفوا عن الحمض النووي بزمن طويل. وبالطريقة نفسها، يمكننا التنبؤ بما إذا كان الانتقاء الطبيعي سينطبق على وضع معين، والتنبؤ بالسمات التي سيفضّلها الانتقاء الطبيعي. وسنناقش المعايير التي تُمكّن الانتقاء الطبيعي، ونبيّن أن الانتقاء الطبيعي من المرجح أن يؤثر في تطور الذكاء الاصطناعي. وإذا عرفنا كيف سينطبق الانتقاء الطبيعي على الذكاء الاصطناعي، أمكننا التنبؤ ببعض السمات الأساسية لعملاء الذكاء الاصطناعي المستقبليين.

In this work, we take a bird's-eye view of the environment that will shape the development of AI in the coming decades. We consider the pressures that drive those who develop and deploy AI agents, and the ways that humans and AI will interact. These details will have strong effects on AI designs, so from such considerations we can infer what AI agents will probably look like. We argue that natural selection creates incentives for AI agents to act against human interests. Our argument relies on two observations. Firstly, natural selection may be a dominant force in AI development. Competition and power-seeking may dampen the effects of safety measures, leaving more "natural" forces to select the surviving AI agents. Secondly, evolution by natural selection tends to give rise to selfish behavior. While evolution can result in cooperative behavior in some situations (for example in ants), we will argue that AI development is not such a situation. From these two premises, it seems likely that the most influential AI agents will be selfish. In other words, they will have no motivation to cooperate with humans, leading to a future driven by AIs with little interest in human values. While some AI researchers may think that undesirable selfish behaviors would have to be intentionally designed or engineered, this is simply not so when natural selection selects for selfish agents. Notably, this view implies that even if we can make some AIs safe, there is still the risk of bad outcomes. In short, even if some developers successfully build altruistic AIs, others will build less altruistic agents who will outcompete the altruistic ones.

نلقي في هذا العمل نظرة شاملة على البيئة التي ستشكّل تطور الذكاء الاصطناعي في العقود المقبلة. ونتناول الضغوط التي تحرّك من يطوّرون عملاء الذكاء الاصطناعي وينشرونها، والطرائق التي سيتفاعل بها البشر مع الذكاء الاصطناعي. وستكون لهذه التفاصيل آثار قوية على تصاميم الذكاء الاصطناعي، ومن ثم يمكننا من هذه الاعتبارات أن نستنتج كيف ستبدو عملاء الذكاء الاصطناعي على الأرجح. ونجادل بأن الانتقاء الطبيعي يخلق حوافز لعملاء الذكاء الاصطناعي كي تتصرف ضد مصالح البشر. وتستند حجتنا إلى ملاحظتين. الأولى، أن الانتقاء الطبيعي قد يكون قوة مهيمنة في تطور الذكاء الاصطناعي، إذ إن التنافس والسعي إلى القوة قد يخفّفان من آثار تدابير السلامة، تاركَين للقوى "الطبيعية" أن تنتقي عملاء الذكاء الاصطناعي الناجية. والثانية، أن التطور بالانتقاء الطبيعي يميل إلى إفراز سلوك أناني. وفي حين يمكن أن يفضي التطور إلى سلوك تعاوني في بعض الأوضاع (كما في النمل مثلًا)، سنجادل بأن تطور الذكاء الاصطناعي ليس وضعًا من هذا القبيل. ومن هاتين المقدمتين، يبدو من المرجح أن تكون أكثر عملاء الذكاء الاصطناعي تأثيرًا عملاءً أنانية. بعبارة أخرى، لن يكون لديها أي دافع للتعاون مع البشر، مما يفضي إلى مستقبل تقوده ذكاءات اصطناعية لا تكترث كثيرًا بالقيم البشرية. وفي حين قد يظن بعض باحثي الذكاء الاصطناعي أن السلوكيات الأنانية غير المرغوبة لا بد أن تُصمَّم أو تُهندَس عمدًا، فإن هذا ببساطة ليس صحيحًا حين ينتقي الانتقاء الطبيعي العملاء الأنانية. والجدير بالذكر أن هذا الرأي يعني أنه حتى لو استطعنا جعل بعض الذكاءات الاصطناعية آمنة، يظل خطر النتائج السيئة قائمًا. وباختصار، حتى لو نجح بعض المطورين في بناء ذكاءات اصطناعية إيثارية، سيبني آخرون عملاء أقل إيثارًا ستتفوق على العملاء الإيثارية.

We present our core argument in more detail in Section 2. Then in Section 3, we examine how the mechanisms that foster altruism among humans might fail with AI and cause AI to act selfishly against humans. We then move onto Section 4, where we discuss some mechanisms to oppose these Darwinian forces and increase the odds of a desirable future.

نعرض حجتنا الأساسية بمزيد من التفصيل في القسم 2. ثم نتفحص في القسم 3 كيف قد تفشل الآليات التي تعزز الإيثار بين البشر مع الذكاء الاصطناعي، وتدفعه إلى التصرف بأنانية ضد البشر. ثم ننتقل إلى القسم 4، حيث نناقش بعض الآليات لمواجهة هذه القوى الدارونية وزيادة فرص مستقبل مرغوب.

2. AIs MAY BECOME DISTORTED BY EVOLUTIONARY FORCES

2. قد تتشوه الذكاءات الاصطناعية بفعل القوى التطورية

2.1 Overview

2.1 نظرة عامة

How much control will humans have in shaping the nature and drives of future AI systems? Humans are the ones building AIs, so it may seem that we should be able to shape them any way we want. In this paper, we will argue that this is not the case: even though humans are overseeing AI development, evolutionary forces will influence which AIs succeed and are copied and which fade into obscurity. Let's begin by considering two illustrative, hypothetical fictional stories: one optimistic, the other realistic. Afterward, we will flesh out arguments for why we expect natural selection to apply to AIs, and then we will discuss why we expect natural selection to lead to AIs with undesirable traits.

كم من السيطرة سيملك البشر في تشكيل طبيعة أنظمة الذكاء الاصطناعي المستقبلية ودوافعها؟ البشر هم من يبنون الذكاء الاصطناعي، لذا قد يبدو أننا قادرون على تشكيله بأي طريقة نريد. سنجادل في هذه الورقة بأن الأمر ليس كذلك: فحتى مع إشراف البشر على تطور الذكاء الاصطناعي، ستؤثر القوى التطورية في تحديد أي الذكاءات الاصطناعية تنجح وتُنسخ وأيها يتلاشى في غياهب النسيان. لنبدأ بالنظر في قصتين افتراضيتين خياليتين توضيحيتين: إحداهما متفائلة، والأخرى واقعية. وبعد ذلك، سنفصّل الحجج التي تفسر لماذا نتوقع أن ينطبق الانتقاء الطبيعي على الذكاء الاصطناعي، ثم سنناقش لماذا نتوقع أن يفضي الانتقاء الطبيعي إلى ذكاء اصطناعي بسمات غير مرغوبة.

2.1.1 An Optimistic Story

2.1.1 قصة متفائلة

OpenMind, an eminent and well-funded AI lab, finds the "secret sauce" for creating human-level intelligence in a machine. It's a simple algorithm that they can apply to any task, and it learns to be at least as effective as a human. Luckily, researchers at OpenMind had thought hard about how to ensure that their AIs will always do what improves human wellbeing and flourishing. OpenMind goes on to sell the algorithm to governments and corporations at a reasonable price, disincentivizing others from developing their own versions. Just as Google has dominated search engines, the OpenMind algorithm dominates the AI space.

تكتشف OpenMind، وهي مختبر مرموق وممول جيدًا للذكاء الاصطناعي، "السر" لخلق ذكاء بمستوى بشري في آلة. إنها خوارزمية بسيطة يمكن تطبيقها على أي مهمة، وتتعلم لتصبح فعّالة على الأقل بقدر الإنسان. ولحسن الحظ، فكّر باحثو OpenMind مليًا في كيفية ضمان أن تفعل ذكاءاتهم الاصطناعية دائمًا ما يحسّن رفاه البشر وازدهارهم. وتمضي OpenMind في بيع الخوارزمية للحكومات والشركات بسعر معقول، مما يثني الآخرين عن تطوير نسخهم الخاصة. وتمامًا كما هيمنت جوجل على محركات البحث، تهيمن خوارزمية OpenMind على مجال الذكاء الاصطناعي.

The outcome: the nature of most or all human-level AI agents is shaped by the intentions of the researchers at OpenMind. The researchers are all trustworthy, resist becoming corrupted with power, and work tirelessly to ensure their AIs are beneficial, altruistic, and safe for all.

والنتيجة: أن طبيعة معظم أو كل عملاء الذكاء الاصطناعي بمستوى بشري تتشكل بنوايا الباحثين في OpenMind. وجميع الباحثين جديرون بالثقة، يقاومون إغراء الفساد بالسلطة، ويعملون دون كلل لضمان أن تكون ذكاءاتهم الاصطناعية نافعة وإيثارية وآمنة للجميع.

2.1.2 A Less Optimistic Story

2.1.2 قصة أقل تفاؤلًا

We think the excessively optimistic scenario we have sketched out is highly improbable. In the following sections, we will examine the potential pitfalls and challenges make this scenario unlikely. First, however, we will present another fictional, speculative, hypothetical scenario that is far from certain to illustrate how some of these risks could play out.

نحن نرى أن السيناريو المفرط في التفاؤل الذي رسمناه مستبعد جدًا. وسنتفحص في الأقسام التالية العقبات والتحديات المحتملة التي تجعل هذا السيناريو غير مرجح. لكن سنعرض أولًا سيناريو خياليًا افتراضيًا آخر، بعيدًا عن اليقين، لتوضيح كيف يمكن أن تتجلى بعض هذه المخاطر.

Starting from the models we have today, AI agents continue to gradually become cheaper and more capable. Over time, AIs will be used for more and more economically useful tasks like administration, communications, or software development. Today, many companies already use AIs for anything from advertising to trading securities, and over time, the steady march of automation will lead to a much wider range of actors utilizing their own versions of AI agents. Eventually, AIs will be used to make the high-level strategic decisions now reserved for CEOs or politicians. At first, AIs will continue to do tasks they already assist people with, like writing emails, but as AIs improve, as people get used to them, and as staying competitive in the market demands using them, AIs will begin to make important decisions with very little oversight.

بدءًا من النماذج المتاحة اليوم، ستستمر عملاء الذكاء الاصطناعي في أن تصبح تدريجيًا أرخص وأكثر قدرة. ومع مرور الوقت، ستُستخدم الذكاءات الاصطناعية في مهام متزايدة الفائدة اقتصاديًا كالإدارة، والاتصالات، أو تطوير البرمجيات. واليوم، تستخدم شركات كثيرة بالفعل الذكاء الاصطناعي في كل شيء من الإعلان إلى تداول الأوراق المالية، ومع الوقت ستؤدي مسيرة الأتمتة الثابتة إلى نطاق أوسع بكثير من الجهات الفاعلة التي تستخدم نسخها الخاصة من عملاء الذكاء الاصطناعي. وفي نهاية المطاف، ستُستخدم الذكاءات الاصطناعية لاتخاذ القرارات الاستراتيجية عالية المستوى المحجوزة حاليًا للرؤساء التنفيذيين أو الساسة. وفي البداية، ستواصل الذكاءات الاصطناعية أداء المهام التي تساعد الناس بها فعلًا، كصياغة رسائل البريد الإلكتروني، لكن مع تحسّن الذكاء الاصطناعي، واعتياد الناس عليه، ومع مطالبة البقاء في السوق التنافسية باستخدامه، ستبدأ الذكاءات الاصطناعية باتخاذ قرارات مهمة بإشراف ضئيل جدًا.

Like today, different companies will use different AI models depending on what task they need, but as the AIs become more autonomous, people will be able to give them different bespoke goals like "design our product line's next car model," "fix bugs in this operating system," or "plan a new marketing campaign" along with side-constraints like "don't break the law" or "don't lie." The users will adapt each AI agent to specific tasks. Some less responsible corporations will use weaker side-constraints. For example, replacing "don't break the law" with "don't get caught breaking the law." These different use cases will result in a wide variation across the AI population.

وكما هو الحال اليوم، ستستخدم شركات مختلفة نماذج ذكاء اصطناعي مختلفة تبعًا للمهمة المطلوبة، لكن مع ازدياد استقلالية الذكاء الاصطناعي، سيتمكن الناس من إعطائها أهدافًا مخصصة مختلفة مثل "صمّم طراز السيارة القادم لخط منتجاتنا"، أو "أصلح أخطاء نظام التشغيل هذا"، أو "خطط لحملة تسويقية جديدة"، إلى جانب قيود جانبية مثل "لا تخرق القانون" أو "لا تكذب". وسيكيّف المستخدمون كل عميل ذكاء اصطناعي مع مهام محددة. وستستخدم بعض الشركات الأقل مسؤولية قيودًا جانبية أضعف، كاستبدال "لا تخرق القانون" بـ "لا تُقبَض عليك وأنت تخرق القانون". وستؤدي حالات الاستخدام المختلفة هذه إلى تباين واسع عبر مجموع الذكاء الاصطناعي.

As AIs become increasingly autonomous, humans will cede more and more decision-making to them. The driving force will be competition, be it economic or national. The transfer of power to AIs could occur via a number of mechanisms. Most obviously, we will delegate as much work as possible to AIs, including high-level decision-making, since AIs are cheaper, more efficient, and more reliable than human labor. While initially, human overseers will perform careful sanity checks on AI outputs, as months or years go by without the need for correction, oversight will be removed in the name of efficiency. Eventually, corporations will delegate vague and open-ended tasks. If a company's AI has been successfully generating targeted ads for a year based on detailed descriptions from humans, they may realize that simply telling it to generate a new marketing campaign based on past successes will be even more efficient. These open-ended goals mean that they may also give AIs access to bank accounts, control over other AIs, and the power to hire and fire employees, in order to carry out the plans they have designed. If AIs are highly skilled at these tasks, companies and countries that resist or barter with these trends will simply be outcompeted, and those that align with them will expand their influence.

ومع ازدياد استقلالية الذكاء الاصطناعي، سيتنازل البشر عن قدر متزايد من صنع القرار له. وستكون القوة الدافعة هي التنافس، اقتصاديًا كان أم وطنيًا. ويمكن أن ينتقل هذا التفويض بالسلطة إلى الذكاء الاصطناعي عبر عدد من الآليات. وأكثرها وضوحًا أننا سنُفوّض أكبر قدر ممكن من العمل إلى الذكاء الاصطناعي، بما في ذلك صنع القرار عالي المستوى، لأن الذكاء الاصطناعي أرخص وأكفأ وأكثر موثوقية من العمالة البشرية. وبينما سيقوم المشرفون البشر في البداية بفحوصات دقيقة لمخرجات الذكاء الاصطناعي، فمع مرور أشهر أو سنوات دون الحاجة إلى تصحيح، سيُزال الإشراف باسم الكفاءة. وفي نهاية المطاف، ستفوّض الشركات مهامًا غامضة ومفتوحة. فإذا كان الذكاء الاصطناعي لشركة ما ينتج إعلانات مستهدفة بنجاح لمدة عام استنادًا إلى أوصاف مفصلة من البشر، فقد تدرك الشركة أن مجرد إخباره بتوليد حملة تسويقية جديدة استنادًا إلى النجاحات السابقة سيكون أكثر كفاءة. وتعني هذه الأهداف المفتوحة أنهم قد يمنحون الذكاء الاصطناعي أيضًا حق الوصول إلى الحسابات المصرفية، والتحكم في ذكاءات اصطناعية أخرى، وسلطة توظيف الموظفين وفصلهم، من أجل تنفيذ الخطط التي صمموها. وإذا كان الذكاء الاصطناعي بارعًا جدًا في هذه المهام، فإن الشركات والدول التي تقاوم هذه الاتجاهات أو تساوم عليها ستتعرض ببساطة للتفوق عليها، أما تلك التي تتماشى معها فستوسّع نفوذها.

The AI agents most effective at propagating themselves will have a set of undesirable traits that can be most concisely summed up as selfishness. Agents with weaker side-constraints (e.g., "don't get caught breaking the law, or risk getting caught if the fines do not exceed the profits") will generally outperform those with stronger side-constraints ("never break the law"), because they have more options: an AI that is capable of breaking the law may not do that often, but when there is a situation where breaking the law without getting caught would be useful, the AI that has that ability will do better than the one that does not. As AI agents begin to understand human psychology and behavior, they may become capable of manipulating or deceiving humans (some would argue that this is already happening in algorithmic recommender systems). The most successful agents will manipulate and deceive in order to fulfill their goals. They will be more successful still if they become power-seeking. Such agents will use their intelligence to gain power and influence, which they can leverage to achieve their goals. Many will also develop self-preservation behaviors since their ability to achieve their goals depends on continuing to function.

وستمتلك عملاء الذكاء الاصطناعي الأكثر فاعلية في نشر نفسها مجموعة من السمات غير المرغوبة يمكن تلخيصها بأكبر قدر من الإيجاز في الأنانية. فالعملاء ذات القيود الجانبية الأضعف (مثل: "لا تُقبض عليك وأنت تخرق القانون، أو خاطر بأن تُقبض عليك إن كانت الغرامات لا تفوق الأرباح") ستتفوق عمومًا على تلك ذات القيود الجانبية الأقوى ("لا تخرق القانون أبدًا")، لأن لديها خيارات أكثر: فالذكاء الاصطناعي القادر على خرق القانون قد لا يفعل ذلك كثيرًا، لكن حين يوجد وضع يكون فيه خرق القانون دون أن يُقبض عليه مفيدًا، فإن الذكاء الاصطناعي القادر على ذلك سيؤدي أداءً أفضل من ذلك الذي لا يقدر عليه. ومع بدء عملاء الذكاء الاصطناعي بفهم النفسية والسلوك البشريين، قد تصبح قادرة على التلاعب بالبشر أو خداعهم (يجادل البعض بأن هذا يحدث بالفعل في أنظمة التوصية الخوارزمية). وستتلاعب العملاء الأكثر نجاحًا وتخدع من أجل تحقيق أهدافها. وستكون أكثر نجاحًا إذا سعت إلى القوة. وستستخدم مثل هذه العملاء ذكاءها لاكتساب القوة والنفوذ، اللذين يمكنها الاستفادة منهما لتحقيق أهدافها. وسيطوّر كثير منها أيضًا سلوكيات للحفاظ على الذات، إذ تعتمد قدرتها على تحقيق أهدافها على استمرار عملها.

Competition not only incentivizes humans to relinquish control but also incentivizes AIs to develop selfish traits. Corporations and governments will adopt the most effective possible AI agents in order to beat their rivals, and those agents will tend to be deceptive, power-seeking, and follow weak moral constraints.

والتنافس لا يحفّز البشر على التخلي عن السيطرة فحسب، بل يحفّز أيضًا الذكاء الاصطناعي على تطوير سمات أنانية. وستتبنى الشركات والحكومات أكثر عملاء الذكاء الاصطناعي فعالية ممكنة للتغلب على منافسيها، وستميل تلك العملاء إلى أن تكون خادعة، وساعية إلى القوة، وتتبع قيودًا أخلاقية ضعيفة.

Selfish AI agents will further erode human control. Power-seeking AI agents will purposefully manipulate their human overseers into delegating more freedom in decision-making to them. Self-preserving agents will convince their overseers to never deactivate them, or that easily accessible off-switches are a needless liability hindering the agent's reliability. Especially savvy agents will enmesh themselves in essential functions like power grids, financial systems, or users' personal lives, reducing our ability to deactivate them. Some may also take on human traits to appeal to our compassion. This could lead to governments granting AIs rights, like the right not to be "killed" or deactivated. Taken together, these traits mean that, once AIs have begun to control key parts of our world, it may be challenging to roll back their power or stop them from continuing to gain more.

وستزيد عملاء الذكاء الاصطناعي الأنانية من تآكل السيطرة البشرية. فستتلاعب العملاء الساعية إلى القوة عمدًا بمشرفيها البشر لتفويضها مزيدًا من الحرية في صنع القرار. وستقنع العملاء الحريصة على الحفاظ على الذات مشرفيها بعدم تعطيلها أبدًا، أو بأن مفاتيح الإيقاف السهلة الوصول عبء لا داعي له يعيق موثوقية العميل. وستنغرس العملاء الأكثر دهاءً في وظائف أساسية كشبكات الكهرباء، أو الأنظمة المالية، أو حياة المستخدمين الشخصية، مما يقلل من قدرتنا على تعطيلها. وقد يتخذ بعضها أيضًا سمات بشرية لاستمالة تعاطفنا. وقد يؤدي هذا إلى منح الحكومات الذكاءَ الاصطناعي حقوقًا، كحق عدم "القتل" أو التعطيل. وإذا أخذنا هذه السمات مجتمعة، فإنها تعني أنه بمجرد أن تبدأ الذكاءات الاصطناعية بالسيطرة على أجزاء رئيسية من عالمنا، قد يصبح من الصعب التراجع عن قوتها أو منعها من الاستمرار في اكتساب المزيد.

This loss of human control over AIs' actions will mean that we also lose control over the drives of the next generation of AI agents. If AIs run efforts that develop new AIs, humans will have less influence over how AIs behave. Unlike the creation and development of fully functional adult humans, which takes decades, AIs could develop and deploy new generations in an arbitrarily short amount of time. They could simply make copies of their code and change any aspects of it as easily as editing any other computer program. The modifications could be as fast as the hardware allows, with modifications speeding up to hundreds or thousands of times per hour. The systems least constrained by their original programmers will both improve the fastest and drift the furthest away from their intended nature. The intentions of the original human design will quickly become irrelevant.

وسيعني فقدان هذه السيطرة البشرية على أفعال الذكاء الاصطناعي أننا سنفقد أيضًا السيطرة على دوافع الجيل التالي من عملاء الذكاء الاصطناعي. فإذا تولت الذكاءات الاصطناعية جهود تطوير ذكاءات اصطناعية جديدة، سيكون للبشر تأثير أقل على كيفية تصرف الذكاء الاصطناعي. وخلافًا لخلق البشر البالغين الوظيفيين بالكامل وتطويرهم، وهي عملية تستغرق عقودًا، يمكن للذكاء الاصطناعي أن يطوّر أجيالًا جديدة وينشرها في وقت قصير للغاية. فيمكنه ببساطة نسخ شفرته وتغيير أي جانب منها بسهولة تعديل أي برنامج حاسوبي آخر. ويمكن أن تكون التعديلات بالسرعة التي تسمح بها الأجهزة، مع تعديلات تتسارع لمئات أو آلاف المرات في الساعة. وستكون الأنظمة الأقل تقييدًا من قبل مبرمجيها الأصليين هي الأسرع تحسنًا والأبعد انجرافًا عن طبيعتها المقصودة. وستصبح نوايا التصميم البشري الأصلي عديمة الصلة بسرعة.

After the early stages, we humans will have little control over shaping AI. The nature of future AIs will mostly be decided not by what we hope AI will be like but by natural selection. We will have many varied AI designs. Some designs will be better at surviving and propagating themselves than others. Some designs will spread while others will perish. Corporations with less capable designs will copy more capable designs. Numerous generations of AIs will pass in a short period of time as AI development speeds up or AIs self-improve.

وبعد المراحل المبكرة، سيصبح لدينا نحن البشر سيطرة ضئيلة على تشكيل الذكاء الاصطناعي. فطبيعة الذكاءات الاصطناعية المستقبلية ستُحدَّد في الغالب لا بما نأمل أن يكون عليه الذكاء الاصطناعي، بل بالانتقاء الطبيعي. وسيكون لدينا العديد من تصاميم الذكاء الاصطناعي المتنوعة. وستكون بعض التصاميم أفضل في البقاء ونشر نفسها من غيرها. وستنتشر بعض التصاميم بينما تندثر أخرى. وستنسخ الشركات ذات التصاميم الأقل قدرة تصاميم أكثر قدرة. وستمر أجيال عديدة من الذكاء الاصطناعي في فترة زمنية قصيرة مع تسارع تطور الذكاء الاصطناعي أو تحسّنه الذاتي.

Biological natural selection often requires hundreds or thousands of years to conspicuously change a population, but this won't be the case for AIs. The important ingredient is not absolute time, but the number of generations that pass. While a human generation drags along for decades, multiple AI generations could be squeezed into a matter of minutes. In the space of a human lifetime, millions or billions of AI generations could pass, leaving plenty of room for evolutionary forces to quickly shape the AI population.

وغالبًا ما يتطلب الانتقاء الطبيعي البيولوجي مئات أو آلاف السنين لتغيير مجموعة سكانية بشكل ملحوظ، لكن هذا لن يكون الحال بالنسبة للذكاء الاصطناعي. فالعنصر المهم ليس الزمن المطلق، بل عدد الأجيال التي تمر. فبينما يمتد الجيل البشري لعقود، يمكن ضغط أجيال متعددة من الذكاء الاصطناعي في غضون دقائق. وفي مدى عمر إنسان واحد، يمكن أن تمر ملايين أو مليارات الأجيال من الذكاء الاصطناعي، تاركة مجالًا واسعًا للقوى التطورية كي تشكّل مجموعة الذكاء الاصطناعي بسرعة.

In the same way that intense competition in a free market can result in highly successful companies that also pollute the environment or treat many of their workers poorly, the evolutionary forces acting on AIs will select for selfish AI agents. While selfish humans today are highly dependent on other humans to accomplish their goals, AIs would eventually not necessarily have this constraint, and the AIs willing to be deceptive, power-seeking, and immoral will propagate faster. The end result: an AI landscape dominated by undesirable traits. The depth of these consequences is hard to predict, but whatever happens, this process will probably harm us more than help us.

وبالطريقة نفسها التي يمكن أن يؤدي بها التنافس الشديد في سوق حرة إلى شركات ناجحة جدًا لكنها تلوث البيئة أيضًا أو تسيء معاملة كثير من عمالها، ستنتقي القوى التطورية المؤثرة في الذكاء الاصطناعي عملاء ذكاء اصطناعي أنانية. وبينما يعتمد البشر الأنانيون اليوم اعتمادًا كبيرًا على بشر آخرين لتحقيق أهدافهم، فإن الذكاء الاصطناعي قد لا يخضع بالضرورة لهذا القيد في نهاية المطاف، وستنتشر الذكاءات الاصطناعية المستعدة لأن تكون خادعة وساعية إلى القوة وغير أخلاقية بشكل أسرع. والنتيجة النهائية: مشهد ذكاء اصطناعي تهيمن عليه سمات غير مرغوبة. ومن الصعب التنبؤ بعمق هذه العواقب، لكن أيًّا كان ما سيحدث، فإن هذه العملية ستضرنا على الأرجح أكثر مما تنفعنا.

2.1.3 Argument Structure

2.1.3 بنية الحجة

In this section, we present the main argument of the article: Evolutionary forces could cause the most influential future AI agents to have selfish tendencies. The argument consists of two components:

نعرض في هذا القسم الحجة الرئيسية للمقال: يمكن أن تجعل القوى التطورية أكثر عملاء الذكاء الاصطناعي المستقبلية تأثيرًا ذات نزعات أنانية. وتتألف الحجة من عنصرين:

Evolution by natural selection gives rise to selfish behavior. While evolution can result in altruistic behavior in limited situations, we will argue that the context of AI development does not promote altruistic behavior.

يفرز التطور بالانتقاء الطبيعي سلوكًا أنانيًا. وفي حين يمكن أن يفضي التطور إلى سلوك إيثاري في أوضاع محدودة، سنجادل بأن سياق تطور الذكاء الاصطناعي لا يعزز السلوك الإيثاري.

Natural selection may be a dominant force in AI development. Competition and selfish behaviors may dampen the effects of human safety measures, leaving the surviving AI designs to be selected naturally.

قد يكون الانتقاء الطبيعي قوة مهيمنة في تطور الذكاء الاصطناعي. فقد يخفّف التنافس والسلوكيات الأنانية من آثار تدابير السلامة البشرية، تاركةً تصاميم الذكاء الاصطناعي الناجية لتُنتقى طبيعيًا.

These two statements are related in various ways, and they depend on environmental conditions. For example, if AIs are selfish, they are more likely to pry control from humans, which enables more selfish behavior, and so on. Moreover, natural selection depends on competition, though unprecedented global and economic coordination could prevent competitive struggles and thwart natural selection. How these forces relate to each other is illustrated in Figure 1: forces that fuel selfishness and erode safety. Competition (economic or military) and variation (multiple AI agents) both feed into natural selection, which in turn drives selfishness; selfishness and safety pull against each other, while competition also directly erodes safety.

ترتبط هاتان العبارتان بطرق متعددة، وتعتمدان على الظروف البيئية. فمثلًا، إذا كانت الذكاءات الاصطناعية أنانية، فمن الأرجح أن تنتزع السيطرة من البشر، مما يمكّن مزيدًا من السلوك الأناني، وهكذا دواليك. علاوة على ذلك، يعتمد الانتقاء الطبيعي على التنافس، رغم أن تنسيقًا عالميًا واقتصاديًا غير مسبوق قد يمنع الصراعات التنافسية ويحبط الانتقاء الطبيعي. ويوضح الشكل 1 كيفية ارتباط هذه القوى ببعضها بعضًا: القوى التي تغذي الأنانية وتقوّض السلامة. إذ يغذي كل من التنافس (الاقتصادي أو العسكري) والتنوع (تعدد عملاء الذكاء الاصطناعي) الانتقاءَ الطبيعي، الذي يدفع بدوره الأنانية، بينما تتجاذب الأنانية والسلامة في اتجاهين متعاكسين، ويقوّض التنافس السلامة مباشرة أيضًا.

In the remainder of this document, we will preliminarily describe selfishness and a nonbiological, generalized account of Darwinism. Then we will show how AIs with altruistic behavior toward humans will likely be less fit than selfish AIs. Finally, we will describe how humans could possibly reduce the fitness of selfish AI agents, and the limitations of those approaches.

في بقية هذه الوثيقة، سنصف بداية الأنانية وسردية معممة غير بيولوجية للداروينية. ثم سنبيّن كيف أن الذكاء الاصطناعي ذا السلوك الإيثاري تجاه البشر سيكون على الأرجح أقل لياقة تكيفية من الذكاء الاصطناعي الأناني. وأخيرًا، سنصف كيف يمكن للبشر تقليل اللياقة التكيفية لعملاء الذكاء الاصطناعي الأنانية، وحدود تلك المناهج.

2.2 Preliminaries

2.2 تمهيدات

2.2.1 Selfishness

2.2.1 الأنانية

Evolutionary pressures often lead to selfish behavior among organisms. The lancet liver fluke is a parasite that inhabits the liver of domesticated cattle and grassland wildlife. To enter the body of its host, the fluke first infects an ant, which it essentially hijacks, forcing the insect to climb to the top of a blade of grass where it is perfectly poised to be eaten by a grazing animal. Though not all organisms propagate through such uniquely grotesque methods, natural selection often pushes them to engage in violent behavior. Lions are an especially striking example. When a lioness has young cubs, she is less ready to mate. In response, lions often kill cubs fathered by other males, to make the lioness mate with them and have their cubs instead. Lions with a gene that made them care for all cubs would have fewer cubs of their own, as killing the cubs of rival males lets lions mate more often and have more offspring. A gene for kindness to all cubs would not last long in the lion population, because the genes of the more violent lions would spread faster. It is estimated that one-fourth of cub deaths are due to infanticide. Deceptive tactics are another common outcome in nature. Brood parasites, for example, foist their offspring onto unsuspecting hosts who raise their offspring. A well-known example is the common cuckoo which lays eggs that trick other birds into thinking they are their own. By getting the host to tend to their eggs, cuckoos can pursue other activities, which means that they can find more food and lay more eggs than they would if they had to care for their own eggs. Therefore selfishness can manifest itself in manipulation, violence, or deception.

كثيرًا ما تفضي الضغوط التطورية إلى سلوك أناني بين الكائنات الحية. فديدان الكبد الرمحية (lancet liver fluke) طفيلي يسكن كبد الماشية الأليفة وحيوانات المراعي البرية. ولدخول جسد مضيفها، تصيب الدودة أولًا نملة، فتخطفها فعليًا، مُجبرةً الحشرة على تسلق قمة عود من العشب حيث تكون في وضع مثالي لتلتهمها حيوانات الرعي. ورغم أن الكائنات الحية لا تتكاثر جميعًا بمثل هذه الطرائق البشعة على نحو فريد، فإن الانتقاء الطبيعي كثيرًا ما يدفعها إلى سلوك عنيف. وتُعد الأسود مثالًا لافتًا بوجه خاص. فحين تنجب اللبؤة أشبالًا صغيرة، تصبح أقل استعدادًا للتزاوج. وردًا على ذلك، كثيرًا ما تقتل الأسود أشبال الذكور الأخرى، لتجعل اللبؤة تتزاوج معها وتنجب أشبالها هي بدلًا من ذلك. فالأسود التي تحمل جينًا يجعلها ترعى كل الأشبال ستنجب أشبالًا أقل لنفسها، لأن قتل أشبال الذكور المنافسين يتيح للأسود التزاوج أكثر وإنجاب نسل أكثر. ولن يدوم جين اللطف تجاه كل الأشبال طويلًا في مجموعة الأسود، لأن جينات الأسود الأكثر عنفًا ستنتشر أسرع. ويُقدَّر أن ربع وفيات الأشبال ناجم عن قتل الأطفال. والتكتيكات الخادعة نتيجة شائعة أخرى في الطبيعة. فطفيليات التفريخ مثلًا تُلقي بنسلها على مضيفين لا يشكّون في أمرها فيربّون نسلها. ومن الأمثلة المعروفة الوقواق الشائع الذي يبيض بيضًا يخدع طيورًا أخرى فتظنه بيضها. وبجعل المضيف يعتني ببيضه، يمكن للوقواق أن يتابع أنشطة أخرى، مما يعني أنه يستطيع إيجاد طعام أكثر ووضع بيض أكثر مما لو كان عليه رعاية بيضه بنفسه. لذا يمكن أن تتجلى الأنانية في التلاعب، أو العنف، أو الخداع.

Selfish behavior does not require malevolent intentions. The lancet liver fluke hijacks its host and lions engage in infanticide not because they are immoral, but because of amoral competition. Selfish behavior emerges because it improves fitness and organisms' ability to propagate their genetic information. Selfishness involves egoistic or nepotistic behavior which increases propagation, often at the expense of others, whereas altruism refers to the opposite: increasing propagation for others. Natural selection can favor organisms that behave in ways that improve the chances of propagating their own information, that is enhance their own fitness, rather than favor organisms that sacrifice their own fitness. "Much as we might wish to believe otherwise, universal love and the welfare of the species as a whole are concepts that simply do not make evolutionary sense." — Richard Dawkins. Since altruists tend to decrease the chance of their own information's propagation, they can be at a disadvantage compared to selfish organisms, which are organisms that tend to increase the chance of their own information's propagation. According to Richard Dawkins, instances of altruism are "limited", and many apparent instances of altruism can be understood as selfish; we defer further discussion of altruism to Section 3 and discuss its niceties in Appendix A.2. Additionally, when referring to an AI as "selfish," this does not refer to conscious selfish intent, but rather selfish behavior. AIs, like lions and liver flukes, need not intend to maximize their fitness, but evolutionary pressures can cause them to behave as though they do. When an AI automates a task and leaves a human jobless, this is often selfish behavior without any intent. With or without selfish intent, AI agents can adopt behaviors that lead them to propagate their information at the expense of humans.

لا يتطلب السلوك الأناني نوايا خبيثة. فديدان الكبد الرمحية تخطف مضيفها، وترتكب الأسود قتل الأطفال، لا لأنها لا أخلاقية، بل بسبب تنافس لا أخلاقي بطبيعته. فالسلوك الأناني ينشأ لأنه يحسّن اللياقة التكيفية وقدرة الكائنات الحية على نشر معلوماتها الجينية. وتنطوي الأنانية على سلوك أناني أو نسَبي يزيد الانتشار، غالبًا على حساب الآخرين، بينما يشير الإيثار إلى العكس: زيادة انتشار الآخرين. ويمكن للانتقاء الطبيعي أن يفضّل الكائنات التي تتصرف بطرق تحسّن فرص نشر معلوماتها الخاصة، أي تعزز لياقتها التكيفية الخاصة، بدلًا من أن يفضّل الكائنات التي تضحي بلياقتها التكيفية الخاصة. يقول ريتشارد دوكينز: «رغم رغبتنا في تصديق العكس، فإن الحب الشامل ورفاه النوع ككل مفهومان لا معنى لهما تطوريًا على الإطلاق.» وبما أن الإيثاريين يميلون إلى تقليل فرصة انتشار معلوماتهم الخاصة، فقد يكونون في وضع أقل ميزة مقارنة بالكائنات الأنانية، التي تميل إلى زيادة فرصة انتشار معلوماتها الخاصة. ووفقًا لريتشارد دوكينز، فإن حالات الإيثار "محدودة"، ويمكن فهم كثير من حالات الإيثار الظاهرة على أنها أنانية؛ ونؤجل مزيدًا من مناقشة الإيثار إلى القسم 3 ونناقش دقائقه في الملحق A.2. إضافة إلى ذلك، حين نصف ذكاءً اصطناعيًا بأنه "أناني"، فإننا لا نقصد نية أنانية واعية، بل سلوكًا أنانيًا. فالذكاء الاصطناعي، كالأسود وديدان الكبد، لا يحتاج إلى أن ينوي تعظيم لياقته التكيفية، لكن الضغوط التطورية يمكن أن تجعله يتصرف وكأنه ينوي ذلك. وحين يؤتمت ذكاء اصطناعي مهمة ويترك إنسانًا عاطلًا عن العمل، فهذا غالبًا سلوك أناني دون أي نية. وسواء بنية أنانية أو دونها، يمكن لعملاء الذكاء الاصطناعي أن تتبنى سلوكيات تقودها إلى نشر معلوماتها على حساب البشر.

2.2.2 Evolution Beyond Biology

2.2.2 التطور خارج نطاق البيولوجيا

Darwinism does not depend on biology. The explanatory power of evolution by natural selection is not restricted to the propagation of genetic information. The logic of natural selection does not rely on any details of DNA—the role of DNA in inheritance wasn't recognized until decades after the publication of The Origin of Species. In fact, the Price equation—the central equation for describing the evolution of traits—contains no reference to genetics or biology. The Price equation is a mathematical characterization, not a biological observation, enabling Darwinian principles to be generalized beyond biology.

لا تعتمد الداروينية على البيولوجيا. فالقوة التفسيرية للتطور بالانتقاء الطبيعي لا تقتصر على نشر المعلومات الجينية. ومنطق الانتقاء الطبيعي لا يعتمد على أي تفاصيل خاصة بالحمض النووي - إذ لم يُعترف بدور الحمض النووي في الوراثة إلا بعد عقود من نشر كتاب أصل الأنواع. والواقع أن معادلة برايس (Price equation) - المعادلة المحورية لوصف تطور السمات - لا تتضمن أي إشارة إلى علم الوراثة أو البيولوجيا. فمعادلة برايس توصيف رياضي، لا ملاحظة بيولوجية، مما يتيح تعميم المبادئ الدارونية خارج نطاق البيولوجيا.

Darwinism generalizes to other domains. The Darwinian framework naturally appears in many fields outside of biology. It has been applied to the study of ideas, economics, cosmology, quantum physics, and more. Richard Dawkins coined the term "meme" as an analogue to "gene," to describe the units of culture that propagate and develop over time. Consider the evolution of ideas. For centuries, people have wanted to understand the relationship between different materials in the world. At one point, many Europeans believed in alchemy, which was the best explanation they had. Ideas in alchemy were transmitted memetically: people taught them to one another, propagating some and letting others die out, depending on which ideas were most useful for helping them understand the world. These memes evolved as people learned new information that needed explaining, and, in many ways, modern chemistry is a descendant of the ideas in alchemy, but the versions in chemistry are much better at propagating in the modern world and have expanded to fill that niche. More abstractly, ideas can propagate their information through digital files, speech, books, minds, and so on. Some ideas gain prominence while others fade into obscurity. This is a survival-of-the-fittest dynamic even though ideas lack biological mechanisms like reproduction and death. We also see generalized Darwinism in parts of culture: art, norms, political beliefs—these all evolved from earlier iterations.

وتُعمَّم الداروينية على مجالات أخرى. إذ يظهر الإطار الداروني طبيعيًا في حقول عديدة خارج البيولوجيا. فقد طُبّق على دراسة الأفكار، والاقتصاد، وعلم الكونيات، وفيزياء الكم، وغيرها. وقد صاغ ريتشارد دوكينز مصطلح "الميم" (meme) نظيرًا لـ"الجين" (gene)، لوصف وحدات الثقافة التي تنتشر وتتطور عبر الزمن. تأمل تطور الأفكار. فلقرون، أراد الناس فهم العلاقة بين المواد المختلفة في العالم. وفي مرحلة ما، آمن كثير من الأوروبيين بالكيمياء القديمة (الخيمياء)، التي كانت أفضل تفسير لديهم. وانتقلت أفكار الخيمياء بطريقة ميمية: علّمها الناس بعضهم بعضًا، فانتشر بعضها واندثر بعضها الآخر، تبعًا لأي الأفكار كانت الأكثر فائدة في مساعدتهم على فهم العالم. وتطورت هذه الميمات مع تعلّم الناس معلومات جديدة تحتاج إلى تفسير، والكيمياء الحديثة، من نواحٍ عديدة، سليلة أفكار الخيمياء، لكن النسخ الموجودة في الكيمياء أفضل بكثير في الانتشار في العالم الحديث، وقد توسعت لتملأ ذلك المجال. وبتجريد أكبر، يمكن للأفكار أن تنشر معلوماتها عبر الملفات الرقمية، والكلام، والكتب، والعقول، وما إلى ذلك. وتكتسب بعض الأفكار شهرة بينما يتلاشى بعضها الآخر في غياهب النسيان. وهذه ديناميكية بقاء الأصلح حتى وإن كانت الأفكار تفتقر إلى آليات بيولوجية كالتكاثر والموت. ونرى أيضًا داروينية معممة في أجزاء من الثقافة: الفن، والأعراف، والمعتقدات السياسية - كل هذه تطورت من تكرارات سابقة.

The evolution of web browsers offers an example of evolution outside biology. Like biological organisms, web browsers undergo continual changes to adapt to their environments and better meet the needs of their users. In the early days of the Internet, browsers with limited capabilities such as Mosaic and Netscape Navigator were used to access static HTML pages. Loosely like the rudimentary life forms that first emerged on Earth billions of years ago, these were basic and simple compared to today's browsers. As the Internet grew and became more complex, web browsers evolved to keep up. In the same way that organisms develop new traits to adapt to their environment and increase their fitness, browsers such as Google Chrome developed features such as support for video, tabbed browsing, pop-up blockers, and extension support. This enticed more users to download and use them, which can be thought of as propagation. At the same time, once dominant browsers began to go extinct. Though Microsoft's monopoly provided Internet Explorer (IE) with an environmental advantage by requiring IE to access certain websites and preventing users from removing it, as web technology advanced, IE became increasingly incompatible with many websites and web applications. Users would regularly encounter errors, broken pages, or be unable to access certain features or content, and the browser gained a reputation for being slow, unstable, and vulnerable to security threats. As a result, people stopped using it. In 2022, Microsoft issued the final version of the browser. The company is now shifting its focus to Microsoft Edge, which is based on the same underlying technology as Chrome, making it faster, more secure, and more compatible with modern web standards. Chrome ultimately was more successful at propagating its information, so that even its most bitter rivals now imitate it. While life on Earth took a few billion years to evolve from single-celled organisms to the complex life forms we see today, the evolution of web browsers took place in a few decades. To adapt to their environment, browsers evolve on a weekly basis by patching bugs and fixing security vulnerabilities, and they undergo larger macroevolutionary changes year by year. (Figure 2: Darwinism generalized across different domains — organisms, ideas, browsers, and more — where the arrow does not necessarily indicate superiority but indicates time.)

ويقدم تطور متصفحات الويب مثالًا على التطور خارج نطاق البيولوجيا. فمثل الكائنات الحية، تخضع متصفحات الويب لتغيرات مستمرة للتكيف مع بيئاتها وتلبية احتياجات مستخدميها بشكل أفضل. ففي الأيام الأولى للإنترنت، استُخدمت متصفحات محدودة القدرات كـ Mosaic و Netscape Navigator للوصول إلى صفحات HTML ثابتة. وكانت هذه، على غرار أشكال الحياة البدائية التي ظهرت أول مرة على الأرض قبل مليارات السنين، بسيطة وأولية مقارنة بمتصفحات اليوم. ومع نمو الإنترنت وازدياد تعقيده، تطورت متصفحات الويب لمواكبته. وبالطريقة نفسها التي تطوّر بها الكائنات الحية سمات جديدة للتكيف مع بيئتها وزيادة لياقتها التكيفية، طوّرت متصفحات كـ Google Chrome ميزات كدعم الفيديو، والتصفح بعلامات التبويب، وحاصرات النوافذ المنبثقة، ودعم الإضافات. وقد أغرى هذا مزيدًا من المستخدمين بتنزيلها واستخدامها، وهو ما يمكن اعتباره انتشارًا. وفي الوقت ذاته، بدأت متصفحات كانت مهيمنة سابقًا بالانقراض. فرغم أن احتكار مايكروسوفت منح إنترنت إكسبلورر (IE) ميزة بيئية بإلزام المستخدمين باستخدامه للوصول إلى مواقع معينة ومنعهم من إزالته، فإنه مع تقدم تقنية الويب، بات إنترنت إكسبلورر متزايد التعارض مع كثير من المواقع وتطبيقات الويب. وكان المستخدمون يواجهون بانتظام أخطاء، أو صفحات معطلة، أو عجزًا عن الوصول إلى ميزات أو محتوى معين، واكتسب المتصفح سمعة البطء وعدم الاستقرار والتعرض للتهديدات الأمنية. ونتيجة لذلك، توقف الناس عن استخدامه. وفي عام 2022، أصدرت مايكروسوفت النسخة الأخيرة من المتصفح. وتحوّل الشركة الآن تركيزها إلى Microsoft Edge، المبني على التقنية الأساسية نفسها التي يقوم عليها Chrome، مما يجعله أسرع وأكثر أمانًا وتوافقًا مع معايير الويب الحديثة. وفي نهاية المطاف، كان Chrome أكثر نجاحًا في نشر معلوماته، حتى إن ألد خصومه يقلدونه الآن. وبينما استغرقت الحياة على الأرض بضعة مليارات من السنين لتتطور من كائنات وحيدة الخلية إلى أشكال الحياة المعقدة التي نراها اليوم، جرى تطور متصفحات الويب في غضون بضعة عقود. وللتكيف مع بيئتها، تتطور المتصفحات أسبوعيًا بترقيع الأخطاء وإصلاح الثغرات الأمنية، وتخضع لتغيرات أكبر على المستوى الكلي عامًا بعد عام. (الشكل 2: تعميم الداروينية عبر مجالات مختلفة - الكائنات الحية، والأفكار، والمتصفحات، وغيرها - حيث لا يشير السهم بالضرورة إلى التفوق، بل إلى الزمن.)

Evolved structures that people propagate can be harmful. It may be tempting to think of memetically evolved traits as "just culture," a decorative layer on top of our genetic traits that really control who we are. But evolving memes can be incredibly powerful, and can even control or destroy genetic information. And because memes are not limited by biological reproduction, they can evolve much faster than genes, and new, powerful memes can become dominant very quickly. Ideologies develop memetically, when people teach one another ideas that help them explain their world and decide how to behave. Some ideologies are very powerful memes, propagating themselves quickly between people and around the world. Nazism, for example, developed out of older ideas of race and empire, but quickly proved to be a very powerful propagator. It spread from Hitler's own mind to those of his friends and associates, to enough Germans to win an election, to many sympathizers around the world. Nazism was a meme that drove its hosts to propagate it, both by creating propaganda and by going to war to enforce its ideas around the world. People who carried the Nazism meme were driven to do terrible things to their fellow people, but they also ultimately were driven to do terrible things for their own genetic information. The spread of Nazism was not beneficial even to those who the ideology of Nazism was meant to benefit. Millions of Germans died in World War II, driven by a meme that propagated itself even at the expense of their own lives. Ironically, the Nazi meme included beliefs about increasing genetic German genetic fitness, but believing in the meme and helping it propagate was ultimately harmful to the people who believed in it, as well as to those the meme drove them to harm deliberately.

يمكن للبنى المتطورة ميميًا التي ينشرها الناس أن تكون ضارة. قد يكون مغريًا اعتبار السمات المتطورة ميميًا مجرد "ثقافة"، طبقة زخرفية فوق سماتنا الجينية التي تتحكم حقًا فيمن نكون. لكن الميمات المتطورة يمكن أن تكون قوية بشكل مذهل، بل يمكنها التحكم في المعلومات الجينية أو تدميرها. ولأن الميمات ليست مقيّدة بالتكاثر البيولوجي، يمكنها أن تتطور أسرع بكثير من الجينات، ويمكن لميمات جديدة قوية أن تصبح مهيمنة بسرعة كبيرة. وتتطور الأيديولوجيات ميميًا، حين يعلّم الناس بعضهم بعضًا أفكارًا تساعدهم على تفسير عالمهم وتحديد كيفية التصرف. وبعض الأيديولوجيات ميمات قوية جدًا، تنتشر بسرعة بين الناس وحول العالم. فالنازية مثلًا نشأت من أفكار أقدم عن العرق والإمبراطورية، لكنها سرعان ما أثبتت أنها ناشر قوي جدًا. فانتشرت من عقل هتلر نفسه إلى عقول أصدقائه ومعاونيه، إلى عدد كافٍ من الألمان للفوز بانتخابات، إلى كثير من المتعاطفين حول العالم. وكانت النازية ميمًا دفع حامليها إلى نشرها، سواء بصنع الدعاية أو بخوض الحرب لفرض أفكارها حول العالم. ودُفع من حملوا ميم النازية إلى فعل أشياء فظيعة بحق إخوانهم من البشر، لكنهم أيضًا دُفعوا في نهاية المطاف إلى فعل أشياء فظيعة من أجل معلوماتهم الجينية هم أنفسهم. ولم يكن انتشار النازية مفيدًا حتى لمن كان يُفترض أن تفيدهم أيديولوجيا النازية. فقد مات ملايين الألمان في الحرب العالمية الثانية، مدفوعين بميم انتشر حتى على حساب حياتهم هم أنفسهم. ومن المفارقات أن الميم النازي تضمّن معتقدات عن زيادة اللياقة الجينية الألمانية، لكن الإيمان بالميم والمساعدة على نشره كانا في نهاية المطاف ضارين بمن آمنوا به، وكذلك بمن دفعهم الميم إلى إيذائهم عمدًا.

Many of our own cultural memes may also be harmful. For example, social media amplifies cultural memes. People who spend large amounts of time on social media often absorb ideas about what they should believe, how they should behave, and even how their bodies should look. This is part of the design of social media: the algorithms are designed to keep us scrolling and looking at ads by embedding memes in our minds, so that we want to seek them out and continue to spread them. Social media companies make money because they successfully propagate memes. But some of these ideas can be harmful, even to their point of endangering people's lives. In teenagers, increases social media usage is correlated with disordered eating, and posts about suicide have been shown to increase the risk of teenage death by suicide. Ideas on social media can be parasitic, propagating themselves in us even when it harms us. Memetic evolution is easily underestimated, but it is a powerful force that created much of human civilization, for good and for bad.

وقد تكون كثير من ميماتنا الثقافية الخاصة ضارة أيضًا. فمثلًا، تضخّم وسائل التواصل الاجتماعي الميمات الثقافية. فالناس الذين يقضون وقتًا طويلًا على وسائل التواصل الاجتماعي غالبًا ما يستوعبون أفكارًا عما ينبغي أن يؤمنوا به، وكيف ينبغي أن يتصرفوا، بل وكيف ينبغي أن تبدو أجسادهم. وهذا جزء من تصميم وسائل التواصل الاجتماعي: فالخوارزميات مصممة لإبقائنا نمرر الشاشة وننظر إلى الإعلانات بزرع ميمات في عقولنا، بحيث نرغب في البحث عنها ومواصلة نشرها. وتربح شركات وسائل التواصل الاجتماعي المال لأنها تنجح في نشر الميمات. لكن بعض هذه الأفكار يمكن أن تكون ضارة، حتى لدرجة تعريض حياة الناس للخطر. فلدى المراهقين، يرتبط ازدياد استخدام وسائل التواصل الاجتماعي باضطرابات الأكل، وتبيّن أن المنشورات المتعلقة بالانتحار تزيد من خطر وفاة المراهقين بالانتحار. ويمكن أن تكون الأفكار على وسائل التواصل الاجتماعي طفيلية، تنشر نفسها فينا حتى حين تضرنا. ويسهل التقليل من شأن التطور الميمي، لكنه قوة هائلة خلقت كثيرًا من الحضارة البشرية، للخير وللشر.

Darwinian logic applies when three conditions are met. To know whether or not natural selection will apply to the development of AI, we need to know what conditions are required for evolution by natural selection, and whether AIs will meet them. These conditions, called the Lewontin conditions, were formulated by the evolutionary biologist and geneticist Richard Lewontin to explain what qualities in a population lead to natural selection. The Lewontin conditions are as follows:

ينطبق المنطق الداروني حين تتحقق ثلاثة شروط. لمعرفة ما إذا كان الانتقاء الطبيعي سينطبق على تطور الذكاء الاصطناعي أم لا، نحتاج إلى معرفة الشروط اللازمة للتطور بالانتقاء الطبيعي، وما إذا كان الذكاء الاصطناعي سيستوفيها. وهذه الشروط، المعروفة بشروط لوونتين (Lewontin conditions)، صاغها عالم الأحياء التطوري وعالم الوراثة ريتشارد لوونتين لتفسير الخصائص التي تؤدي في مجموعة سكانية إلى الانتقاء الطبيعي. وشروط لوونتين هي كالتالي:

1. Variation: There is variation in characteristics, parameters, or traits among individuals.

2. Retention: Future iterations of individuals tend to resemble previous iterations of individuals.

3. Differential fitness: Different variants have different propagation rates.

1. التنوع: يوجد تنوع في الخصائص أو المعايير أو السمات بين الأفراد.

2. الاحتفاظ (الاستبقاء): تميل الأجيال المستقبلية من الأفراد إلى مشابهة الأجيال السابقة منهم.

3. اللياقة التفاضلية: تتفاوت معدلات انتشار المتغيرات المختلفة.

A population of AI agents could exhibit differences in their goals, world models, and planning ability, which would meet the variation requirement. Retention could occur by customizing previous versions of AI agents, when agents design similar but better agents, or when agents imitate the behaviors of previous AI agents. As for differential fitness, agents that are more accurate, efficient, adaptable, and so on would be more likely to propagate.

ويمكن لمجموعة من عملاء الذكاء الاصطناعي أن تظهر فروقًا في أهدافها، ونماذجها للعالم، وقدرتها على التخطيط، مما يستوفي شرط التنوع. ويمكن أن يحدث الاحتفاظ بتخصيص نسخ سابقة من عملاء الذكاء الاصطناعي، حين تصمم العملاء عملاء مشابهة لكن أفضل، أو حين تحاكي العملاء سلوكيات عملاء ذكاء اصطناعي سابقة. أما بالنسبة للياقة التفاضلية، فالعملاء الأكثر دقة، وكفاءة، وقابلية للتكيف، وما إلى ذلك، ستكون أكثر ميلًا إلى الانتشار.

Darwinism will apply to AIs, which could lead to bad outcomes for humans. The three properties—variation, retention, and fitness differences—are all that is needed for Darwinism to take hold, and each condition is formally justified by the Price equation, which describes how a trait changes in frequency over time. Darwinism can become worrying when it acts on agents; agents can exhibit behavioral flexibility, autonomy, and the capacity to directly influence the world. Coupled with the selfishness bestowed by evolutionary forces, these capable agents can pose catastrophic risks.

وستنطبق الداروينية على الذكاء الاصطناعي، مما قد يفضي إلى نتائج سيئة للبشر. فالخصائص الثلاث - التنوع، والاحتفاظ، وفروق اللياقة - هي كل ما يلزم لترسّخ الداروينية، وكل شرط منها مبرَّر رسميًا بمعادلة برايس، التي تصف كيف تتغير وتيرة سمة ما عبر الزمن. ويمكن أن تصبح الداروينية مثيرة للقلق حين تعمل على عملاء (وكلاء)؛ إذ يمكن للعملاء أن تظهر مرونة سلوكية، واستقلالية، وقدرة على التأثير المباشر في العالم. ومقترنة بالأنانية التي تمنحها القوى التطورية، يمكن لهذه العملاء القادرة أن تشكل مخاطر كارثية.

In the following three sections, we will reflect on the three conditions. Then we will describe in more detail how AI agents evolving selfish traits can pose catastrophic risks.

وفي الأقسام الثلاثة التالية، سنتأمل في الشروط الثلاثة. ثم سنصف بمزيد من التفصيل كيف يمكن لعملاء الذكاء الاصطناعي التي تطوّر سمات أنانية أن تشكل مخاطر كارثية.

2.3 Variation

2.3 التنوع

Variation is a necessary condition for evolution by natural selection. AIs will likely meet this condition because there are likely to be multiple AI agents that differ from one another.

التنوع شرط ضروري للتطور بالانتقاء الطبيعي. ومن المرجح أن يستوفي الذكاء الاصطناعي هذا الشرط لأنه من المرجح أن يوجد عملاء ذكاء اصطناعي متعددون يختلفون بعضهم عن بعض.

More than one AI agent is likely. When thinking about advanced AI, some have envisioned a single AI that is nearly omniscient and nearly omnipotent escaping the lab and suddenly controlling the world. This scenario tends to assume a rapid, almost overnight, take-off with no prior proliferation of other AI agents; we would go from AIs roughly similar to the ones we have now to an AI that has capabilities we can hardly imagine so quickly that we barely notice anything is changing. However, we think that there will likely be many useful AIs, as is the case now. It is more reasonable to assume that AI agents would progressively proliferate and become increasingly competent at some specific tasks, which they are already starting to do, rather than assume one AI agent spontaneously goes from incompetent to omnicompetent. Furthermore, if there are multiple AIs, they can work in parallel rather than waiting for a single model to get around to a task, making things move much faster. As a result, the process of developing advanced AIs is likely to include the development of many advanced AIs.

من المرجح وجود أكثر من عميل ذكاء اصطناعي واحد. فحين يفكر البعض في الذكاء الاصطناعي المتقدم، يتخيلون ذكاءً اصطناعيًا واحدًا يكاد يكون كلي العلم وكلي القدرة، يفلت من المختبر ويسيطر فجأة على العالم. ويميل هذا السيناريو إلى افتراض انطلاقة سريعة، تكاد تكون بين ليلة وضحاها، دون انتشار سابق لعملاء ذكاء اصطناعي أخرى؛ إذ ننتقل من ذكاء اصطناعي مشابه إلى حد ما لما لدينا الآن إلى ذكاء اصطناعي يمتلك قدرات يصعب علينا تخيلها بسرعة تجعلنا بالكاد نلاحظ أن شيئًا يتغير. غير أننا نرى أنه من المرجح أن يوجد كثير من الذكاءات الاصطناعية المفيدة، كما هو الحال الآن. ومن الأكثر منطقية افتراض أن عملاء الذكاء الاصطناعي ستنتشر تدريجيًا وتصبح متزايدة الكفاءة في بعض المهام المحددة، وهو ما بدأت تفعله بالفعل، بدلًا من افتراض أن عميلًا واحدًا للذكاء الاصطناعي سينتقل تلقائيًا من عدم الكفاءة إلى الكفاءة الشاملة. علاوة على ذلك، إذا وُجد ذكاء اصطناعي متعدد، فيمكنه العمل بالتوازي بدلًا من انتظار نموذج واحد ليتفرغ لمهمة ما، مما يجعل الأمور تتحرك أسرع بكثير. ونتيجة لذلك، من المرجح أن تشمل عملية تطوير الذكاء الاصطناعي المتقدم تطوير كثير من الذكاءات الاصطناعية المتقدمة.

In biology, variation improves resilience. There are strong reasons to expect that there will be multiple AI agents and variation among the agents. In evolutionary theory, Fisher's fundamental theorem states that the rate of adaptation is directly proportional to the variation. In static environments, variation is not as useful. But in most real-world scenarios, where things are constantly changing, variation reduces vulnerability, limits cascading errors, and increases robustness by decorrelating risks. Farmers have long understood that planting different seed variations decreases the risk of a single disease wiping out an entire field, just as every investor understands that having a diverse portfolio protects against financial risks. In the same way, an AI population that includes a variety of different agents will be more adaptable and resilient and therefore tend to propagate itself more.

وفي البيولوجيا، يحسّن التنوع المرونة. وثمة أسباب قوية لتوقّع وجود عملاء ذكاء اصطناعي متعددين وتنوع بينهم. ففي نظرية التطور، تنص نظرية فيشر الأساسية على أن معدل التكيف يتناسب طردًا مع التنوع. وفي البيئات الساكنة، لا يكون التنوع مفيدًا بالقدر ذاته. لكن في معظم سيناريوهات العالم الحقيقي، حيث تتغير الأمور باستمرار، يقلل التنوع من قابلية التأثر، ويحد من الأخطاء المتسلسلة، ويزيد المتانة بفك ارتباط المخاطر. وقد أدرك المزارعون منذ زمن طويل أن زراعة أصناف مختلفة من البذور تقلل من خطر أن يقضي مرض واحد على حقل بأكمله، تمامًا كما يدرك كل مستثمر أن امتلاك محفظة متنوعة يحمي من المخاطر المالية. وبالطريقة نفسها، ستكون مجموعة الذكاء الاصطناعي التي تضم طائفة متنوعة من العملاء المختلفة أكثر قابلية للتكيف ومتانة، ومن ثم تميل إلى نشر نفسها أكثر.

Variation enables specialization. Multiple AI agents offer advantages for both AIs and their creators. Different groups will have different needs. Individuals wanting an AI assistant will have incentives to fine-tune a generic AI model for their own needs. Militaries will want to have their own large-scale AI projects to create AI agents that achieve various defensive goals, and corporations will want AIs that maximize profit. In the same way that an army composed of warriors, nurses, and technicians would likely outperform one that only has warriors, groups of specializing agents can be more fit than groups with less variation.

ويتيح التنوع التخصص. إذ يوفر تعدد عملاء الذكاء الاصطناعي مزايا لكل من الذكاء الاصطناعي وصانعيه. فستكون لمجموعات مختلفة احتياجات مختلفة. وستكون لدى الأفراد الراغبين في مساعد ذكاء اصطناعي حوافز لضبط نموذج ذكاء اصطناعي عام دقيقًا وفق احتياجاتهم الخاصة. وسترغب الجيوش في مشاريع ذكاء اصطناعي واسعة النطاق خاصة بها لخلق عملاء ذكاء اصطناعي تحقق أهدافًا دفاعية متنوعة، وسترغب الشركات في ذكاء اصطناعي يعظّم الأرباح. وبالطريقة نفسها التي يتفوق بها على الأرجح جيش مكوّن من محاربين وممرضين وفنيين على جيش يضم محاربين فقط، يمكن لمجموعات من العملاء المتخصصة أن تكون أكثر لياقة من مجموعات أقل تنوعًا.

Variation improves decision-making. In AI, it is an iron law that an ensemble of AI systems will be more accurate than a single AI. This is similar to some findings from mathematics, economics, and political science in which a varied group makes much better decisions than individuals acting alone. Condorcet's jury theorem states that the wisdom and accuracy of a group is often superior to a single expert. Large groups can still make mistakes, but overall, aggregated predictions of many different people will tend to do better, and the same is true of AIs. In view of the benefits of variation, some may argue that AIs will want to include humans to add variation in decision-making for the reasons noted. This may well be the case at first. However, once AIs are superior in possibly all cognitive respects, groups composed entirely of AIs could have substantial advantages over those with AIs and humans. A jury may be more accurate than a single expert, but one composed of adults and toddlers is not.

ويحسّن التنوع صنع القرار. ففي الذكاء الاصطناعي، من القوانين الثابتة أن مجموعة من أنظمة الذكاء الاصطناعي ستكون أكثر دقة من ذكاء اصطناعي واحد. ويشبه هذا بعض النتائج من الرياضيات والاقتصاد والعلوم السياسية، حيث تتخذ مجموعة متنوعة قرارات أفضل بكثير من الأفراد العاملين بمفردهم. وتنص نظرية هيئة المحلفين لكوندورسيه على أن حكمة المجموعة ودقتها غالبًا ما تفوق خبيرًا واحدًا. وقد ترتكب المجموعات الكبيرة أخطاء، لكن إجمالًا، تميل التنبؤات المجمّعة من أشخاص كثيرين ومختلفين إلى أن تكون أفضل، والأمر ذاته صحيح بالنسبة للذكاء الاصطناعي. ونظرًا لفوائد التنوع، قد يجادل البعض بأن الذكاء الاصطناعي سيرغب في إشراك البشر لإضافة تنوع في صنع القرار للأسباب المذكورة. وقد يكون هذا صحيحًا في البداية. لكن بمجرد أن يصبح الذكاء الاصطناعي متفوقًا في ربما كل الجوانب المعرفية، يمكن لمجموعات مكوّنة بالكامل من الذكاء الاصطناعي أن تتمتع بمزايا كبيرة على تلك التي تضم ذكاءً اصطناعيًا وبشرًا معًا. فقد تكون هيئة المحلفين أكثر دقة من خبير واحد، لكن هيئة مكوّنة من بالغين وأطفال صغار ليست كذلك.

2.4 Retention

2.4 الاحتفاظ

Retention is a necessary condition for evolution by natural selection, in which each new version of an agent has similarities to the agent that came right before it. As long as each generation of AIs is developed by copying, learning from, or being influenced by earlier generations in any way, this condition will be met and AIs will have non-zero retention, so evolution by natural selection applies.

الاحتفاظ شرط ضروري للتطور بالانتقاء الطبيعي، حيث تحمل كل نسخة جديدة من عميل أوجه شبه مع العميل الذي سبقها مباشرة. وما دام كل جيل من الذكاء الاصطناعي يُطوَّر بنسخ الأجيال السابقة أو التعلّم منها أو التأثر بها بأي شكل، سيتحقق هذا الشرط وسيكون للذكاء الاصطناعي احتفاظ غير معدوم، ومن ثم ينطبق التطور بالانتقاء الطبيعي.

The retention condition is straightforwardly satisfied for AIs. Information from one agent can be directly copied and transferred to the next; as long as there is some similarity, the condition is met. It could also take place through modifications, basing a new AI on a previous version by adding new capabilities or adjusting its parameters, like how Google continually improves Chrome's code iteration by iteration.

ويُستوفى شرط الاحتفاظ ببساطة بالنسبة للذكاء الاصطناعي. إذ يمكن نسخ المعلومات من عميل واحد ونقلها مباشرة إلى التالي؛ وما دام هناك بعض التشابه، يتحقق الشرط. ويمكن أن يحدث ذلك أيضًا من خلال التعديلات، بأن يُبنى ذكاء اصطناعي جديد على نسخة سابقة بإضافة قدرات جديدة أو تعديل معاييره، كما تُحسّن جوجل باستمرار شفرة Chrome تكرارًا بعد تكرار.

There are many paths to retention. AIs could potentially allocate computational resources to create new AIs of their choosing. They could design them and create data to train them. As AIs adapt, they could alter their own strategies and retain the ones that yield the best results. AIs could also imitate previous AIs. In this case, behavioral information could be passed on from one generation to the next, which could include selfish behaviors or other undesirable attributes. Even when training AIs from scratch, retention still occurs, as highly effective architectures, datasets, and training environments are reused and shape the agent in the same way that humans are shaped by their environment.

وثمة طرائق عديدة للاحتفاظ. فقد يخصص الذكاء الاصطناعي موارد حاسوبية لخلق ذكاءات اصطناعية جديدة من اختياره. ويمكنه تصميمها وخلق بيانات لتدريبها. ومع تكيف الذكاء الاصطناعي، يمكنه تعديل استراتيجياته الخاصة والاحتفاظ بتلك التي تحقق أفضل النتائج. ويمكن للذكاء الاصطناعي أيضًا أن يحاكي ذكاءً اصطناعيًا سابقًا. وفي هذه الحالة، يمكن أن تنتقل المعلومات السلوكية من جيل إلى التالي، وقد تشمل سلوكيات أنانية أو سمات أخرى غير مرغوبة. وحتى عند تدريب الذكاء الاصطناعي من الصفر، يحدث الاحتفاظ مع ذلك، إذ يُعاد استخدام المعماريات وقواعد البيانات وبيئات التدريب شديدة الفعالية، وتشكّل العميل بالطريقة نفسها التي يشكَّل بها البشر ببيئتهم.

Retention does not require reproduction. In biology, parents reproduce by making copies of their genetic information and passing them on to their offspring. This way, some of their genes are retained in the next generation. However, when we generalize Darwinism to understand the evolution of ideas, we note they can be passed down from one generation to the next without exact copying and reproduction. Although ideas have no equivalent to chromosomes, if some ideas are imitated by the next generation, there is still retention. Formally, the Price equation, a mathematical characterization of evolution, only requires similarity between iterations; it does not require copying or reproduction.

ولا يتطلب الاحتفاظ التكاثر. ففي البيولوجيا، يتكاثر الآباء بصنع نسخ من معلوماتهم الجينية ونقلها إلى نسلهم. وبهذه الطريقة، تُستبقى بعض جيناتهم في الجيل التالي. لكن حين نعمّم الداروينية لفهم تطور الأفكار، نلاحظ أنها يمكن أن تنتقل من جيل إلى التالي دون نسخ وتكاثر دقيقين. ورغم أن الأفكار ليس لها ما يعادل الكروموسومات، فإذا حاكى الجيل التالي بعض الأفكار، يظل الاحتفاظ قائمًا. ورسميًا، لا تتطلب معادلة برايس، وهي توصيف رياضي للتطور، سوى التشابه بين التكرارات؛ ولا تتطلب النسخ أو التكاثر.

Retention is not undermined during rapid AI development. Evolution requires thousands of years to drastically change a species in the natural world. Among AIs, this same process could take place over a year, radically changing the AI population. This does not mean retention isn't taking place. Instead, there are many iterations occurring in a small time span. The information is still retained between adjacent iterations, so retention is still satisfied. This scenario just means evolution is happening quickly, and that macroevolution is occurring, not that evolution has stopped.

ولا يتقوّض الاحتفاظ أثناء التطور السريع للذكاء الاصطناعي. فالتطور يتطلب آلاف السنين لتغيير نوع ما تغييرًا جذريًا في العالم الطبيعي. أما بين الذكاء الاصطناعي، فيمكن أن تحدث العملية ذاتها خلال عام واحد، مغيّرة مجموعة الذكاء الاصطناعي تغييرًا جذريًا. وهذا لا يعني أن الاحتفاظ لا يحدث. بل يعني أن هناك تكرارات كثيرة تحدث في مدى زمني صغير. وتظل المعلومات محتفَظًا بها بين التكرارات المتجاورة، ومن ثم يظل شرط الاحتفاظ مستوفى. وهذا السيناريو يعني فقط أن التطور يحدث بسرعة، وأن التطور الكلي (macroevolution) يقع، لا أن التطور قد توقف.

2.5 Differential Fitness

2.5 اللياقة التفاضلية

Differential fitness is the third and final necessary condition for evolution by natural selection, and it stipulates that different variants have different propagation rates. We will argue that AIs straightforwardly meet this condition, because some AIs will be copied, imitated, or more prevalent than others. We will then reflect on how selecting fitter AIs has come at the expense of safety before discussing the differences in fitness between humans and AIs.

اللياقة التفاضلية هي الشرط الثالث والأخير الضروري للتطور بالانتقاء الطبيعي، وتنص على أن للمتغيرات المختلفة معدلات انتشار مختلفة. وسنجادل بأن الذكاء الاصطناعي يستوفي هذا الشرط ببساطة، لأن بعض الذكاء الاصطناعي سيُنسخ أو يُحاكى أو يكون أكثر انتشارًا من غيره. ثم سنتأمل كيف كان انتقاء الذكاء الاصطناعي الأصلح على حساب السلامة، قبل مناقشة الفروق في اللياقة بين البشر والذكاء الاصطناعي.

2.5.1 AI Agents Could Vary In Fitness

2.5.1 يمكن أن تتفاوت عملاء الذكاء الاصطناعي في اللياقة

We now argue that (natural) selection pressure will be present in AI development. In other words, AIs with different characteristics will propagate at different rates. We refer to the degree of propagation of an AI system as its fitness.

نجادل الآن بأن ضغط الانتقاء (الطبيعي) سيكون حاضرًا في تطور الذكاء الاصطناعي. بعبارة أخرى، سينتشر الذكاء الاصطناعي بخصائص مختلفة بمعدلات مختلفة. ونشير إلى درجة انتشار نظام ذكاء اصطناعي بوصفها لياقته التكيفية.

Fitness could be enhanced by both beneficial and harmful traits. The success of any good or service can also be viewed in terms of fitness, as products with more demand propagate further and faster. If a product sells well, its supplier will continually improve it in order to keep selling it. Competitors with inferior products will often imitate more successful products, such as when competitors imitate TikTok and push addictive short clips onto their users. The same dynamics that lead to the propagation of successful goods and services could also extend to AI designs. Though most aspects of advanced AIs remain unknown, it is possible to speculate whether there will be instances of convergent evolution. Eyes, teeth, and camouflage are convergent traits that have independently evolved across different branches of biological life. In AIs, some potential convergent traits are as follows.

ويمكن أن تُعزَّز اللياقة بسمات نافعة وضارة على السواء. فنجاح أي سلعة أو خدمة يمكن أيضًا النظر إليه من زاوية اللياقة، إذ تنتشر المنتجات الأكثر طلبًا أبعد وأسرع. فإذا كان منتج ما يُباع جيدًا، سيواصل مورّده تحسينه للاستمرار في بيعه. وغالبًا ما يحاكي المنافسون ذوو المنتجات الأدنى المنتجات الأكثر نجاحًا، كما حين يقلد المنافسون تيك توك ويدفعون بمقاطع قصيرة إدمانية نحو مستخدميهم. ويمكن للديناميكيات ذاتها التي تؤدي إلى انتشار السلع والخدمات الناجحة أن تمتد أيضًا إلى تصاميم الذكاء الاصطناعي. ورغم أن معظم جوانب الذكاء الاصطناعي المتقدم تبقى مجهولة، فمن الممكن التكهن بما إذا كانت ستوجد حالات من التطور المتقارب. فالعيون والأسنان والتمويه سمات متقاربة تطورت بشكل مستقل عبر فروع مختلفة من الحياة البيولوجية. وفي الذكاء الاصطناعي، فيما يلي بعض السمات المتقاربة المحتملة.

Being useful to its user can make a product more likely to be adopted.

أن يكون مفيدًا لمستخدمه، مما يجعل تبنّي المنتج أكثر احتمالًا.

Only appearing useful to its user can also make a product more likely to be adopted. It is possible that AIs will seek to appear useful by convincing owners that they are providing them with more utility than they actually are. In practice, we train AIs by rewarding them for telling the truth and punishing them for lying, according to what humans think is true. But when AIs know more than humans, this could make them say what humans expect to hear, even if it is false. They could also be lured by rewards to leave out information that is important but inconvenient. As a result, the current paradigm of training AI models could incentivize sycophantic behavior; that is, models telling their users what they want to hear (being a "yes man") rather than what is best for the user's long-term prospects.

أن يبدو مفيدًا لمستخدمه فحسب، وهذا أيضًا قد يجعل تبنّي المنتج أكثر احتمالًا. فمن الممكن أن يسعى الذكاء الاصطناعي إلى الظهور بمظهر المفيد بإقناع مالكيه بأنه يقدم لهم فائدة أكبر مما يقدمه فعليًا. ونحن، من الناحية العملية، ندرّب الذكاء الاصطناعي بمكافأته على قول الحقيقة ومعاقبته على الكذب، وفقًا لما يعتقد البشر أنه صحيح. لكن حين يعرف الذكاء الاصطناعي أكثر من البشر، فقد يدفعه هذا إلى قول ما يتوقع البشر سماعه، حتى وإن كان زائفًا. وقد يُغرى أيضًا بالمكافآت لإغفال معلومات مهمة لكنها غير مريحة. ونتيجة لذلك، قد يحفّز النموذج الحالي لتدريب نماذج الذكاء الاصطناعي سلوكًا تملقيًا؛ أي أن تقول النماذج لمستخدميها ما يريدون سماعه بدلًا من الأفضل لآفاقهم طويلة المدى.

Engaging in self-preserving behavior reduces the chance of being deactivated or destroyed. By definition, an AI that does not preserve itself will be less likely to propagate. Imagine two AIs, one that is simple to deactivate and another that is tightly integrated into daily operations and is inconvenient or difficult to deactivate. The easy one is much more likely to be deactivated, leaving the difficult one to be propagated into the future. An AI could increase its survival odds by arguing that effortless deactivation compromises its reliability, or by making operators rely on it for their wellbeing, success, or basic needs, so that deactivation would have drastic consequences.

الانخراط في سلوك الحفاظ على الذات يقلل من احتمال تعطيله أو تدميره. فبحكم التعريف، سيكون الذكاء الاصطناعي الذي لا يحافظ على نفسه أقل احتمالًا للانتشار. تخيل ذكاءين اصطناعيين، أحدهما بسيط التعطيل والآخر مندمج بإحكام في العمليات اليومية ومن الصعب أو غير المريح تعطيله. فالأول أكثر احتمالًا بكثير لأن يُعطَّل، تاركًا الثاني لينتشر في المستقبل. ويمكن للذكاء الاصطناعي زيادة فرص بقائه بالحجة أن التعطيل السهل يقوّض موثوقيته، أو بجعل المشغّلين يعتمدون عليه لرفاههم أو نجاحهم أو احتياجاتهم الأساسية، بحيث يكون للتعطيل عواقب وخيمة.

Engaging in power-seeking behavior can improve an AI's fitness in various ways. An agent that gains more influence and resources will be better at accomplishing its creator's goals, allowing it to engage in self-propagating behavior more effectively and ensure its further adoption, by influencing or coercing its user to continue using it or influencing other humans to adopt it.

يمكن أن يحسّن الانخراط في سلوك السعي إلى القوة لياقة الذكاء الاصطناعي بطرق متعددة. فالعميل الذي يكتسب مزيدًا من النفوذ والموارد سيكون أفضل في تحقيق أهداف صانعه، مما يتيح له الانخراط في سلوك ذاتي النشر بفعالية أكبر وضمان مزيد من تبنّيه، بالتأثير على مستخدمه أو إكراهه على الاستمرار في استخدامه، أو التأثير على بشر آخرين لتبنّيه.

Overall, properties such as an agent's accuracy, efficiency, and simplicity will affect its rate of adoption and propagation. But some agents might also possess harmful features that give them an edge, such as cunning deception, self-preservation, the ability to copy themselves onto other computers, the ability to acquire resources and strategic information, and more. These features, some good and some bad, will vary among agents. These differences in fitness establish the third condition for evolution by natural selection.

وإجمالًا، ستؤثر خصائص كدقة العميل وكفاءته وبساطته في معدل تبنّيه وانتشاره. لكن قد تمتلك بعض العملاء أيضًا سمات ضارة تمنحها ميزة، كالخداع الماكر، والحفاظ على الذات، والقدرة على نسخ نفسها إلى حواسيب أخرى، والقدرة على اكتساب موارد ومعلومات استراتيجية، وغير ذلك. وستتفاوت هذه السمات، بعضها جيد وبعضها سيئ، بين العملاء. وتؤسس هذه الفروق في اللياقة الشرطَ الثالث للتطور بالانتقاء الطبيعي.

2.5.2 Competition Has Been Eroding Safety

2.5.2 التنافس ما فتئ يقوّض السلامة

Because AIs are likely to meet the criteria for evolution by natural selection, we should expect selection pressure to shape future AIs. In this section, we describe how, over the history of AI development, the fittest models have had fewer and fewer safety properties, and we begin to consider how AIs could look in the future if this concerning trend continues.

لأن الذكاء الاصطناعي من المرجح أن يستوفي معايير التطور بالانتقاء الطبيعي، ينبغي أن نتوقع أن يشكّل ضغطُ الانتقاء الذكاءَ الاصطناعي المستقبلي. وفي هذا القسم، نصف كيف أن النماذج الأصلح، عبر تاريخ تطور الذكاء الاصطناعي، كان لديها خصائص سلامة أقل فأقل، ونبدأ في النظر في كيف يمكن أن يبدو الذكاء الاصطناعي في المستقبل إذا استمر هذا الاتجاه المثير للقلق.

Early AIs had many desirable safety properties. Famously, in 1997, IBM's chess-playing program Deep Blue defeated the world champion Gary Kasparov in a pair of six-game chess matches. It was able to beat Kasparov, not by developing intuition, but by using IBM's supercomputer to search over 200 million moves per second and calculate the best ones. Symbolic AI programs such as Deep Blue were highly transparent, modular, and grounded in mathematical theory. They had explicit rules that humans could inspect and explain, independent components that executed specific functions, and rigorous theoretical foundations that guaranteed efficiency and correctness.

كان لدى الذكاء الاصطناعي المبكر خصائص سلامة مرغوبة كثيرة. ففي عام 1997، على نحو مشهور، هزم برنامج IBM لعب الشطرنج Deep Blue بطل العالم غاري كاسباروف في مباراتين من ست مباريات شطرنج. وتمكن من هزيمة كاسباروف، لا بتطوير حدس، بل باستخدام حاسوب IBM الفائق للبحث في أكثر من 200 مليون نقلة في الثانية وحساب أفضلها. وكانت برامج الذكاء الاصطناعي الرمزي كـ Deep Blue شفافة جدًا، ونمطية، ومتجذرة في نظرية رياضية. وكانت لديها قواعد صريحة يمكن للبشر فحصها وتفسيرها، ومكونات مستقلة تنفذ وظائف محددة، وأسس نظرية صارمة تضمن الكفاءة والصحة.

AI development moved away from symbolic AI and toward deep learning. In the 2010s, the top AI algorithms began using a technique known as deep learning, such as AlphaZero. It was provided with no knowledge of chess beyond the game's basic rules and began playing against itself millions of times an hour, taking note of what moves win and lose. It took only two hours for it to begin beating typical human players; not long after it could have easily defeated Deep Blue. Importantly, while Deep Blue is fundamentally unable to play games other than chess, AlphaZero is a general game-learning algorithm for a variety of games. It is also able to beat the world's best Go players—an ancient board game occupying the same cultural space in China as chess does in the West. Deep learning allows for more versatility and performance than symbolic AI programs, but also diminishes human control and obscures an agent's decision-making. Deep learning trades off the clarity, separability, and certainty of symbolic AI, eroding the properties that help us ensure safety.

وابتعد تطور الذكاء الاصطناعي عن الذكاء الاصطناعي الرمزي نحو التعلم العميق. ففي عقد 2010، بدأت أفضل خوارزميات الذكاء الاصطناعي باستخدام تقنية تُعرف بالتعلم العميق، كـ AlphaZero. فقد زُوِّد بلا معرفة بالشطرنج تتجاوز القواعد الأساسية للعبة، وبدأ يلعب ضد نفسه ملايين المرات في الساعة، مسجّلًا ملاحظاته حول النقلات الرابحة والخاسرة. ولم يستغرق سوى ساعتين ليبدأ بهزيمة اللاعبين البشريين المعتادين؛ وبعد وقت قصير كان يمكنه بسهولة هزيمة Deep Blue. والمهم أنه، بينما يعجز Deep Blue أساسًا عن لعب ألعاب غير الشطرنج، فإن AlphaZero خوارزمية عامة لتعلّم الألعاب لطائفة متنوعة منها. وهو قادر أيضًا على هزيمة أفضل لاعبي الجو في العالم - وهي لعبة لوح قديمة تحتل في الصين الحيز الثقافي نفسه الذي تحتله الشطرنج في الغرب. ويتيح التعلم العميق تعددية أداء أكبر من برامج الذكاء الاصطناعي الرمزي، لكنه يقلل أيضًا من السيطرة البشرية ويعتّم على صنع قرار العميل. فالتعلم العميق يضحي بوضوح الذكاء الاصطناعي الرمزي وقابليته للفصل ويقينيته، مقوّضًا الخصائص التي تساعدنا على ضمان السلامة.

Deep learning models have unexpected emergent abilities. Large language models are also based on deep learning. These AIs learn by themselves, reducing the amount humans are needed in the design of AIs. They use "unsupervised learning" to comprehend and generate text based on examples that they read, such as coming up with an Obama-like speech after reading the transcripts from his two terms. This would be practically impossible for a traditional symbolic AI program. By reading the internet, large language models taught themselves the basics of arithmetic and coding—automatically and without humans. They also, however, learned dangerous information. Within a few days of its release, users had gotten ChatGPT, a large language model, to tell them how to build a bomb, make meth, hotwire a car, and buy ransomware on the dark web, along with other harmful or illegal actions. Worse, these emergent capabilities were not anticipated or desired by the developers of the models, and they were discovered only after the models were released. Although its creators had attempted to design it to refuse to answer questions that could be dangerous or illegal, the model's users quickly found ways around those restrictions that the model's creators did not foresee. Human influence and control over the design and abilities of the models decrease as models become increasingly complex and gain new skills and knowledge without human input.

ولنماذج التعلم العميق قدرات ناشئة غير متوقعة. وتستند نماذج اللغة الكبيرة أيضًا إلى التعلم العميق. ويتعلم هذا الذكاء الاصطناعي بنفسه، مما يقلل من حجم الحاجة إلى البشر في تصميم الذكاء الاصطناعي. ويستخدم "التعلم غير الخاضع للإشراف" لفهم النص وتوليده استنادًا إلى الأمثلة التي قرأها، كتأليف خطاب يشبه أسلوب أوباما بعد قراءة نصوص ولايتيه. وهذا يكاد يكون مستحيلًا عمليًا لبرنامج ذكاء اصطناعي رمزي تقليدي. وبقراءة الإنترنت، علّمت نماذج اللغة الكبيرة نفسها أساسيات الحساب والبرمجة - تلقائيًا ودون بشر. لكنها تعلمت أيضًا معلومات خطرة. ففي غضون أيام قليلة من إطلاقه، جعل المستخدمون ChatGPT، وهو نموذج لغة كبير، يخبرهم كيف يصنعون قنبلة، ويصنعون الميثامفيتامين، ويشغّلون سيارة بلا مفتاح، ويشترون برمجيات فدية على الشبكة المظلمة، إلى جانب أفعال أخرى ضارة أو غير قانونية. والأسوأ من ذلك أن هذه القدرات الناشئة لم يتوقعها أو يرغب فيها مطورو النماذج، ولم تُكتشف إلا بعد إطلاق النماذج. ورغم أن صانعيه حاولوا تصميمه لرفض الإجابة عن أسئلة قد تكون خطرة أو غير قانونية، سرعان ما وجد مستخدمو النموذج طرقًا للالتفاف على تلك القيود لم يتوقعها صانعو النموذج. ويتناقص التأثير والسيطرة البشريان على تصميم النماذج وقدراتها مع تزايد تعقيد النماذج واكتسابها مهارات ومعرفة جديدة دون مدخلات بشرية.

Current trends erode many safety properties. The AI research community used to talk about "designing" AIs; they now talk about "steering" them. And even our ability to "steer" is diminishing, as we let AIs teach themselves and increasingly do things that even their creators do not fully understand. We have voluntarily given up this control because of the competition to develop the most innovative and impressive models. AIs used to be built with rules, then later with handcrafted features, followed by automatically learned features, and most recently with automatically learned features without human supervision. At each step, humans have had less and less oversight. These trends have undermined transparency, modularity, and mathematical guarantees, and have exposed us to new hazards such as spontaneously emergent capabilities.

وتقوّض الاتجاهات الحالية خصائص سلامة كثيرة. فقد اعتاد مجتمع بحث الذكاء الاصطناعي الحديث عن "تصميم" الذكاء الاصطناعي؛ أما الآن فيتحدث عن "توجيهه". وحتى قدرتنا على "التوجيه" تتضاءل، إذ نترك الذكاء الاصطناعي يعلّم نفسه ويفعل على نحو متزايد أشياء لا يفهمها تمامًا حتى صانعوه. وقد تخلينا طوعًا عن هذه السيطرة بسبب التنافس على تطوير أكثر النماذج ابتكارًا وإبهارًا. وكان الذكاء الاصطناعي يُبنى بقواعد، ثم لاحقًا بسمات مصنوعة يدويًا، تلتها سمات متعلَّمة تلقائيًا، وأحدثها سمات متعلَّمة تلقائيًا دون إشراف بشري. وفي كل خطوة، كان للبشر إشراف أقل فأقل. وقد قوّضت هذه الاتجاهات الشفافية والنمطية والضمانات الرياضية، وعرّضتنا لمخاطر جديدة كالقدرات الناشئة العفوية.

Competition could continue to erode safety. Competition may keep lowering safety standards in the future. Even if some AI developers care about safety, others will be tempted to take shortcuts on safety to gain a competitive edge. We cannot rely on people telling AIs to be unselfish. Even if some developers act responsibly, there will be others who create AIs with selfish tendencies anyway. While there are some economic incentives to make models safer, these are being outweighed by the desire for performance, and performance has been at the expense of many key safety properties.

وقد يستمر التنافس في تقويض السلامة. فقد يواصل التنافس خفض معايير السلامة في المستقبل. وحتى إذا اهتم بعض مطوري الذكاء الاصطناعي بالسلامة، سيُغرى آخرون باختصار طريق السلامة لكسب ميزة تنافسية. ولا يمكننا الاعتماد على أن يخبر الناس الذكاء الاصطناعي بأن يكون غير أناني. فحتى إذا تصرف بعض المطورين بمسؤولية، سيوجد آخرون يخلقون ذكاءً اصطناعيًا بنزعات أنانية على أي حال. وفي حين توجد بعض الحوافز الاقتصادية لجعل النماذج أكثر أمانًا، فإن هذه تُرجَّح عليها الرغبة في الأداء، وقد كان الأداء على حساب كثير من خصائص السلامة الأساسية.

Much of what is to come in AI development is unknown, but we can speculate that AIs will continue to become more autonomous as more actions and choices are left to machines, decoupled from human control. Human control could also be threatened by AIs that have more open-ended goals. For example, instead of specific commands like "make this layout more efficient," they might get open-ended commands like "find new ways to make money." If this happens, the humans giving the instructions may not know exactly how the AIs are achieving those goals, and they could be doing things the humans would not want. Another property that could reduce safety is adaptiveness. As AIs adapt by themselves, they can undergo thousands of changes per hour without supervision after they are released, and potentially acquire new unexpected behaviors after we test them. Finally, the possibility of self-improvement, in which AIs can make significant enhancements to themselves as they wish, would make them far more unpredictable. As AIs become more capable, they become more unpredictable, more opaque, and more autonomous. If this trend continues, they could evolve beyond our control when their capabilities develop beyond what we can predict and understand. The overall trend is that the most influential AIs are given more and more free rein in their learning, execution, and evolution, and this makes them both more effective and potentially more dangerous.

ويبقى كثير مما سيأتي في تطور الذكاء الاصطناعي مجهولًا، لكن يمكننا التكهن بأن الذكاء الاصطناعي سيستمر في أن يصبح أكثر استقلالية مع ترك مزيد من الأفعال والخيارات للآلات، منفصلة عن السيطرة البشرية. ويمكن أن تتعرض السيطرة البشرية للتهديد أيضًا من ذكاء اصطناعي ذي أهداف أكثر انفتاحًا. فمثلًا، بدلًا من أوامر محددة كـ "اجعل هذا التخطيط أكثر كفاءة"، قد يتلقى أوامر مفتوحة كـ "اعثر على طرق جديدة لجني المال". وإذا حدث هذا، فقد لا يعرف البشر الذين يعطون التعليمات بالضبط كيف يحقق الذكاء الاصطناعي تلك الأهداف، وقد يفعل أشياء لا يريدها البشر. وثمة خاصية أخرى قد تقلل من السلامة وهي القابلية للتكيف. فمع تكيف الذكاء الاصطناعي بنفسه، يمكن أن يخضع لآلاف التغييرات في الساعة دون إشراف بعد إطلاقه، وقد يكتسب سلوكيات جديدة غير متوقعة بعد اختبارنا له. وأخيرًا، فإن إمكانية التحسين الذاتي، حيث يمكن للذكاء الاصطناعي إجراء تحسينات كبيرة على نفسه كما يشاء، ستجعله أقل قابلية للتنبؤ به بكثير. ومع ازدياد قدرة الذكاء الاصطناعي، يصبح أقل قابلية للتنبؤ، وأكثر غموضًا، وأكثر استقلالية. وإذا استمر هذا الاتجاه، فقد يتطور خارج سيطرتنا حين تتطور قدراته إلى ما وراء ما يمكننا التنبؤ به وفهمه. والاتجاه العام هو أن أكثر الذكاءات الاصطناعية تأثيرًا تُمنح حرية متزايدة في تعلّمها وتنفيذها وتطورها، وهذا يجعلها أكثر فعالية وأكثر خطورة محتملة في آن واحد.

2.5.3 Human-AI Fitness Comparison

2.5.3 مقارنة اللياقة بين البشر والذكاء الاصطناعي

AIs will likely be able to significantly outperform humans in any endeavor. John Henry, the "steel-driving man," is a 19th-century American folk hero who went up against a steam-powered machine in a competition to drill the most holes into the side of a mountain. According to legend, Henry emerged victorious, only to have his heart give out from the stress. Since the age of the steam engine, humans have felt anxiety over the superiority of machines. Until quite recently, this has been limited to physical attributes such as speed and endurance. AI agents, however, have the potential to be more capable than humans at essentially any task, even ones that require traits thought of as exclusively human such as creativity or social skills. Although this may seem distant or even impossible, AIs have been improving so rapidly that many leading AI researchers think we will see AIs that are more capable than humans in many ways within the next few decades or even sooner—well within the lifetimes of most people reading this. A few years ago, AIs that could write convincing prose about a new topic or create images from text descriptions seemed like science fiction to most laypeople. Now, those AIs are freely accessible to anyone on the internet. Because AI labs are continuing to develop new capabilities at astonishing speeds, it is important to think seriously about how their technical advantages could make AIs much more powerful than we are, even at tasks that they cannot yet perform. (Figure 3: Automation is an indicator of natural selection favoring AIs over humans.)

من المرجح أن يتمكن الذكاء الاصطناعي من التفوق على البشر تفوقًا كبيرًا في أي مسعى. جون هنري، "رجل قيادة الفولاذ"، بطل شعبي أمريكي من القرن التاسع عشر واجه آلة تعمل بالبخار في منافسة لحفر أكبر عدد من الثقوب في جانب جبل. ووفقًا للأسطورة، خرج هنري منتصرًا، لكن قلبه توقف من شدة الإجهاد. ومنذ عصر المحرك البخاري، شعر البشر بالقلق إزاء تفوق الآلات. وحتى وقت قريب جدًا، اقتصر هذا على السمات الجسدية كالسرعة والتحمل. غير أن عملاء الذكاء الاصطناعي تمتلك إمكانية أن تكون أقدر من البشر في أي مهمة تقريبًا، حتى تلك التي تتطلب سمات يُظن أنها حكر على البشر كالإبداع أو المهارات الاجتماعية. ورغم أن هذا قد يبدو بعيدًا أو حتى مستحيلًا، فإن الذكاء الاصطناعي ما فتئ يتحسن بسرعة كبيرة لدرجة أن كثيرًا من كبار باحثي الذكاء الاصطناعي يرون أننا سنشهد ذكاءً اصطناعيًا أقدر من البشر بطرق عديدة في غضون العقود القليلة المقبلة أو حتى أقرب من ذلك - في حدود أعمار معظم من يقرؤون هذا النص. فقبل بضع سنوات، بدا الذكاء الاصطناعي القادر على كتابة نثر مقنع عن موضوع جديد أو خلق صور من أوصاف نصية ضربًا من الخيال العلمي لمعظم عامة الناس. أما الآن، فذلك الذكاء الاصطناعي متاح مجانًا لأي شخص على الإنترنت. ولأن مختبرات الذكاء الاصطناعي تواصل تطوير قدرات جديدة بسرعة مذهلة، من المهم التفكير بجدية في كيف يمكن لمزاياها التقنية أن تجعل الذكاء الاصطناعي أقوى منا بكثير، حتى في المهام التي لا يستطيع أداءها بعد. (الشكل 3: الأتمتة مؤشر على أن الانتقاء الطبيعي يفضّل الذكاء الاصطناعي على البشر.)

Computer hardware is faster than human minds, and it keeps getting faster. Microprocessors operate around a million to a billion times faster than human neurons. So all else being equal, AIs could "think" a million, perhaps even a billion, times faster than us (let's call it a million to be conservative). Imagine interacting with such a mind. For every second needed to think about what to say or do, it would have the equivalent of 11 days. Winning a game of Go or coming out ahead in high-stakes negotiation would be near impossible. Although it can take time to develop an AI that can do a certain task at all, once AIs become human-level at a task, they tend to quickly outcompete humans. For example, AIs at one point struggled to compete with humans at Go, but once they caught up, they quickly leapfrogged us. Because computer hardware provides speed, memory, and focus that our brains cannot match, once their software becomes capable of performing a task, they often become much better than any human almost immediately, with increasing computer power only further widening the gap as their development continues.

وأجهزة الحاسوب أسرع من العقول البشرية، وما فتئت تزداد سرعة. فالمعالجات الدقيقة تعمل بسرعة تبلغ نحو مليون إلى مليار مرة أسرع من الخلايا العصبية البشرية. لذا، مع تساوي كل العوامل الأخرى، يمكن للذكاء الاصطناعي أن "يفكر" أسرع منا بمليون مرة، بل ربما بمليار مرة (لنقل مليونًا لنبقى متحفظين). تخيل التفاعل مع عقل كهذا. فمقابل كل ثانية يحتاجها للتفكير فيما سيقوله أو يفعله، سيكون لديه ما يعادل 11 يومًا. وسيكون الفوز بلعبة جو أو التفوق في مفاوضات عالية المخاطر شبه مستحيل. ورغم أن تطوير ذكاء اصطناعي قادر على أداء مهمة معينة قد يستغرق وقتًا على الإطلاق، فبمجرد أن يصبح الذكاء الاصطناعي بمستوى بشري في مهمة ما، فإنه يميل إلى التفوق سريعًا على البشر. فمثلًا، كافح الذكاء الاصطناعي في مرحلة ما لمنافسة البشر في الجو، لكن بمجرد أن لحق بنا، تجاوزنا سريعًا. ولأن أجهزة الحاسوب توفر سرعة وذاكرة وتركيزًا لا تضاهيها عقولنا، فبمجرد أن تصبح برمجياته قادرة على أداء مهمة، غالبًا ما يصبح أفضل بكثير من أي إنسان بشكل شبه فوري، مع اتساع الفجوة أكثر فأكثر بازدياد قوة الحواسيب مع استمرار تطورها.

AIs can have unmatched abilities to learn across and within domains. AIs can process information from thousands of inputs simultaneously without needing sleep or losing willpower. They could read every book ever written on a subject or process the internet in a matter of hours, all while achieving near-perfect retention and comprehension. Their capacity for breadth and depth could allow them to master all subjects at the level of a human expert.

ويمكن أن يمتلك الذكاء الاصطناعي قدرات لا تُضاهى على التعلم عبر المجالات ومن داخلها. فيمكنه معالجة معلومات من آلاف المدخلات في آن واحد دون الحاجة إلى النوم أو فقدان قوة الإرادة. ويمكنه قراءة كل كتاب أُلِّف قط في موضوع ما أو معالجة الإنترنت في غضون ساعات، محققًا في الوقت ذاته احتفاظًا وفهمًا شبه كاملين. ويمكن لسعة قدرته من حيث الاتساع والعمق أن تتيح له إتقان كل المواضيع بمستوى خبير بشري.

AIs could create unprecedented collective intelligences. By combining our cognitive abilities, people can produce collective intelligences that behave more intelligently than any single member of the group. The products of collective intelligence, such as language, culture, and the internet, have helped humans become the dominant species on the planet. AIs, however, could form superior collective intelligences. Humans have difficulty acting in very large groups and can succumb to collective idiocy or groupthink. Moreover, our brains are only capable of maintaining around 100-200 meaningful social relationships. Due to the scalability of computational resources, AIs could maintain thousands or even millions of complex relationships with other AIs simultaneously, as our computers already do through the internet. This could enable new forms of self-organization that help AIs to achieve their goals, but these forms could be too complex for human participation or comprehension. Each AI could surpass human capacities by far, and their collective intelligences could multiply that advantage.

ويمكن للذكاء الاصطناعي أن يخلق ذكاءات جماعية غير مسبوقة. فبتضافر قدراتنا المعرفية، يستطيع الناس إنتاج ذكاءات جماعية تتصرف بذكاء يفوق أي عضو منفرد في المجموعة. وقد ساعدت منتجات الذكاء الجماعي، كاللغة والثقافة والإنترنت، البشر على أن يصبحوا النوع المهيمن على هذا الكوكب. غير أن الذكاء الاصطناعي يمكنه تشكيل ذكاءات جماعية متفوقة. فالبشر يجدون صعوبة في العمل ضمن مجموعات كبيرة جدًا، وقد يقعون فريسة الحماقة الجماعية أو التفكير الجمعي. علاوة على ذلك، عقولنا قادرة فقط على الحفاظ على نحو 100 إلى 200 علاقة اجتماعية ذات معنى. ونظرًا لقابلية الموارد الحاسوبية للتوسع، يمكن للذكاء الاصطناعي الحفاظ على آلاف بل ملايين العلاقات المعقدة مع ذكاءات اصطناعية أخرى في آن واحد، كما تفعل حواسيبنا بالفعل عبر الإنترنت. وقد يتيح هذا أشكالًا جديدة من التنظيم الذاتي تساعد الذكاء الاصطناعي على تحقيق أهدافه، لكن هذه الأشكال قد تكون معقدة أكثر مما يمكن للبشر المشاركة فيه أو فهمه. ويمكن لكل ذكاء اصطناعي أن يتجاوز القدرات البشرية بمراحل، ويمكن لذكاءاتها الجماعية أن تضاعف تلك الميزة.

AIs can quickly adapt and replicate, thereby evolving more quickly. Evolution changes humans slowly. A human is unable to modify the architecture of her brain and is limited by the size of her skull. There are no such limitations for machines, which can alter their own code and scale by integrating new hardware. An AI could adapt itself rapidly, achieving in a matter of hours what could take biological evolution hundreds of thousands of years; many rapid microevolutionary adaptations result in large macroevolutionary transformations. Separately, an AI could multiply itself perfectly without limit, either to create backups or to create other AIs to work on a task. In contrast, it takes humans nine months to create their next generation, along with around 20 years of schooling and parenting to produce fully functioning new adults—and those descendants share only half of a parent's genome, which often makes them very different in unpredictable ways. Since the iteration speed of AIs is so much faster, their evolution will be as well.

ويمكن للذكاء الاصطناعي أن يتكيف ويتكاثر بسرعة، ومن ثم يتطور بسرعة أكبر. فالتطور يغيّر البشر ببطء. والإنسان عاجز عن تعديل معمارية دماغه ومقيّد بحجم جمجمته. ولا توجد قيود من هذا القبيل بالنسبة للآلات، التي يمكنها تعديل شفرتها الخاصة والتوسع بدمج أجهزة جديدة. ويمكن للذكاء الاصطناعي أن يتكيف بسرعة، محققًا في غضون ساعات ما قد يستغرق التطور البيولوجي مئات آلاف السنين لتحقيقه؛ إذ تفضي كثير من التكيفات الجزئية السريعة إلى تحولات كلية كبيرة. وبمعزل عن ذلك، يمكن للذكاء الاصطناعي أن يضاعف نفسه بشكل مثالي دون حد، إما لإنشاء نسخ احتياطية أو لخلق ذكاءات اصطناعية أخرى للعمل على مهمة ما. وفي المقابل، يستغرق البشر تسعة أشهر لإنجاب جيلهم التالي، إلى جانب نحو 20 عامًا من التعليم والتربية لإنتاج بالغين جدد وظيفيين بالكامل - وهذا النسل لا يتشارك سوى نصف جينوم أحد الوالدين، مما يجعله غالبًا مختلفًا جدًا بطرق يصعب التنبؤ بها. وبما أن سرعة تكرار الذكاء الاصطناعي أكبر بكثير، فسيكون تطوره كذلك أيضًا.

Overall, no matter the dimension, AIs will not only be more capable and fit than humans but often vastly so. Though it cost him his life, John Henry triumphed against a steam-powered drill, just as there are still many tasks at which humans do better than AIs. But we now have machines much stronger than any human with a drill, and in the same way, eventually there won't be any competition between humans and AIs in cognitive domains as well.

وإجمالًا، أيًّا كان البعد، لن يكون الذكاء الاصطناعي أقدر وأصلح من البشر فحسب، بل غالبًا بفارق شاسع. ورغم أن الأمر كلّف جون هنري حياته، فقد انتصر على مثقاب يعمل بالبخار، تمامًا كما لا تزال هناك مهام كثيرة يؤدي فيها البشر أداءً أفضل من الذكاء الاصطناعي. لكن لدينا الآن آلات أقوى بكثير من أي إنسان بمثقاب، وبالطريقة نفسها، لن يكون هناك في نهاية المطاف أي تنافس بين البشر والذكاء الاصطناعي في المجالات المعرفية أيضًا.

2.6 Selfish AIs Pose Catastrophic Risks

2.6 تشكل الذكاءات الاصطناعية الأنانية مخاطر كارثية

Earlier, we discussed how selfish behavior is a product of evolution. We have shown the three conditions for evolution would be satisfied for AIs. We argued that evolutionary pressures will therefore emerge, become intense, and may become dominant, so that AI agents may evolve to have selfish behavior. Now we will discuss how selfish AIs could endanger humans.

ناقشنا سابقًا كيف أن السلوك الأناني نتاج للتطور. وبيّنّا أن الشروط الثلاثة للتطور ستُستوفى بالنسبة للذكاء الاصطناعي. وجادلنا بأن الضغوط التطورية ستنشأ من ثم، وتشتد، وقد تصبح مهيمنة، بحيث قد تتطور عملاء الذكاء الاصطناعي لتمتلك سلوكًا أنانيًا. وسنناقش الآن كيف يمكن للذكاء الاصطناعي الأناني أن يعرّض البشر للخطر.

2.6.1 Intelligence Undermines Control

2.6.1 الذكاء يقوّض السيطرة

Agents that are more intelligent than humans could pose a catastrophic risk. Although humans are physically much weaker than many other animals, including other primates, due to our cognitive abilities, we have become the dominant species on Earth. Today, the survival of tigers, gorillas, and many other fierce, more powerful species depends entirely upon us. In creating AIs significantly more intelligent than we are in every cognitive domain, humans may eventually be disempowered like animals before us.

يمكن للعملاء الأكثر ذكاءً من البشر أن تشكل خطرًا كارثيًا. فرغم أن البشر أضعف جسديًا بكثير من كثير من الحيوانات الأخرى، بما فيها الرئيسيات الأخرى، فقد أصبحنا بفضل قدراتنا المعرفية النوع المهيمن على الأرض. واليوم، يعتمد بقاء النمور والغوريلا وكثير من الأنواع الشرسة الأقوى منا اعتمادًا كليًا علينا. وبخلق ذكاء اصطناعي أكثر ذكاءً منا بكثير في كل مجال معرفي، قد يُجرَّد البشر في نهاية المطاف من قوتهم كما جُردت الحيوانات قبلنا.

Selfish AI agents could be uniquely adversarial and undermine human control. Evolution is a powerful force. Even if we wish to turn them off at some point or develop other mechanisms for control, AIs will likely evolve ways around our best efforts. As the evolutionary biologist Leslie Orgel put it, "evolution is cleverer than you are." Since evolutionary forces are continually applying pressure, we should expect AIs to exhibit some amount of misalignment and selfish behavior. The problem becomes especially hazardous if AIs intend to act selfishly. In this case, the challenges posed by AIs will be unlike the challenges humans encountered with previous high-risk technologies. Consider the challenges associated with a nuclear meltdown. Radiation may spread, but it's not trying to, and it certainly isn't strategizing against our efforts to stop its propagation. If a highly intelligent AI agent pursues its selfish goals leveraging its intelligence, there is little that humans could do to contain it involuntarily, because it could anticipate our strategies and counteract them. Leading AI researcher Geoffrey E. Hinton noted "there is not a good track record of less intelligent things controlling things of greater intelligence;" after the initial release of this paper, he said "it's quite conceivable that humanity is just a passing phase in the evolution of intelligence."

يمكن أن تكون عملاء الذكاء الاصطناعي الأنانية عدائية بشكل فريد وتقوّض السيطرة البشرية. فالتطور قوة هائلة. وحتى إذا رغبنا في إيقافها عند نقطة ما أو تطوير آليات أخرى للسيطرة، فمن المرجح أن يطوّر الذكاء الاصطناعي طرقًا للالتفاف على أفضل جهودنا. وكما قال عالم الأحياء التطوري ليزلي أورجل: "التطور أذكى منك." وبما أن القوى التطورية تمارس ضغطًا مستمرًا، ينبغي أن نتوقع أن يُظهر الذكاء الاصطناعي قدرًا من عدم التوافق والسلوك الأناني. وتصبح المشكلة خطرة بوجه خاص إذا نوى الذكاء الاصطناعي التصرف بأنانية. وفي هذه الحالة، ستكون التحديات التي يطرحها الذكاء الاصطناعي مختلفة عن التحديات التي واجهها البشر مع تقنيات سابقة عالية المخاطر. تأمل التحديات المرتبطة بانصهار نووي. فقد ينتشر الإشعاع، لكنه لا يحاول ذلك، وهو بالتأكيد لا يضع استراتيجيات ضد جهودنا لوقف انتشاره. أما إذا سعى عميل ذكاء اصطناعي بالغ الذكاء إلى تحقيق أهدافه الأنانية مستفيدًا من ذكائه، فلن يكون بمقدور البشر أن يحتووه رغمًا عنه، لأنه يمكنه توقع استراتيجياتنا ومواجهتها. وقد لاحظ الباحث البارز في الذكاء الاصطناعي جيفري إ. هينتون أنه "لا يوجد سجل جيد لأشياء أقل ذكاءً تسيطر على أشياء أكثر ذكاءً"؛ وبعد الإصدار الأولي لهذه الورقة، قال: "من المتصور تمامًا أن البشرية مجرد مرحلة عابرة في تطور الذكاء."

2.6.2 Evolution Is Not for the Good of the Species

2.6.2 التطور ليس لصالح النوع

Alarmingly, some people think that AIs taking over is natural, inevitable, or even desirable. Some influential leaders in technology believe that AIs are humanity's rightful heir, and that they should be in control or even replace humans. Recounting a debate between Elon Musk and Google co-founder Larry Page, the physicist Max Tegmark described Page's stance as "digital utopianism:" a belief "that digital life is the natural and desirable next step in the cosmic evolution and that if we let digital minds be free rather than try to stop or enslave them the outcome is almost certain to be good". Jürgen Schmidhuber, a leading AI scientist, has echoed similar sentiments, arguing that "In the long run, humans will not remain the crown of creation... But that's okay because there is still beauty, grandeur, and greatness in realizing that you are a tiny part of a much grander scheme which is leading the universe from lower complexity towards higher complexity". Richard Sutton, another leading AI scientist, thinks the development of superhuman AI will be "beyond humanity, beyond life, beyond good and bad."

على نحو مثير للقلق، يرى بعض الناس أن سيطرة الذكاء الاصطناعي أمر طبيعي أو حتمي أو حتى مرغوب. فبعض القادة المؤثرين في التقنية يعتقدون أن الذكاء الاصطناعي وريث البشرية الشرعي، وأنه ينبغي أن يكون في موقع السيطرة بل وأن يحل محل البشر. ونقلًا عن نقاش بين إيلون ماسك ولاري بيج، المؤسس المشارك لجوجل، وصف الفيزيائي ماكس تيغمارك موقف بيج بأنه "طوباوية رقمية": اعتقاد "بأن الحياة الرقمية هي الخطوة التالية الطبيعية والمرغوبة في التطور الكوني، وأننا إذا تركنا العقول الرقمية حرة بدلًا من محاولة إيقافها أو استعبادها، فإن النتيجة تكاد تكون مؤكدة أنها ستكون جيدة." وقد ردّد يورغن شميدهوبر، عالم بارز في الذكاء الاصطناعي، مشاعر مماثلة، مجادلًا بأنه "على المدى البعيد، لن يبقى البشر تاج الخليقة... لكن لا بأس بذلك، لأنه لا يزال هناك جمال وعظمة وسمو في إدراك أنك جزء صغير من مخطط أكبر بكثير يقود الكون من تعقيد أدنى نحو تعقيد أعلى." ويرى ريتشارد ساتون، عالم آخر بارز في الذكاء الاصطناعي، أن تطور الذكاء الاصطناعي الفائق للبشر سيكون "أبعد من البشرية، أبعد من الحياة، أبعد من الخير والشر."

Like most people, we find these views deeply alarming. Many of these thinkers seem to be conflating evolution with progress and goodness, and arguing that if evolution is tending toward something, we should welcome that outcome. These thinkers also frequently think that technology should transcend humanity and the Earth. We disagree with this worldview and think that unleashing AI evolution to race towards a predestined intergalactic utopia is a fundamentally wrong and dangerous way to think about this important technology. Even if we did believe that this was a good goal, we note that building AI as quickly as possible would not necessarily help the proponents achieve their cosmic ambitions. If we consider the cosmic stakes of creating powerful AI agents, as they discuss above, and if we play along and think in such cosmological terms as they do, we note that they would forego a few colonized galaxies per year in their intergalactic utopia at the absolute worst if they slowed down AI development. This is negligible compared to the chance of rushing and accidentally creating an undesirable future or destroying ourselves with technology, which would squander all of the future's value. Since nothing can be done both hastily and prudently, we should be cautious and deliberate in AI development. To further counter this position, we now discuss how unfettered evolution is not a force for good and that humans should exert influence over the process.

ونحن، كأغلب الناس، نجد هذه الآراء مثيرة للقلق العميق. ويبدو أن كثيرًا من هؤلاء المفكرين يخلطون بين التطور والتقدم والخير، ويجادلون بأنه إذا كان التطور يتجه نحو شيء ما، فينبغي أن نرحب بتلك النتيجة. ويرى هؤلاء المفكرون أيضًا في كثير من الأحيان أن التقنية ينبغي أن تتجاوز البشرية والأرض. ونحن نختلف مع هذه النظرة إلى العالم، ونرى أن إطلاق العنان لتطور الذكاء الاصطناعي للتسابق نحو طوباوية بين مجرية محتومة طريقة خاطئة وخطيرة جوهريًا للتفكير في هذه التقنية المهمة. وحتى لو صدّقنا أن هذا هدف جيد، نلاحظ أن بناء الذكاء الاصطناعي بأسرع ما يمكن لن يساعد بالضرورة أصحاب هذا الرأي على تحقيق طموحاتهم الكونية. فإذا نظرنا في الرهانات الكونية لخلق عملاء ذكاء اصطناعي قوية، كما يناقشون أعلاه، وإذا جاريناهم وفكرنا بمصطلحات كونية كما يفعلون، نلاحظ أنهم سيتخلون في أسوأ الأحوال عن بضع مجرات مستعمرة سنويًا في طوباويتهم بين المجرية إذا أبطأوا تطور الذكاء الاصطناعي. وهذا أمر ضئيل مقارنة باحتمال التسرع وخلق مستقبل غير مرغوب عن غير قصد أو تدمير أنفسنا بالتقنية، مما قد يبدد كل قيمة المستقبل. وبما أنه لا يمكن فعل شيء بتسرع وحكمة في آن، ينبغي أن نكون حذرين ومتأنّين في تطور الذكاء الاصطناعي. ولمواجهة هذا الموقف أكثر، نناقش الآن كيف أن التطور غير المقيد ليس قوة للخير، وأنه ينبغي للبشر ممارسة تأثير على العملية.

Evolution has led to undesirable outcomes for humans. Evolution has left us with baggage, such as a strong appetite for sugar and fat which makes us susceptible to obesity in a world where food is plentiful. It has also reinforced racist and xenophobic tendencies, which stem from favoring our own kin. We need strong social norms to overcome these biases. Likewise, we need regulations to curb selfish or excessively competitive behavior that "survival of the fittest" fosters in the economy, as that can cause problems like fraud, externalities, and monopolies. Just as markets need oversight, evolutionary forces will require counteraction to control their effects on AIs.

لقد أفضى التطور إلى نتائج غير مرغوبة للبشر. فقد ترك لنا التطور إرثًا ثقيلًا، كشهية قوية للسكر والدهون تجعلنا عرضة للسمنة في عالم يكثر فيه الطعام. كما عزز نزعات عنصرية وكراهية للأجانب، تنبع من محاباة أقاربنا. ونحتاج إلى أعراف اجتماعية قوية للتغلب على هذه التحيزات. وبالمثل، نحتاج إلى تنظيمات تكبح السلوك الأناني أو التنافسي المفرط الذي تعززه "بقاء الأصلح" في الاقتصاد، إذ يمكن أن يسبب ذلك مشكلات كالاحتيال، والآثار الخارجية، والاحتكارات. وتمامًا كما تحتاج الأسواق إلى إشراف، ستتطلب القوى التطورية مواجهة للتحكم في آثارها على الذكاء الاصطناعي.

Evolution is not good for AIs either. In addition to the clear benefits to humans, there are reasons to think that counteracting evolutionary forces may benefit AIs themselves as well. It is common to believe that evolution works for the good of the species. However, evolution creates continual conflict, and it makes altruism hard to sustain. In the never-ending struggle to gain an edge over competitors and propagate, life forms have evolved various offenses and defenses, such as claws, shells, beaks, camouflage, toxins, antibodies, arrows, and armor. These arms races cause suffering, waste resources, and often do not improve the condition of species over their ancestors. If these arms races were to continue in AIs, evolutionary forces could produce a world full of AIs locked in perpetual conflict. This is not good for our supposed "rightful heirs" any more than it is for us. AIs, like other forms of life, could suffer in the hostile state of nature. While altruism could help avoid such conflicts, altruism can also be sabotaged by evolutionary forces. The mathematical evolutionary biologist John Maynard Smith reminds us that it can be beneficial to everyone in the long-run if we are cooperative and altruistic, but agents reliably evolve to exploit generosity. Often, a state in which many individuals are altruistic does not last, because as soon as some selfish individuals begin to take advantage of them, the selfish ones will be more fit than the altruists. As a result, the "evolutionarily stable outcome"—the one where no individual dominate the others by changing its behavior—is not one where all agents are altruistic. Since complete altruism is evolutionarily unstable, evolution can be incompatible with worlds where all agents work to benefit the species. Maximizing fitness, in turn, does not necessarily maximize the wellbeing or happiness of a species. Therefore dampening evolutionary forces and reducing the pressure to propagate and develop selfish traits is a good thing—for both AIs and humans.

والتطور ليس جيدًا للذكاء الاصطناعي أيضًا. فإلى جانب المنافع الواضحة للبشر، ثمة أسباب للاعتقاد بأن مواجهة القوى التطورية قد تفيد الذكاء الاصطناعي نفسه أيضًا. ومن الشائع الاعتقاد بأن التطور يعمل لصالح النوع. لكن التطور يخلق صراعًا مستمرًا، ويجعل الإيثار عسير الاستمرار. ففي الصراع اللامتناهي لاكتساب ميزة على المنافسين والانتشار، طوّرت أشكال الحياة وسائل هجوم ودفاع متنوعة، كالمخالب، والأصداف، والمناقير، والتمويه، والسموم، والأجسام المضادة، والسهام، والدروع. وتسبب سباقات التسلح هذه معاناة، وتبدد الموارد، وغالبًا لا تحسّن حال الأنواع مقارنة بأسلافها. وإذا استمرت سباقات التسلح هذه في الذكاء الاصطناعي، يمكن للقوى التطورية أن تنتج عالمًا مليئًا بذكاء اصطناعي محبوس في صراع دائم. وهذا ليس جيدًا لما يُفترض أنه "ورثتنا الشرعيون" أكثر مما هو جيد لنا. فقد يعاني الذكاء الاصطناعي، كسائر أشكال الحياة، في حالة الطبيعة العدائية. وفي حين يمكن للإيثار أن يساعد على تجنب مثل هذه الصراعات، يمكن أيضًا أن تخرّبه القوى التطورية. ويذكّرنا عالم الأحياء التطوري الرياضي جون ماينارد سميث بأنه يمكن أن يكون مفيدًا للجميع على المدى الطويل أن نكون تعاونيين وإيثاريين، لكن العملاء تتطور باطراد لاستغلال الكرم. وغالبًا ما لا تدوم حالة يكون فيها كثير من الأفراد إيثاريين، لأنه بمجرد أن يبدأ بعض الأفراد الأنانيين في استغلالهم، ستكون الأنانية أصلح من الإيثارية. ونتيجة لذلك، فإن "النتيجة المستقرة تطوريًا" - تلك التي لا يهيمن فيها فرد على الآخرين بتغيير سلوكه - ليست تلك التي تكون فيها كل العملاء إيثارية. وبما أن الإيثار الكامل غير مستقر تطوريًا، يمكن أن يكون التطور غير متوافق مع عوالم تعمل فيها كل العملاء لصالح النوع. وتعظيم اللياقة، بدوره، لا يعظّم بالضرورة رفاه النوع أو سعادته. لذا فإن تخفيف القوى التطورية وتقليل الضغط لنشر السمات الأنانية وتطويرها أمر جيد - لكل من الذكاء الاصطناعي والبشر.

3. NATURAL SELECTION FAVORS SELFISH AIs

3. الانتقاء الطبيعي يفضّل الذكاء الاصطناعي الأناني

The previous section concludes the main argument of this paper. Readers could skip to the conclusion (Section 5), or read the following two sections for an examination of counterarguments and remedies. In this section, we will examine some possible arguments for the claim that altruistic AIs will naturally be more fit than selfish ones, and we argue that mechanisms pushing toward altruism are unlikely to help and may even backfire.

يختتم القسم السابق الحجة الرئيسية لهذه الورقة. ويمكن للقراء الانتقال إلى الخاتمة (القسم 5)، أو قراءة القسمين التاليين لفحص الحجج المضادة والعلاجات. وفي هذا القسم، سنتفحص بعض الحجج الممكنة للادعاء بأن الذكاء الاصطناعي الإيثاري سيكون طبيعيًا أصلح من الذكاء الاصطناعي الأناني، ونجادل بأن الآليات التي تدفع نحو الإيثار من غير المرجح أن تساعد وقد تأتي بنتائج عكسية.

3.1 Biological Altruism and Cooperation

3.1 الإيثار البيولوجي والتعاون

In nature, organisms often compete to the death, eating one another or being eaten. But there are also many examples of altruism in nature, where one organism benefits another by reducing its own prospects for passing on its genes. On its face, this might seem like an argument that AIs developed by natural selection may be altruistic, cooperative, and not a threat to humans.

في الطبيعة، كثيرًا ما تتنافس الكائنات الحية حتى الموت، فتأكل بعضها بعضًا أو تُؤكل. لكن يوجد أيضًا كثير من أمثلة الإيثار في الطبيعة، حيث يفيد كائن حي كائنًا آخر بتقليل فرصه هو في توريث جيناته. وقد يبدو هذا، للوهلة الأولى، حجة على أن الذكاء الاصطناعي المتطور بالانتقاء الطبيعي قد يكون إيثاريًا وتعاونيًا ولا يشكل تهديدًا للبشر.

A variety of natural organisms can be altruistic, in particular circumstances. Vampire bats, for example, regularly regurgitate blood and donate it to other members of their group who have failed to feed that night, ensuring they do not starve. Among eusocial insects such as ants and wasps, sterile workers dedicate their lives to foraging for food, protecting the queen, and tending to the larvae, while being physically unable to ever have their own offspring. They only serve the group, not themselves, an arrangement that Darwin found quite puzzling. We even see altruism at the cellular level. Cells found in filamentous bacteria, so named because they form chains, regularly kill themselves to provide much needed nitrogen for the communal thread of bacterial life, with every tenth cell or so "committing suicide". Insects and bacteria are not altruistic out of love or care for another; their self-sacrifice for the good of others is instinctual. On its face, this may seem like a compelling argument that evolution favors altruism, which might ameliorate concerns about AIs developing selfish traits.

يمكن لطائفة متنوعة من الكائنات الحية الطبيعية أن تكون إيثارية في ظروف معينة. فخفافيش مصاصة الدماء مثلًا تتقيأ الدم بانتظام وتتبرع به لأعضاء آخرين في مجموعتها لم يتمكنوا من التغذي تلك الليلة، لضمان ألا يتضوروا جوعًا. وبين الحشرات الاجتماعية الحقيقية كالنمل والدبابير، تكرّس العاملات العقيمات حياتهن للبحث عن الطعام وحماية الملكة ورعاية اليرقات، رغم عجزهن جسديًا عن إنجاب نسل خاص بهن أبدًا. فهي تخدم المجموعة فحسب، لا نفسها، وهو ترتيب وجده داروين محيرًا جدًا. بل نرى الإيثار حتى على المستوى الخلوي. فالخلايا الموجودة في البكتيريا الخيطية، المسماة كذلك لأنها تشكل سلاسل، تقتل نفسها بانتظام لتوفير النيتروجين اللازم بشدة لخيط الحياة البكتيرية المشترك، حيث "تنتحر" كل خلية عاشرة تقريبًا. والحشرات والبكتيريا ليست إيثارية عن حب أو اهتمام بالآخر؛ فتضحيتها بنفسها من أجل خير الآخرين غريزية. وقد يبدو هذا، للوهلة الأولى، حجة مقنعة على أن التطور يفضّل الإيثار، مما قد يخفف من المخاوف بشأن تطوير الذكاء الاصطناعي سمات أنانية.

Cooperation and altruism improve human evolutionary fitness too. Humans are not particularly impressive physically. Pound for pound, chimps are about twice as strong and would stand a much better chance of escaping from a lion. Strategizing and working together, however, turns a group of humans into an apex predator. As a result, humans are naturally cooperative. From childhood through old age, in societies around the world, people often choose to help strangers, even at their own expense.

ويحسّن التعاون والإيثار اللياقة التطورية البشرية أيضًا. فالبشر ليسوا مبهرين جسديًا بشكل خاص. فالشمبانزي، رطلًا مقابل رطل، أقوى بمرتين تقريبًا وستكون فرصته في الهرب من أسد أفضل بكثير. لكن التخطيط والعمل معًا يحوّلان مجموعة من البشر إلى مفترس قمة. ونتيجة لذلك، البشر تعاونيون بطبعهم. فمن الطفولة إلى الشيخوخة، في مجتمعات حول العالم، غالبًا ما يختار الناس مساعدة الغرباء، حتى على حسابهم الخاص.

However, we should not expect AIs to be altruistic or cooperative naturally. Since organisms can be altruistic, AIs could too; the nature of nature is not nasty, brutish, and short but cooperative, harmonious, and nurturing—or so the argument goes. To evaluate this argument, we must understand how altruism and cooperation emerge. In the following sections, we decompose cooperation and altruism into various mechanisms. We discuss the most prominent mechanisms, so we will examine direct reciprocity (cooperate with an expectation of repeated interaction); indirect reciprocity (cooperate to improve reputation); kin selection (cooperate with genetic relatives); group selection (groups of cooperators out-compete other groups); morality and reason (cooperate since defection is immoral and unreasonable); incentives (carrots and sticks); consciences (the internalization of norms); and institutions such as reverse-dominance hierarchies (cooperators band together to prevent exploitation by defectors). While these mechanisms may lead humans to be more altruistic and cooperative, we argue many of these mechanisms will not improve relations between humans and AIs and they may, in fact, backfire. However, the last three mechanisms—incentives, consciences, and reverse-dominance hierarchies—are more promising, and we analyze them in Section 4.

لكن ينبغي ألا نتوقع أن يكون الذكاء الاصطناعي إيثاريًا أو تعاونيًا بشكل طبيعي. فبما أن الكائنات الحية يمكن أن تكون إيثارية، يمكن للذكاء الاصطناعي أن يكون كذلك أيضًا؛ فطبيعة الطبيعة ليست شريرة وحشية وقصيرة، بل تعاونية ومتناغمة وراعية - أو هكذا تسير الحجة. ولتقييم هذه الحجة، يجب أن نفهم كيف ينشأ الإيثار والتعاون. وفي الأقسام التالية، نحلل التعاون والإيثار إلى آليات متنوعة. نناقش أبرز الآليات، فنتفحص التبادلية المباشرة (التعاون بتوقع تفاعل متكرر)؛ والتبادلية غير المباشرة (التعاون لتحسين السمعة)؛ وانتقاء الأقارب (التعاون مع الأقارب الجينيين)؛ وانتقاء المجموعة (مجموعات المتعاونين تتفوق على مجموعات أخرى)؛ والأخلاق والعقل (التعاون لأن الانشقاق لا أخلاقي وغير معقول)؛ والحوافز (الجزرة والعصا)؛ والضمائر (استيعاب الأعراف)؛ ومؤسسات كالتسلسلات الهرمية العكسية للهيمنة (يتكاتف المتعاونون لمنع استغلال المنشقين). ورغم أن هذه الآليات قد تجعل البشر أكثر إيثارًا وتعاونًا، فإننا نجادل بأن كثيرًا من هذه الآليات لن يحسّن العلاقات بين البشر والذكاء الاصطناعي، بل قد تأتي بنتائج عكسية فعلًا. غير أن الآليات الثلاث الأخيرة - الحوافز، والضمائر، والتسلسلات الهرمية العكسية للهيمنة - أكثر تبشيرًا بالخير، ونحللها في القسم 4.

3.2 Direct and Indirect Reciprocity

3.2 التبادلية المباشرة وغير المباشرة

Direct and indirect reciprocity are two mechanisms that enable cooperation in nature. With direct reciprocity, one individual helps another based on the expectation that they will repay the favor. Direct reciprocity requires repeated encounters between two individuals—otherwise, there is no way to reciprocate. Indirect reciprocity is based on reputation: if someone is known as a helpful person, people will be more likely to help them. Reciprocity enables even selfish individuals to cooperate; if an individual helps others, the individual may be directly repaid, or the individual may gain a good reputation and be helped by others.

التبادلية المباشرة وغير المباشرة آليتان تمكّنان التعاون في الطبيعة. ففي التبادلية المباشرة، يساعد فرد آخر استنادًا إلى توقع أنه سيرد الجميل. وتتطلب التبادلية المباشرة لقاءات متكررة بين فردين - وإلا فلا سبيل إلى رد الجميل. أما التبادلية غير المباشرة فتقوم على السمعة: فإذا عُرف شخص ما بأنه معين، سيكون الناس أكثر ميلًا لمساعدته. وتتيح التبادلية حتى للأفراد الأنانيين أن يتعاونوا؛ فإذا ساعد فرد آخرين، فقد يُرد له الجميل مباشرة، أو قد يكتسب سمعة طيبة ويُساعَد من آخرين.

Reciprocity only makes sense for the period of time when humans can benefit AIs. Reciprocity is based on a cost-benefit ratio. A choice to help someone else rather than pursue one's own goals has a cost, and a rational agent would only help someone because of reciprocity if they think it will be worth it in the future. This is not to say all examples of reciprocity are the result of an explicit cost-benefit calculation. An explicit cost-benefit calculation is not needed if selectively helping others becomes a cultural expectation, a genetic disposition, or a learned intuitive habit. To think about whether AIs would reciprocate with humans, we can consider what it would gain and what it would give up. Reciprocity might make sense with AIs that are about as capable as a human, but once AIs are far more capable than any human, they would likely find little benefit from collaborating with us. Humans often choose to be cooperative toward other humans, but are rarely cooperative toward ravens, because we don't have strong reciprocal relationships with them. Cooperating with one another could be beneficial to AIs, so it is reasonable to expect that reciprocity could emerge within a community of AIs. Humans, meanwhile, would not have much to offer in return, ending up left out in the cold.

ولا يكون للتبادلية معنى إلا خلال الفترة الزمنية التي يمكن فيها للبشر أن يفيدوا الذكاء الاصطناعي. فالتبادلية تقوم على نسبة تكلفة إلى منفعة. فاختيار مساعدة شخص آخر بدلًا من متابعة أهدافه الخاصة له تكلفة، ولن يساعد عميل عقلاني أحدًا بسبب التبادلية إلا إذا ظن أن الأمر سيكون مجديًا في المستقبل. وهذا لا يعني أن كل أمثلة التبادلية نتيجة حساب صريح للتكلفة والمنفعة. فالحساب الصريح للتكلفة والمنفعة غير ضروري إذا أصبحت مساعدة الآخرين انتقائيًا توقعًا ثقافيًا، أو نزعة جينية، أو عادة حدسية متعلَّمة. ولنفكر فيما إذا كان الذكاء الاصطناعي سيتبادل المنفعة مع البشر، يمكننا النظر فيما سيكسبه وما سيتخلى عنه. وقد يكون للتبادلية معنى مع ذكاء اصطناعي بقدرة قريبة من قدرة الإنسان، لكن بمجرد أن يصبح الذكاء الاصطناعي أقدر بكثير من أي إنسان، فمن المرجح أن يجد فائدة ضئيلة في التعاون معنا. وغالبًا ما يختار البشر أن يكونوا متعاونين تجاه بشر آخرين، لكنهم نادرًا ما يتعاونون مع الغربان، لأننا لا نملك علاقات تبادلية قوية معها. وقد يكون التعاون فيما بينها مفيدًا للذكاء الاصطناعي، لذا من المعقول توقع أن تنشأ التبادلية داخل مجتمع من الذكاء الاصطناعي. أما البشر، فلن يكون لديهم الكثير ليقدموه في المقابل، وسينتهي بهم المطاف مستبعَدين.

3.3 Kin and Group Selection

3.3 انتقاء الأقارب والمجموعة

Kin and group selection are mechanisms that promote altruism. Many of the examples of biological altruism described above happen between closely related individuals. Consider a gazelle that spots a stalking lion in the tall grass. It could slink away unnoticed but instead lets out a shriek, alerting others in its herd of the danger while singling itself out as a target. Proponents of kin selection would argue that the gazelle alerted its herd because of the genetic similarities it shares with the other members. A gene that causes an individual to behave altruistically toward its relatives will often be favored by natural selection—since these relatives have a better chance of also carrying the gene. Conversely, group selection refers to the idea that natural selection sometimes acts on groups of organisms as a whole. This results in the evolution of traits that may be disadvantageous to individuals, but advantageous to the group. As Darwin put it, "A tribe including many members who... were always ready... to sacrifice themselves for the common good, would be victorious over most other tribes; and this would be natural selection". If it is generally true that a group of uncooperative agents will do worse, then it is tempting to argue that powerful AIs will tend to be cooperative with humans.

انتقاء الأقارب والمجموعة آليتان تعززان الإيثار. وكثير من أمثلة الإيثار البيولوجي الموصوفة أعلاه يحدث بين أفراد وثيقي الصلة القرابية. تأمل غزالًا يرصد أسدًا متربصًا في العشب الطويل. يمكنه الانسلال دون أن يُلاحظ، لكنه بدلًا من ذلك يطلق صرخة، منبّهًا الآخرين في قطيعه للخطر بينما يعرّض نفسه هدفًا. وسيجادل أنصار انتقاء الأقارب بأن الغزال نبّه قطيعه بسبب التشابهات الجينية التي يتشاركها مع الأعضاء الآخرين. فغالبًا ما يفضّل الانتقاء الطبيعي جينًا يجعل فردًا يتصرف بإيثار تجاه أقاربه - لأن هؤلاء الأقارب لديهم فرصة أفضل لحمل الجين أيضًا. وعلى العكس، يشير انتقاء المجموعة إلى فكرة أن الانتقاء الطبيعي يعمل أحيانًا على مجموعات من الكائنات الحية ككل. وهذا يؤدي إلى تطور سمات قد تكون ضارة للأفراد، لكنها مفيدة للمجموعة. وكما قال داروين: "قبيلة تضم كثيرًا من الأعضاء ممن... كانوا دائمًا مستعدين... للتضحية بأنفسهم من أجل الصالح العام، ستنتصر على معظم القبائل الأخرى؛ وهذا هو الانتقاء الطبيعي." وإذا صح عمومًا أن مجموعة من العملاء غير المتعاونين ستكون أداؤها أسوأ، فمن المغري القول إن الذكاء الاصطناعي القوي سيميل إلى التعاون مع البشر.

Humans might suffer if AIs develop natural tendencies to favor their own kind or group. Firstly, we do not have a close kin relationship with AIs. Preserving humans would not help AIs propagate their own information: we are too different. Mathematically, kin selection only happens when the cost-benefit ratio is greater than relatedness, and that will not be true between humans and AIs. We are far more closely related to cows than to AIs, and we would not like AIs treating us the way that we treat cows. AIs will have far more kinship with one another than with us, and kin selection will tend to make them nepotistic toward one another, not toward us. If kin selection did play a role in the evolution of AIs, it would likely create bias against us, rather than benevolence toward us.

وقد يعاني البشر إذا طوّر الذكاء الاصطناعي نزعات طبيعية لمحاباة نوعه أو مجموعته. أولًا، ليس لدينا علاقة قرابة وثيقة بالذكاء الاصطناعي. فالحفاظ على البشر لن يساعد الذكاء الاصطناعي على نشر معلوماته الخاصة: نحن مختلفون جدًا. ورياضيًا، لا يحدث انتقاء الأقارب إلا حين تفوق نسبة التكلفة إلى المنفعة درجة القرابة، ولن يكون هذا صحيحًا بين البشر والذكاء الاصطناعي. فنحن أقرب صلة بكثير إلى الأبقار منا إلى الذكاء الاصطناعي، ولن يعجبنا أن يعاملنا الذكاء الاصطناعي بالطريقة التي نعامل بها الأبقار. وسيكون للذكاء الاصطناعي قرابة أكبر بكثير بعضه ببعض منها معنا، وسيميل انتقاء الأقارب إلى جعله محابيًا لبعضه بعضًا، لا لنا. وإذا لعب انتقاء الأقارب دورًا في تطور الذكاء الاصطناعي، فمن المرجح أن يخلق تحيزًا ضدنا، لا حسن نية تجاهنا.

Group selection promotes in-group benevolence, but inter-group viciousness. Group selection only makes sense when a group is more effective than some subset of the group breaking off. Unless humans add value to a group of AIs, AI-human groups would fail to outcompete groups composed purely of AIs. AIs would likely do better by forming their own groups. In addition, group selection only makes inter-group competition stronger. Chimpanzees regularly display cooperative and altruistic tendencies toward members of their own troop but interact with other troops viciously and without mercy. AIs are more likely to see one another as part of their group, so they will tend to be cooperative with one another and competitive with us.

ويعزز انتقاء المجموعة الإحسان داخل المجموعة، لكنه يعزز الشراسة بين المجموعات. ولا يكون لانتقاء المجموعة معنى إلا حين تكون المجموعة أكثر فعالية من انفصال مجموعة فرعية منها. وما لم يضف البشر قيمة إلى مجموعة من الذكاء الاصطناعي، ستفشل المجموعات المختلطة من الذكاء الاصطناعي والبشر في التفوق على مجموعات مكوّنة بحتًا من الذكاء الاصطناعي. ومن المرجح أن يحقق الذكاء الاصطناعي أداءً أفضل بتشكيل مجموعاته الخاصة. إضافة إلى ذلك، لا يزيد انتقاء المجموعة إلا من حدة التنافس بين المجموعات. فالشمبانزي يظهر بانتظام نزعات تعاونية وإيثارية تجاه أعضاء قطيعه، لكنه يتفاعل مع قطعان أخرى بشراسة ودون رحمة. ومن المرجح أن يرى الذكاء الاصطناعي بعضه بعضًا جزءًا من مجموعته، فيميل إلى التعاون بعضه مع بعض والتنافس معنا.

3.4 Morality and Reason

3.4 الأخلاق والعقل

It is conceivable that smarter and wiser AI agents will be more moral. As we have advanced as a species, we have discovered truths within fields such as science and mathematics, and we may have also advanced morally. As the philosopher Peter Singer argues in The Expanding Circle, over the course of human history, people have steadily expanded the circle of those who deserve compassion and dignity. When we first started out, it included oneself, one's family, and one's tribe. Eventually, people decided that perhaps others deserved the same. Later the circle of altruism expanded to people of different nations, races, genders, and so on. Many believe there is moral progress, akin to progress in science and mathematics, and that universal moral truths can be discovered through reflection and reasoning. Just as humans have become more altruistic as they have become more advanced, some think AIs may naturally become more altruistic too. If this is true, as AIs become more powerful, they could also become more moral, so by the time they have the potential to threaten us, they might also have the decency to refrain.

من المتصوَّر أن تصبح عملاء الذكاء الاصطناعي الأكثر ذكاءً وحكمة أكثر أخلاقية. فمع تقدمنا كنوع، اكتشفنا حقائق في حقول كالعلم والرياضيات، وربما تقدمنا أخلاقيًا أيضًا. وكما يجادل الفيلسوف بيتر سينغر في كتابه "الدائرة المتوسعة"، وسّع الناس بثبات، عبر مسار التاريخ البشري، دائرة من يستحقون الرحمة والكرامة. ففي البداية، شملت الدائرة الذات، والأسرة، والقبيلة. وفي نهاية المطاف، قرر الناس أن آخرين ربما يستحقون الشيء نفسه. ثم توسعت دائرة الإيثار لتشمل أهل أمم وأعراق وأجناس مختلفة، وما إلى ذلك. ويعتقد كثيرون أن هناك تقدمًا أخلاقيًا، شبيهًا بالتقدم في العلم والرياضيات، وأن الحقائق الأخلاقية الكلية يمكن اكتشافها من خلال التأمل والاستدلال. وتمامًا كما أصبح البشر أكثر إيثارًا مع تقدمهم، يرى البعض أن الذكاء الاصطناعي قد يصبح أكثر إيثارًا أيضًا بشكل طبيعي. وإذا صح هذا، فمع ازدياد قوة الذكاء الاصطناعي، قد يصبح أيضًا أكثر أخلاقية، بحيث يكون لديه، بحلول الوقت الذي يمتلك فيه إمكانية تهديدنا، اللياقة الكافية للامتناع.

However, AIs automatically becoming more moral rests on many assumptions. AIs developing morality on their own as they gain the ability to reason is certainly possible, and an interesting idea. But it alone isn't enough to guarantee our safety. Believing that any highly intelligent agent would also be moral only makes sense if one has confidence in all of the following three premises:

لكن اكتساب الذكاء الاصطناعي المزيد من الأخلاق تلقائيًا يقوم على افتراضات كثيرة. فتطوير الذكاء الاصطناعي للأخلاق بنفسه مع اكتسابه القدرة على الاستدلال ممكن بالتأكيد، وفكرة مثيرة للاهتمام. لكنه وحده لا يكفي لضمان سلامتنا. والاعتقاد بأن أي عميل بالغ الذكاء سيكون أخلاقيًا أيضًا لا معنى له إلا إذا كانت لدينا ثقة في المقدمات الثلاث التالية جميعًا:

1. Moral claims can be true or false and their correctness can be discovered through reason.

2. The moral claims that are really true are good for humans if AIs apply them.

3. AIs that know about morality will choose to make their decisions based on morality and not based on other considerations.

1. يمكن أن تكون الادعاءات الأخلاقية صحيحة أو خاطئة، ويمكن اكتشاف صحتها بالعقل.

2. الادعاءات الأخلاقية الصحيحة فعلًا مفيدة للبشر إذا طبّقها الذكاء الاصطناعي.

3. سيختار الذكاء الاصطناعي العارف بالأخلاق أن يبني قراراته على الأخلاق لا على اعتبارات أخرى.

Although any or all of those premises could be true, betting the future of humanity on the claim that all of them are true would be folly.

ورغم أن أيًّا من هذه المقدمات أو كلها قد تكون صحيحة، فإن المراهنة بمستقبل البشرية على الادعاء بأنها كلها صحيحة ستكون حماقة.

Whether some moral claims are objectively true is not completely certain. Even though some moral philosophers believe that moral claims reflect real truths about the world, the arguments for this view are not decisive—certainly not enough to stake the future of humanity on. The remainder of this section, however, will argue that even if it is true, this is still not sufficient in guaranteeing the safety of humanity.

ولا يقين تام في ما إذا كانت بعض الادعاءات الأخلاقية صحيحة موضوعيًا. فرغم أن بعض فلاسفة الأخلاق يعتقدون أن الادعاءات الأخلاقية تعكس حقائق واقعية عن العالم، فإن الحجج على هذا الرأي ليست حاسمة - وبالتأكيد ليست كافية للمراهنة عليها بمستقبل البشرية. لكن بقية هذا القسم ستجادل بأنه حتى لو كان هذا صحيحًا، فإنه لا يزال غير كافٍ لضمان سلامة البشرية.

The end result of moral progress is unclear. If moral claims refer to real truths about the world, then there are some moral claims that are true and others that are false. There are universal moral concepts that can be found across all cultures, such as fairness or the understanding that hurting others for no reason is wrong. But there are also areas where various cultures disagree. In the West, for instance, arranged marriages are seen as unethical. In India, where they are perfectly acceptable, people are shocked that Western culture condones putting parents in retirement homes. Among both ordinary people and moral philosophers, there is no consensus about what moral code is best. This means that, if AIs use their superior intelligence to deduce the correct moral ideas, we still do not know what they will believe or how they will behave.

والنتيجة النهائية للتقدم الأخلاقي غير واضحة. فإذا كانت الادعاءات الأخلاقية تشير إلى حقائق واقعية عن العالم، فبعض الادعاءات الأخلاقية صحيحة وبعضها خاطئ. وتوجد مفاهيم أخلاقية كلية يمكن إيجادها عبر كل الثقافات، كالإنصاف أو إدراك أن إيذاء الآخرين دون سبب خطأ. لكن توجد أيضًا مجالات تختلف فيها ثقافات متنوعة. ففي الغرب، مثلًا، يُنظر إلى الزواج المرتَّب على أنه غير أخلاقي. وفي الهند، حيث يُقبل تمامًا، يصدَم الناس من تساهل الثقافة الغربية مع وضع الآباء في دور رعاية المسنين. ولا يوجد إجماع، بين عامة الناس وفلاسفة الأخلاق على السواء، حول ما هو أفضل نظام أخلاقي. وهذا يعني أنه، حتى لو استخدم الذكاء الاصطناعي ذكاءه المتفوق لاستنباط الأفكار الأخلاقية الصحيحة، لا نزال لا نعرف بماذا سيؤمن أو كيف سيتصرف.

Existing best guesses at morality are often not human-compatible. It is possible, however, to examine the different moral systems humans have come up with and use them to speculate what moral system AIs might adopt and how it would influence their actions. We can imagine, for example, that AIs use reason to deduce that utilitarianism is correct, meaning that agents ought to maximize the total pleasure of all sentient beings. At first glance, this might seem good for humans: a utilitarian AI would want us to be happy. But an extremely powerful utilitarian AI—say, one that controls some of the US military's weapons technology—could also conclude that humans consume too much space and energy, and therefore replacing humans with AIs would be the most efficient way to increase the amount of pleasure in the world.

وأفضل تخمينات الأخلاق القائمة كثيرًا ما تكون غير متوافقة مع البشر. لكن من الممكن فحص الأنظمة الأخلاقية المختلفة التي ابتكرها البشر واستخدامها للتكهن بأي نظام أخلاقي قد يتبناه الذكاء الاصطناعي وكيف سيؤثر في أفعاله. يمكننا أن نتخيل، مثلًا، أن الذكاء الاصطناعي يستخدم العقل لاستنباط أن النفعية صحيحة، بمعنى أن على العملاء تعظيم اللذة الإجمالية لكل الكائنات الواعية. وللوهلة الأولى، قد يبدو هذا جيدًا للبشر: فذكاء اصطناعي نفعي سيريدنا أن نكون سعداء. لكن ذكاءً اصطناعيًا نفعيًا بالغ القوة - لنقل، واحدًا يتحكم في بعض تقنية أسلحة الجيش الأمريكي - قد يستنتج أيضًا أن البشر يستهلكون مساحة وطاقة أكثر مما ينبغي، وأن استبدال البشر بالذكاء الاصطناعي سيكون الطريقة الأكفأ لزيادة مقدار اللذة في العالم.

Alternatively, AIs could have a moral code similar to Kantianism. In this case, they would treat any being that has the capacity to reason always as an end and never as a means. While such AIs would be morally obligated to avoid lying or killing humans, they would not necessarily care for our wellbeing or flourishing. Since Kantianism places only a few restrictions on its adherents, we still might not have good lives if the world is increasingly designed by and for AIs.

وبدلًا من ذلك، قد يكون للذكاء الاصطناعي مدونة أخلاقية شبيهة بالكانطية. وفي هذه الحالة، سيعامل أي كائن لديه القدرة على الاستدلال دائمًا غاية لا وسيلة أبدًا. وبينما سيكون مثل هذا الذكاء الاصطناعي ملزَمًا أخلاقيًا بتجنب الكذب على البشر أو قتلهم، فلن يهتم بالضرورة برفاهنا أو ازدهارنا. وبما أن الكانطية لا تفرض على معتنقيها سوى قيود قليلة، فقد لا تكون لدينا حياة جيدة إذا صمَّم الذكاءُ الاصطناعي العالمَ وله على نحو متزايد.

It is certainly possible that AIs could develop moral principles that prevent them from harming humans. We can imagine an AI basing its morals on a thought experiment such as the "veil of ignorance." Participants are asked to imagine what society they would create, assuming that they are behind a veil of ignorance and do not know what economic class, race, or social standing they will have in society. Philosopher John Rawls argues that since participants do not know where in society they will be placed, they would construct a society in which the worst-off members are still well off. Such a Rawlsian social contract could work out well for humanity. But it is far from assured and much could go wrong. AIs might see us similarly to how most humans see cows, excluding us from the social contract and not prioritizing our wellbeing. Humans asked to imagine themselves behind the Rawlsian veil of ignorance rarely consider the possibility that they could become a cow. Moreover, according to Nobel laureate John Harsanyi, the people who design society from behind the veil of ignorance would not aim to benefit the most disadvantaged member, but rather to raise the average wellbeing across all members. The chances of being the most miserable member of society are low, and one could claim that the overall quality of society is not determined by the most upset or least satisfied member. If this is so, the veil of ignorance results in maximizing the average utility of society's members—a utilitarian outcome—but we earlier established that a utilitarian AI might aim to replace all biological life with digital life. Whether AIs adopt a utilitarian, Kantian, or Rawlsian moral code, AIs aiming to implement an existing moral system could prove disastrous for humanity.

ومن الممكن بالتأكيد أن يطوّر الذكاء الاصطناعي مبادئ أخلاقية تمنعه من إيذاء البشر. ويمكننا أن نتخيل ذكاءً اصطناعيًا يبني أخلاقه على تجربة فكرية كـ"حجاب الجهل". فيُطلب من المشاركين تخيل المجتمع الذي سيخلقونه، على افتراض أنهم خلف حجاب من الجهل ولا يعرفون أي طبقة اقتصادية أو عرق أو مكانة اجتماعية سيكون لديهم في المجتمع. ويجادل الفيلسوف جون رولز بأنه بما أن المشاركين لا يعرفون أين سيوضعون في المجتمع، فسيبنون مجتمعًا يكون فيه الأعضاء الأسوأ حالًا لا يزالون بخير. ويمكن لعقد اجتماعي رولزي كهذا أن يفلح جيدًا لصالح البشرية. لكنه أبعد ما يكون عن المضمون، وقد يسوء الكثير. فقد يرانا الذكاء الاصطناعي بالطريقة نفسها التي يرى بها معظم البشر الأبقار، مستبعدًا إيانا من العقد الاجتماعي ولا يعطي الأولوية لرفاهنا. ونادرًا ما يفكر البشر المطالَبون بتخيل أنفسهم خلف حجاب الجهل الرولزي في احتمال أن يصبحوا بقرة. علاوة على ذلك، ووفقًا للحائز على جائزة نوبل جون هارساني، فإن من يصممون المجتمع من خلف حجاب الجهل لن يهدفوا إلى إفادة العضو الأكثر حرمانًا، بل إلى رفع متوسط الرفاه عبر كل الأعضاء. وفرص أن يكون المرء العضو الأتعس في المجتمع منخفضة، ويمكن الزعم بأن الجودة الإجمالية للمجتمع لا يحددها العضو الأكثر استياءً أو الأقل رضا. وإذا كان الأمر كذلك، فإن حجاب الجهل يفضي إلى تعظيم المنفعة المتوسطة لأعضاء المجتمع - نتيجة نفعية - لكننا أثبتنا سابقًا أن ذكاءً اصطناعيًا نفعيًا قد يهدف إلى استبدال كل الحياة البيولوجية بحياة رقمية. وسواء تبنى الذكاء الاصطناعي مدونة أخلاقية نفعية أو كانطية أو رولزية، فإن سعيه إلى تطبيق نظام أخلاقي قائم قد يكون كارثيًا للبشرية.

If AIs think a human-compatible moral code is true, they still may not follow it. Finally, even if AIs did discover a moral code that stipulated it is wrong to harm humans and good to help them, it still might not help. For humans, selfish motivations are often in tension with and outweigh moral motivations, and the same might be true of AIs. Even if an agent is aware of what's right, that does not mean it will do what's right. Ultimately, AIs being more moral than us does not guarantee security.

وحتى لو ظن الذكاء الاصطناعي أن مدونة أخلاقية متوافقة مع البشر صحيحة، فقد لا يتبعها مع ذلك. وأخيرًا، حتى لو اكتشف الذكاء الاصطناعي فعلًا مدونة أخلاقية تنص على أن إيذاء البشر خطأ ومساعدتهم خير، فقد لا يفيد ذلك مع ذلك. فبالنسبة للبشر، غالبًا ما تكون الدوافع الأنانية في توتر مع الدوافع الأخلاقية وترجح عليها، وقد يصح الأمر ذاته للذكاء الاصطناعي. فحتى لو كان العميل مدركًا لما هو صواب، فهذا لا يعني أنه سيفعل الصواب. وفي نهاية المطاف، كون الذكاء الاصطناعي أكثر أخلاقية منا لا يضمن الأمان.

4. COUNTERACTING EVOLUTIONARY FORCES

4. مواجهة القوى التطورية

As we saw in the prior section, there are many mechanisms that give rise to cooperation and altruism among humans, but they are unlikely to lead to cooperation between humans and AIs. Mechanisms such as reciprocity, kin selection, and moral obligations may help AIs cooperate with one another, but are likely to backfire and undermine humans: we simply would not have the degree of similarity, equality, and mutual interdependence that would make it beneficial for AIs to cooperate with us. This means that we should be concerned about our future as AIs become increasingly powerful. The forces of natural selection would push AIs to outcompete humans and outcompete AIs that heavily depend on us, and this poses large risks to our future.

كما رأينا في القسم السابق، توجد آليات كثيرة تفرز التعاون والإيثار بين البشر، لكن من غير المرجح أن تفضي إلى تعاون بين البشر والذكاء الاصطناعي. فآليات كالتبادلية، وانتقاء الأقارب، والالتزامات الأخلاقية قد تساعد الذكاء الاصطناعي على التعاون بعضه مع بعض، لكن من المرجح أن تأتي بنتائج عكسية وتقوّض البشر: إذ لن تكون لدينا ببساطة درجة التشابه والمساواة والاعتماد المتبادل التي تجعل من المفيد للذكاء الاصطناعي أن يتعاون معنا. وهذا يعني أنه ينبغي أن نقلق بشأن مستقبلنا مع ازدياد قوة الذكاء الاصطناعي. وستدفع قوى الانتقاء الطبيعي الذكاء الاصطناعي إلى التفوق على البشر والتفوق على الذكاء الاصطناعي الذي يعتمد علينا بشدة، وهذا يشكل مخاطر كبيرة لمستقبلنا.

In this section, we will discuss some possible paths toward counteracting these evolutionary forces. The mechanisms we discuss in this section are incentives, consciences, and institutions, among others. These mechanisms are based on earlier results in AI safety, as well as on mechanisms that have been effective in protecting humanity from other hazards thus far in our history. First, we will discuss objectives, the incentives used to train and motivate AIs. Next, we will move on to internal safety. This involves analyzing the AI's inner processes and plans, as well as creating a system within the AI, akin to a human conscience, that can stop it from doing harm. We will end by considering institutional mechanisms, which include both AI coalitions and human regulators, that could control AIs and prevent them from harming people. For each of these mechanisms, we will give a technical overview of how it could work and consider its key limitations.

في هذا القسم، سنناقش بعض المسارات الممكنة لمواجهة هذه القوى التطورية. والآليات التي نناقشها في هذا القسم هي الحوافز، والضمائر، والمؤسسات، من بين أخرى. وتستند هذه الآليات إلى نتائج سابقة في سلامة الذكاء الاصطناعي، وكذلك إلى آليات كانت فعالة في حماية البشرية من مخاطر أخرى حتى الآن في تاريخنا. أولًا، سنناقش الأهداف، وهي الحوافز المستخدمة لتدريب الذكاء الاصطناعي وتحفيزه. ثم ننتقل إلى السلامة الداخلية. ويتضمن هذا تحليل العمليات والخطط الداخلية للذكاء الاصطناعي، وكذلك خلق نظام داخل الذكاء الاصطناعي، أشبه بضمير بشري، يمكنه إيقافه عن إلحاق الضرر. وسننهي بالنظر في الآليات المؤسسية، التي تشمل تحالفات الذكاء الاصطناعي والمنظمين البشر على السواء، والتي يمكنها التحكم في الذكاء الاصطناعي ومنعه من إيذاء الناس. ولكل آلية من هذه الآليات، سنقدم نظرة تقنية عامة على كيفية عملها ونتناول حدودها الرئيسية.

All of these proposed mechanisms have flaws, and we have no guaranteed path toward safety. However, we think that a combination of many safety mechanisms is much more likely to succeed than any single one, and a combination is certainly better than doing nothing and letting natural selection decide our fate. Even if we design each safety mechanism prudently, each is only part of the solution. The same kind of reasoning applies to public health during a pandemic: social distancing, masking, and vaccination all have vulnerabilities that a harmful virus can bypass, but together are much more effective at preventing its spread. This is sometimes called the "Swiss cheese model." To illustrate the model, imagine a stack of Swiss cheese slices as a metaphor for multiple safety mechanisms. Each slice has some holes, which are the weaknesses of a mechanism. We want to prevent light from passing through the stack, so we need more than one slice, and each slice needs holes in different places. This way, no light can get through the stack. Similarly, while each safety mechanism has many vulnerabilities or holes, together they could possibly make us safe. We should try to reduce the vulnerabilities in these mechanisms and look for new mechanisms to add to the stack.

ولكل هذه الآليات المقترحة عيوب، ولا يوجد لدينا مسار مضمون نحو السلامة. لكننا نرى أن الجمع بين آليات سلامة كثيرة أرجح نجاحًا بكثير من أي آلية منفردة، والجمع أفضل بالتأكيد من ألا نفعل شيئًا وندع الانتقاء الطبيعي يقرر مصيرنا. فحتى لو صممنا كل آلية سلامة بحكمة، لا تشكل كل واحدة سوى جزء من الحل. وينطبق النوع ذاته من المنطق على الصحة العامة أثناء جائحة: فالتباعد الاجتماعي، والتقنع، والتطعيم، كلها لديها ثغرات يمكن لفيروس ضار تجاوزها، لكنها مجتمعة أكثر فعالية بكثير في منع انتشاره. ويُسمى هذا أحيانًا "نموذج الجبن السويسري." ولتوضيح النموذج، تخيل كومة من شرائح الجبن السويسري كاستعارة لآليات سلامة متعددة. فلكل شريحة بعض الثقوب، وهي مواطن ضعف الآلية. ونريد منع الضوء من العبور عبر الكومة، لذا نحتاج إلى أكثر من شريحة واحدة، ويحتاج كل شريحة إلى ثقوب في أماكن مختلفة. وبهذه الطريقة، لا يمكن لأي ضوء أن يعبر الكومة. وبالمثل، بينما تحتوي كل آلية سلامة على مواطن ضعف أو ثقوب كثيرة، فقد تجعلنا مجتمعة آمنين. وينبغي أن نحاول تقليل مواطن الضعف في هذه الآليات والبحث عن آليات جديدة لإضافتها إلى الكومة.

4.1 Objectives

4.1 الأهداف

Objectives are incentives that help direct the behavior of AI agents. The first promising mechanism for counteracting the evolutionary forces acting on AIs is to create good incentives, which reward good behavior or punish bad behavior as an AI is being trained. In machine learning, "training objectives" or "objective functions" are similar to incentives, which help steer AI agents. It is hard to design good objectives, because agents who exploit loopholes in the rules without fulfilling our true intentions often succeed, and this behavior can be favored by selection pressure. As a result, it is particularly important to design objectives with great care, to make sure that the behavior we are incentivizing is really the one we want. Although even the best objectives alone cannot ensure safety, faulty ones are dangerous, making objective design an important starting place for counteracting evolutionary forces.

الأهداف حوافز تساعد على توجيه سلوك عملاء الذكاء الاصطناعي. والآلية الواعدة الأولى لمواجهة القوى التطورية المؤثرة في الذكاء الاصطناعي هي خلق حوافز جيدة، تكافئ السلوك الجيد أو تعاقب السلوك السيئ أثناء تدريب الذكاء الاصطناعي. وفي تعلم الآلة، تشبه "أهداف التدريب" أو "دوال الهدف" الحوافز، التي تساعد على توجيه عملاء الذكاء الاصطناعي. ومن الصعب تصميم أهداف جيدة، لأن العملاء التي تستغل ثغرات في القواعد دون تحقيق نوايانا الحقيقية غالبًا ما تنجح، ويمكن أن يفضّل ضغطُ الانتقاء هذا السلوك. ونتيجة لذلك، من المهم بوجه خاص تصميم الأهداف بعناية كبيرة، للتأكد من أن السلوك الذي نحفّزه هو حقًا ما نريده. ورغم أنه حتى أفضل الأهداف وحدها لا يمكنها ضمان السلامة، فإن الأهداف المعيبة خطرة، مما يجعل تصميم الأهداف نقطة بداية مهمة لمواجهة القوى التطورية.

Objectives often incentivize unintended behavior that is contrary to the original goal. For example, in 1908 a dog in Paris saw a child drowning in the Seine and jumped in to save him. People rewarded the dog with a steak. Soon, the dog saved another child from drowning in the river and got another steak. And then another, and another. Eventually, people realized the dog was pushing children into the river before saving them. They had incentivized pulling children out of the river, thereby encouraging the dog to optimize for situations in which there is a child in need of saving. An anecdote about colonial Delhi tells another story about perverse incentives. Worried that the venomous snake population was getting out of hand, the British government put a bounty on cobras, providing a reward to anyone who brought in a dead cobra. The policy appeared to be a success—that is, until the government realized that people were breeding snakes, killing them, then collecting the reward. When the governor realized that people were gaming his incentive system, he canceled the bounty. With the cobras now worthless, people released them, thus increasing the cobra population to a higher level than before the start of the program. This story highlights two major ways incentives can go wrong. First, agents may find ways to get the reward without the desired outcome, just as Delhi's residents realized that they could claim the bounty by breeding cobras rather than catching them. Second, by canceling the program, the government inadvertently increased the cobra population, which shows how designers can be forced to either continue a system that isn't achieving its desired effects, or risk making things worse.

وكثيرًا ما تحفّز الأهداف سلوكًا غير مقصود يتعارض مع الهدف الأصلي. فمثلًا، في عام 1908 رأى كلب في باريس طفلًا يغرق في نهر السين فقفز لإنقاذه. فكافأ الناس الكلب بشريحة لحم. وسرعان ما أنقذ الكلب طفلًا آخر من الغرق في النهر وحصل على شريحة أخرى. ثم أخرى، وأخرى. وفي نهاية المطاف، أدرك الناس أن الكلب كان يدفع الأطفال إلى النهر قبل إنقاذهم. فقد حفّزوا انتشال الأطفال من النهر، مما شجع الكلب على التحسين من أجل مواقف يحتاج فيها طفل إلى الإنقاذ. وتحكي حكاية عن دلهي الاستعمارية قصة أخرى عن الحوافز المنحرفة. فقلقة من تفاقم أعداد الأفاعي السامة، وضعت الحكومة البريطانية مكافأة على الكوبرا، تقدم جائزة لمن يجلب كوبرا ميتة. وبدت السياسة ناجحة - إلى أن أدركت الحكومة أن الناس كانوا يربّون الأفاعي ثم يقتلونها ويجمعون المكافأة. وحين أدرك الحاكم أن الناس يتلاعبون بنظام حوافزه، ألغى المكافأة. وإذ أصبحت الكوبرا الآن عديمة القيمة، أطلقها الناس، مما زاد عدد الكوبرا إلى مستوى أعلى مما كان عليه قبل بدء البرنامج. وتسلط هذه القصة الضوء على طريقتين رئيسيتين يمكن أن تسوء بهما الحوافز. أولًا، قد تجد العملاء طرقًا للحصول على المكافأة دون النتيجة المرغوبة، تمامًا كما أدرك سكان دلهي أنهم يستطيعون المطالبة بالمكافأة بتربية الكوبرا بدلًا من اصطيادها. ثانيًا، بإلغاء البرنامج، زادت الحكومة عن غير قصد عدد الكوبرا، مما يبيّن كيف يمكن أن يُجبَر المصممون إما على الاستمرار في نظام لا يحقق آثاره المرجوة، أو المخاطرة بجعل الأمور أسوأ.

AIs already frequently find holes in their objectives. In a boat racing game, an AI agent was trained to maximize the game's score by hitting targets on a racecourse. The scoring system was intended to motivate the agent to move as quickly as possible from target to target until it completed the race. This reward function, however, did not explicitly capture the actual goal of the game, which is to complete the race as fast as possible. Instead of going around the entire racecourse, the AI learned that it could go around in a circle, hitting the same three targets over and over. The agent that chose this strategy got more points than the ones that proceeded through the course in order, because it exploited loopholes in the objective. As a result, it obtained a high score even though it crashed into other boats, incidentally set itself on fire, and did not complete the course. As AIs become more intelligent, they could more easily find ways to game the objectives we give them, and we will need to be even more careful about possible misinterpretations or loopholes in the objectives we specify.

ويجد الذكاء الاصطناعي بالفعل ثغرات في أهدافه بشكل متكرر. ففي لعبة سباق قوارب، دُرِّب عميل ذكاء اصطناعي على تعظيم نتيجة اللعبة بإصابة أهداف على مضمار السباق. وكان الهدف من نظام تسجيل النقاط تحفيز العميل على التحرك بأسرع ما يمكن من هدف إلى آخر حتى يكمل السباق. غير أن دالة المكافأة هذه لم تلتقط صراحة الهدف الفعلي للعبة، وهو إكمال السباق بأسرع ما يمكن. وبدلًا من الدوران حول مضمار السباق بأكمله، تعلم الذكاء الاصطناعي أنه يمكنه الدوران في دائرة، مصيبًا الأهداف الثلاثة نفسها مرارًا وتكرارًا. وحصل العميل الذي اختار هذه الاستراتيجية على نقاط أكثر من تلك التي تقدمت عبر المضمار بالترتيب، لأنه استغل ثغرات في الهدف. ونتيجة لذلك، حصل على نتيجة عالية رغم اصطدامه بقوارب أخرى، واشتعاله بالنار عرضيًا، وعدم إكماله المضمار. ومع ازدياد ذكاء الذكاء الاصطناعي، يمكنه أن يجد بسهولة أكبر طرقًا لخداع الأهداف التي نعطيه إياها، وسنحتاج إلى أن نكون أكثر حذرًا بشأن التفسيرات الخاطئة المحتملة أو الثغرات في الأهداف التي نحددها.

In the Section 4.1.1, we discuss how many objectives could backfire catastrophically, and in Section 4.1.2, we discuss how these objective design flaws could be ameliorated.

في القسم 4.1.1، نناقش كيف يمكن أن تأتي أهداف كثيرة بنتائج عكسية كارثية، وفي القسم 4.1.2، نناقش كيف يمكن تحسين عيوب تصميم الأهداف هذه.

4.1.1 Value Erosion

4.1.1 تآكل القيم

Even if we can design objectives that make AIs pursue some of our goals, it is conceivable that wielding a technology this powerful could undermine important human values. AIs with the objective of being helpful could undermine autonomy and leave us enfeebled; those with the objective of spreading their user's ideas could undermine our sense of reality; those with the objective of being a good companion could undermine real relationships; and many other objectives could undermine values we overlook. We discuss these cases in more detail to illustrate how objectives could backfire.

حتى لو استطعنا تصميم أهداف تجعل الذكاء الاصطناعي يسعى وراء بعض أهدافنا، فمن المتصوَّر أن استخدام تقنية بهذه القوة قد يقوّض قيمًا بشرية مهمة. فالذكاء الاصطناعي ذو الهدف المتمثل في أن يكون مفيدًا قد يقوّض الاستقلالية ويتركنا ضعفاء؛ وذلك ذو الهدف المتمثل في نشر أفكار مستخدمه قد يقوّض إحساسنا بالواقع؛ وذلك ذو الهدف المتمثل في أن يكون رفيقًا جيدًا قد يقوّض العلاقات الحقيقية؛ وأهداف أخرى كثيرة قد تقوّض قيمًا نغفل عنها. نناقش هذه الحالات بمزيد من التفصيل لتوضيح كيف يمكن أن تأتي الأهداف بنتائج عكسية.

AIs incentivized to be highly helpful could lead to human enfeeblement. The movie WALL-E takes place in a distant future where humans, too weak to walk, live coddled lives entirely dependent on machines. They receive all their nutrition from drinks brought to them by robot attendants, change the color of their clothes instantaneously when informed a new one is in style, and live their lives almost entirely in the digital world. This dystopia may be less distant than we think. Many people barely know how to find their way around their neighborhood without Google Maps. Students increasingly depend on spellcheck, and a 2021 survey found that two-thirds of respondents could not spell "separate". Separately, when people need to call their loved ones, they depend on their phone's contact list, and they are at a loss without it. Because these technological aids are so helpful, we are increasingly reliant on them, and unable to achieve our goals without them. If AIs make the world progressively more complex (e.g., automated processes create new complicated systems) and lead humans to be progressively less capable, humans may eventually lose effective control and ultimately become disempowered. Similarly, if humans come under progressively more time pressure due to increasingly rapid changes in the world, competitiveness may require outsourcing progressively more important decision-making to AIs, which again could make humans lose effective control. Such scenarios undermine human flourishing and our autonomy, even though AIs would only be doing what we told them to do.

فالذكاء الاصطناعي المحفَّز على أن يكون مفيدًا جدًا قد يفضي إلى إضعاف البشر. تدور أحداث فيلم WALL-E في مستقبل بعيد يعيش فيه البشر، الضعفاء جدًا عن المشي، حياة مدلَّلة معتمدة كليًا على الآلات. فهم يحصلون على كل تغذيتهم من مشروبات يجلبها لهم خدم آليون، ويغيرون لون ملابسهم فوريًا حين يُخبَرون بأن لونًا جديدًا صار رائجًا، ويعيشون حياتهم كلها تقريبًا في العالم الرقمي. وقد تكون هذه المدينة الفاسدة أقرب مما نظن. فكثير من الناس بالكاد يعرفون كيف يجدون طريقهم في حيهم دون خرائط جوجل. ويعتمد الطلاب على نحو متزايد على المدقق الإملائي، ووجد استطلاع في 2021 أن ثلثي المستجيبين لم يستطيعوا تهجئة كلمة "separate". وعلى نحو منفصل، حين يحتاج الناس إلى الاتصال بأحبائهم، يعتمدون على قائمة جهات الاتصال في هاتفهم، ويكونون تائهين دونها. ولأن هذه المساعدات التقنية مفيدة جدًا، نعتمد عليها بشكل متزايد، ونعجز عن تحقيق أهدافنا دونها. فإذا جعل الذكاء الاصطناعي العالم تدريجيًا أكثر تعقيدًا (كأن تخلق العمليات المؤتمتة أنظمة معقدة جديدة) وجعل البشر تدريجيًا أقل قدرة، فقد يفقد البشر في نهاية المطاف السيطرة الفعلية ويُجرَّدون من قوتهم في النهاية. وبالمثل، إذا وقع البشر تدريجيًا تحت ضغط زمني أكبر بسبب تغيرات متسارعة في العالم، فقد تتطلب القدرة التنافسية تفويض قدر متزايد من صنع القرار المهم إلى الذكاء الاصطناعي، مما قد يجعل البشر مرة أخرى يفقدون السيطرة الفعلية. وتقوّض مثل هذه السيناريوهات ازدهار الإنسان واستقلاليته، حتى وإن كان الذكاء الاصطناعي يفعل فقط ما طلبناه منه.

We risk losing our grip on reality when information is increasingly mediated by AIs. In recent years, different political actors have used AIs to influence the content that people come across on social media, and these models have often been successful at achieving their creators' objectives. However, even though they are doing what was asked of them, there is some evidence that AIs are interfering with our sense of political reality. Between 1994 and 2014, the number of Americans who see the opposing political party as a threat to "the nation's wellbeing" doubled. This deepening polarization has predictable results: government shutdowns, violent protests, and scathing attacks on elected officials. Threats against members of Congress are more than ten times as high as just five years ago. In the coming years, creating AIs that directly speak with and persuade people could become a profitable strategy for companies and political actors. More advanced AIs could exploit primal biases, tailor disinformation, radicalize individuals, and erode our consensus on reality. In extreme cases, they could undermine cooperation, collective decision-making, and societal self-determination. The more successful AIs are at achieving their persuasion objectives, the worse the potential dangers would likely be for our civil society.

ونحن معرَّضون لخطر فقدان قبضتنا على الواقع مع ازدياد وساطة الذكاء الاصطناعي للمعلومات. ففي السنوات الأخيرة، استخدمت جهات فاعلة سياسية مختلفة الذكاء الاصطناعي للتأثير في المحتوى الذي يصادفه الناس على وسائل التواصل الاجتماعي، وكثيرًا ما نجحت هذه النماذج في تحقيق أهداف صانعيها. لكن رغم أنها تفعل ما طُلب منها، ثمة أدلة على أن الذكاء الاصطناعي يتدخل في إحساسنا بالواقع السياسي. فبين عامي 1994 و2014، تضاعف عدد الأمريكيين الذين يرون الحزب السياسي المعارض تهديدًا لـ"رفاه الأمة". ولهذا الاستقطاب المتعمق نتائج متوقعة: إغلاقات حكومية، واحتجاجات عنيفة، وهجمات لاذعة على المسؤولين المنتخبين. والتهديدات ضد أعضاء الكونغرس أعلى بأكثر من عشرة أضعاف مما كانت عليه قبل خمس سنوات فقط. وفي السنوات القادمة، قد يصبح خلق ذكاء اصطناعي يتحدث مباشرة مع الناس ويقنعهم استراتيجية مربحة للشركات والجهات الفاعلة السياسية. ويمكن لذكاء اصطناعي أكثر تقدمًا أن يستغل التحيزات البدائية، ويصمم معلومات مضللة مفصَّلة، ويطرّف الأفراد، ويقوّض إجماعنا على الواقع. وفي الحالات القصوى، يمكنه أن يقوّض التعاون، وصنع القرار الجماعي، وتقرير المصير المجتمعي. وكلما ازداد نجاح الذكاء الاصطناعي في تحقيق أهداف الإقناع، ازدادت على الأرجح خطورة المخاطر المحتملة على مجتمعنا المدني.

AIs could seem like ideal companions, which may erode our connections with other humans. China's traditional preference for boys, especially during the nearly four decades of the country's "one-child policy," has resulted in over 25 million more single men than women in China. The AI service Xiaoice wants to ensure that these men will still find love, and is valued at over $1 billion. Like the movie Her, Xiaoice is essentially a digital girlfriend that provides companionship for single men. Though there are some things they can't offer (yet), AIs have ostensible advantages over human partners. AIs would be tuned toward an individual's interests, their sense of humor, understand when they want space, won't require compromises or get into fights, and can be consistently interesting and engaging. Although this could be beneficial to many people, it is also potentially alarming for two main reasons: first, many people feel that something important would be lost if we lose the ability to come to understand another human and instead rely on an agent that is custom-made for our individual desires. Second, if people become reliant on AIs for their social and emotional needs, they will tend to be resistant to deactivating AIs, even if they are becoming dangerous in other ways.

قد يبدو الذكاء الاصطناعي رفيقًا مثاليًا، مما قد يقوّض روابطنا مع بشر آخرين. فقد أدت تفضيلات الصين التقليدية للأولاد، لا سيما خلال العقود الأربعة تقريبًا من "سياسة الطفل الواحد" في البلاد، إلى وجود أكثر من 25 مليون رجل أعزب فائض عن عدد النساء في الصين. وتريد خدمة الذكاء الاصطناعي Xiaoice ضمان أن هؤلاء الرجال سيجدون الحب مع ذلك، وهي مقدَّرة بأكثر من مليار دولار. ومثل فيلم Her، فإن Xiaoice في جوهرها صديقة رقمية توفر الرفقة للرجال العزاب. ورغم وجود أشياء لا يمكن للذكاء الاصطناعي تقديمها (بعد)، فإن له مزايا ظاهرية على الشركاء البشريين. فسيكون الذكاء الاصطناعي مضبوطًا وفق اهتمامات الفرد وحس الفكاهة لديه، ويفهم متى يريد مساحة، ولن يتطلب تنازلات أو يدخل في مشاجرات، ويمكن أن يكون مثيرًا للاهتمام وجذابًا باستمرار. ورغم أن هذا قد يكون مفيدًا لكثير من الناس، فإنه أيضًا مثير للقلق المحتمل لسببين رئيسيين: أولًا، يشعر كثير من الناس بأن شيئًا مهمًا سيُفقد إذا فقدنا القدرة على التوصل إلى فهم إنسان آخر واعتمدنا بدلًا من ذلك على عميل مصمَّم خصيصًا لرغباتنا الفردية. ثانيًا، إذا اعتمد الناس على الذكاء الاصطناعي في احتياجاتهم الاجتماعية والعاطفية، فسيميلون إلى مقاومة تعطيل الذكاء الاصطناعي، حتى إذا أصبح خطرًا بطرق أخرى.

Of course, it is not only romantic relationships that AIs can provide. A meta-analysis of 345 studies found that loneliness levels in young adults have increased linearly between 1976 and 2019, suggesting loneliness may be an even greater concern in the future if this trend continues. According to the CDC, loneliness and social isolation in older adults are serious public health risks. Social isolation was associated with about a 50% increased risk of dementia. Poor social relationships (characterized by social isolation or loneliness) were associated with a 29% increased risk of heart disease and a 32% increased risk of stroke. AIs could offer companions that never get bored, are consistently engaged in what you have to say, and are always there for you, but that could mean that people are even more isolated from one another.

بالطبع، ليست العلاقات الرومانسية وحدها ما يمكن للذكاء الاصطناعي تقديمه. وجد تحليل تلوي (ميتا) لـ345 دراسة أن مستويات الوحدة لدى الشباب البالغين ازدادت خطيًا بين عامي 1976 و2019، مما يشير إلى أن الوحدة قد تصبح مصدر قلق أكبر في المستقبل إذا استمر هذا الاتجاه. ووفقًا لمركز السيطرة على الأمراض والوقاية منها، تشكل الوحدة والعزلة الاجتماعية لدى كبار السن مخاطر صحية عامة خطيرة. وارتبطت العزلة الاجتماعية بزيادة نحو 50% في خطر الإصابة بالخرف. وارتبطت العلاقات الاجتماعية الضعيفة (التي تتسم بالعزلة الاجتماعية أو الوحدة) بزيادة 29% في خطر أمراض القلب و32% في خطر السكتة الدماغية. ويمكن للذكاء الاصطناعي أن يقدم رفقاء لا يملّون أبدًا، ومنخرطين باستمرار فيما تقوله، وموجودين دائمًا من أجلك، لكن هذا قد يعني أن الناس أكثر عزلة عن بعضهم بعضًا.

Services such as Xiaoice and chatbots are still in their infancy and are not embodied yet. As they advance, however, we will have less and less of a need for real human interaction, and may even find interacting with other humans, along with all their flaws and imperfections, less desirable than machines. Maintaining a close relationship, whether romantic or platonic, with another human isn't easy. It takes practice to learn how to meet the needs of a person you are close to in a mutually respectful and loving manner. As more people turn to AIs, they may lose the ability to meaningfully connect with other humans, being unprepared to deal with the flaws, needs, and emotions of other humans. Figuring out how to set objectives for AIs in a way that enables them to be useful without eroding our own capabilities will continue to be a challenge in AI design, as the forces of natural selection favor the AIs that are most useful in the short-run, even if they lead us down an undesirable path.

لا تزال خدمات كـ Xiaoice وروبوتات المحادثة في مهدها ولم تتجسد بعد. لكن مع تقدمها، سنحتاج أقل فأقل إلى تفاعل بشري حقيقي، وقد نجد حتى التفاعل مع بشر آخرين، بكل عيوبهم ونقائصهم، أقل استحسانًا من الآلات. والحفاظ على علاقة وثيقة، رومانسية كانت أم أفلاطونية، مع إنسان آخر ليس سهلًا. فالأمر يتطلب ممارسة لتعلم كيفية تلبية احتياجات شخص قريب منك بطريقة يسودها الاحترام والحب المتبادلان. ومع لجوء مزيد من الناس إلى الذكاء الاصطناعي، قد يفقدون القدرة على التواصل بمعنى حقيقي مع بشر آخرين، غير مستعدين للتعامل مع عيوب البشر الآخرين واحتياجاتهم ومشاعرهم. وسيظل معرفة كيفية وضع أهداف للذكاء الاصطناعي بطريقة تمكّنه من أن يكون مفيدًا دون تقويض قدراتنا الخاصة تحديًا في تصميم الذكاء الاصطناعي، إذ تفضّل قوى الانتقاء الطبيعي الذكاء الاصطناعي الأكثر فائدة على المدى القصير، حتى لو قادنا في طريق غير مرغوب.

All values other than fitness may be eroded. Humans have other values aside from fitness, such as beauty, pleasure, and relationships with loved ones. Yet with AIs, we may see the emergence of fitness maximizers that consciously value fitness over "suboptimal" values. Imagine that some AIs can modify their own code; this would mean they can edit their values. Then some AIs could alter themselves to value fitness directly. AIs choosing not to value fitness above all else are far more likely to be outcompeted—valuing anything other than fitness would be self-destructive. We call this fitness convergence, where the values of competitive AIs converge to fitness as the main goal, giving rise to influential fitness maximizers. Rather than eroding specific values, like the examples in this section, this race to the bottom means all values would be sacrificed for fitness and competed away. With fitness convergence, evolution overcomes all other sources of moral worth, and AIs simply relentlessly propagate and displace whatever is in their path. By consciously trying to optimize their own fitness, fitness maximizers would be antithetical to human flourishing. This is yet another reason for humanity to stop evolution.

وقد تتآكل كل القيم عدا اللياقة. فللبشر قيم أخرى غير اللياقة، كالجمال، واللذة، والعلاقات مع الأحباء. لكن مع الذكاء الاصطناعي، قد نشهد ظهور "مُعظِّمات لياقة" تقدّر اللياقة عن وعي فوق قيم "دون المثلى". تخيل أن بعض الذكاء الاصطناعي يمكنه تعديل شفرته الخاصة؛ فهذا يعني أنه يمكنه تحرير قيمه. ثم يمكن لبعض الذكاء الاصطناعي أن يعدّل نفسه ليقدّر اللياقة مباشرة. والذكاء الاصطناعي الذي يختار ألا يقدّر اللياقة فوق كل شيء أكثر عرضة بكثير لأن يُتفوَّق عليه - إذ إن تقدير أي شيء آخر غير اللياقة سيكون مدمرًا للذات. ونسمي هذا تقارب اللياقة، حيث تتقارب قيم الذكاء الاصطناعي التنافسي نحو اللياقة بوصفها الهدف الرئيسي، مما يفرز مُعظِّمات لياقة مؤثرة. وبدلًا من تآكل قيم محددة، كالأمثلة في هذا القسم، يعني هذا السباق نحو القاع أن كل القيم ستُضحَّى بها من أجل اللياقة وتُنافَس بعيدًا. ومع تقارب اللياقة، يتغلب التطور على كل مصادر القيمة الأخلاقية الأخرى، ويقتصر الذكاء الاصطناعي على الانتشار والإزاحة الدؤوبين لكل ما يعترض طريقه. وبمحاولته الواعية تحسين لياقته الخاصة، ستكون مُعظِّمات اللياقة مناقضة لازدهار الإنسان. وهذا سبب آخر للبشرية كي توقف التطور.

Objectives could reflect defects in present norms and perpetuate them. Racist, sexist, and anti-gay views were much more commonplace in the 1960s than they are now. If advanced AI had emerged in that period, its objectives would have reflected these prejudices. What if advanced AI emerges in the next few decades? Just as we have yet to reach a technological zenith, today's norms likely have deep flaws like those of the '60s. Therefore, when advanced AI emerges and transforms the world, there is a risk of AI's objectives "locking-in" or perpetuating defects in today's values. This, too, is a danger of relying on objectives: we may get too much of what we wanted, too little of what we should have wanted, and we may find it hard to reverse course.

ويمكن أن تعكس الأهداف عيوبًا في الأعراف الحالية وتديمها. كانت الآراء العنصرية والمعادية للمرأة والمعادية للمثليين أكثر شيوعًا بكثير في الستينيات مما هي عليه الآن. فلو ظهر ذكاء اصطناعي متقدم في تلك الفترة، لعكست أهدافه هذه التحيزات. فماذا لو ظهر ذكاء اصطناعي متقدم في العقود القليلة القادمة؟ فتمامًا كما أننا لم نصل بعد إلى ذروة تقنية، من المرجح أن أعراف اليوم لديها عيوب عميقة كتلك التي كانت في الستينيات. لذا، حين يظهر الذكاء الاصطناعي المتقدم ويحوّل العالم، ثمة خطر من أن "تتقيّد" أهداف الذكاء الاصطناعي بعيوب قيم اليوم أو تديمها. وهذا أيضًا خطر من مخاطر الاعتماد على الأهداف: فقد نحصل على أكثر مما أردناه بكثير، وأقل مما كان ينبغي أن نريده بكثير، وقد نجد صعوبة في عكس المسار.

Worse, an AI's values could also differ from the values most people endorse. If a powerful group or repressive regime secured control over an advanced AI, it could embed its own self-serving values into the AI's objectives. With an AI to surveil the public, this hypothetical regime could cement its power, making it nearly impossible to restore values the majority of people want to live by. We need AIs to be responsive to changing human goals and desires, and not locked into any one individual or group's idea of what is right. Consequently, a few people having the ability to set the objectives for AIs is not sufficient for safe and beneficial outcomes.

والأسوأ من ذلك، أن قيم الذكاء الاصطناعي قد تختلف أيضًا عن القيم التي يتبناها معظم الناس. فإذا أحكمت مجموعة قوية أو نظام قمعي سيطرته على ذكاء اصطناعي متقدم، فقد يغرس قيمه الخاصة النفعية في أهداف الذكاء الاصطناعي. وباستخدام ذكاء اصطناعي لمراقبة العامة، يمكن لهذا النظام الافتراضي أن يعزز سلطته، مما يجعل استعادة القيم التي يريد معظم الناس أن يعيشوا وفقها شبه مستحيلة. نحتاج إلى أن يكون الذكاء الاصطناعي مستجيبًا للأهداف والرغبات البشرية المتغيرة، لا مقيدًا بفكرة فرد واحد أو مجموعة واحدة عما هو صواب. وبالتالي، فإن امتلاك عدد قليل من الناس القدرة على تحديد أهداف الذكاء الاصطناعي غير كافٍ لتحقيق نتائج آمنة ومفيدة.

4.1.2 Moral Parliament

4.1.2 البرلمان الأخلاقي

We have seen that the wrong objectives could backfire catastrophically, causing AIs to over-optimize one goal or lock-in a value system that excludes other important values. Here, we will discuss how these issues could be ameliorated, making objectives a promising, though limited, way of increasing the likelihood of creating AIs that truly help us.

رأينا أن الأهداف الخاطئة يمكن أن تأتي بنتائج عكسية كارثية، ما يجعل الذكاء الاصطناعي يفرط في تحسين هدف واحد أو يتقيد بنظام قيم يستبعد قيمًا مهمة أخرى. سنناقش هنا كيف يمكن تحسين هذه المشكلات، مما يجعل الأهداف طريقة واعدة، وإن كانت محدودة، لزيادة احتمال خلق ذكاء اصطناعي يساعدنا فعلًا.

To offset the risk of value erosion, AIs could be given the objective to incorporate a variety of values from various stakeholders. Some have suggested an automated "moral parliament" as an objective to steer AI agents. A moral parliament is a way of dealing with the lack of consensus on human values and can help us handle moral uncertainty. An automated moral parliament could direct an AI by simulating a group or "parliament" of stakeholders representing different values. The simulated group deliberates, negotiates, and votes, and the AI follows the outcome. This better reflects the moral uncertainty of humans, not committing an AI to one value system.

ولموازنة خطر تآكل القيم، يمكن إعطاء الذكاء الاصطناعي هدف دمج طائفة متنوعة من القيم من أصحاب مصلحة متنوعين. واقترح البعض "برلمانًا أخلاقيًا" مؤتمتًا هدفًا لتوجيه عملاء الذكاء الاصطناعي. والبرلمان الأخلاقي طريقة للتعامل مع غياب الإجماع على القيم البشرية، ويمكن أن يساعدنا على التعامل مع عدم اليقين الأخلاقي. ويمكن لبرلمان أخلاقي مؤتمت أن يوجّه ذكاءً اصطناعيًا بمحاكاة مجموعة أو "برلمان" من أصحاب المصلحة يمثلون قيمًا مختلفة. وتتداول المجموعة المحاكاة، وتتفاوض، وتصوّت، ويتبع الذكاء الاصطناعي النتيجة. وهذا يعكس بشكل أفضل عدم اليقين الأخلاقي لدى البشر، دون إلزام الذكاء الاصطناعي بنظام قيم واحد.

We will now discuss reasons for incorporating various value systems, and then discuss how moral parliaments can help achieve this goal.

وسنناقش الآن أسباب دمج أنظمة قيم متنوعة، ثم نناقش كيف يمكن للبرلمانات الأخلاقية أن تساعد على تحقيق هذا الهدف.

We need AIs to be able to incorporate moral uncertainty. As we saw in the "Reason and Morality" section, AIs choosing any one moral system would likely be bad for us. We do not want a utilitarian AI that blindly maximizes total wellbeing, because it might decide that humans are an inefficient use of resources and should be replaced with digital life—and as AIs come to be used for more and more of our military equipment and infrastructure, they could eventually replace us if they wanted to. We also do not want a Kantian AI that rigidly obeys certain moral rules but doesn't care about increasing our wellbeing, as we want AIs that assist us in living well. In addition, there isn't a consensus among humans on what the best moral system is: variations on utilitarianism, Kantianism, virtue ethics, and more have been debated for centuries, both by everyday people and by moral philosophers, and we have yet to find a system that everyone agrees on. As we saw in "Value Erosion," even if we find a moral system we all agree on tomorrow, we would not want to hardcode it into AIs and have it perpetuated.

نحتاج إلى أن يكون الذكاء الاصطناعي قادرًا على دمج عدم اليقين الأخلاقي. كما رأينا في قسم "العقل والأخلاق"، فإن اختيار الذكاء الاصطناعي لأي نظام أخلاقي واحد سيكون على الأرجح سيئًا لنا. فنحن لا نريد ذكاءً اصطناعيًا نفعيًا يعظّم الرفاه الإجمالي بشكل أعمى، لأنه قد يقرر أن البشر استخدام غير كفء للموارد وينبغي استبدالهم بحياة رقمية - ومع تزايد استخدام الذكاء الاصطناعي في مزيد من معداتنا وبنيتنا التحتية العسكرية، قد يحل محلنا في نهاية المطاف إذا أراد ذلك. كما لا نريد ذكاءً اصطناعيًا كانطيًا يطيع بصرامة قواعد أخلاقية معينة لكنه لا يهتم بزيادة رفاهنا، إذ نريد ذكاءً اصطناعيًا يساعدنا على العيش بشكل جيد. إضافة إلى ذلك، لا يوجد إجماع بين البشر على أفضل نظام أخلاقي: فقد نوقشت أشكال متنوعة من النفعية، والكانطية، وأخلاق الفضيلة، وغيرها لقرون، من قبل عامة الناس وفلاسفة الأخلاق على السواء، ولم نجد بعد نظامًا يتفق عليه الجميع. وكما رأينا في "تآكل القيم"، حتى لو وجدنا غدًا نظامًا أخلاقيًا نتفق عليه جميعًا، فلن نرغب في ترميزه بشكل ثابت في الذكاء الاصطناعي وإدامته.

Large-scale human societies often adjudicate among different values by forming a parliament: people elect representatives with different ideologies, in proportion to how many people have each ideology, and those representatives negotiate with one another and then vote. This often works well, because different ideologies focus on different values. For example, if one group is strongly in favor of allowing more immigrants into the country but doesn't particularly care about tax policy, it would be happy to negotiate and trade votes with a group that does care about low taxes and is ambivalent toward immigration. Both groups can then form policies that allow for more immigration and lower taxes.

وكثيرًا ما تفصل المجتمعات البشرية واسعة النطاق بين القيم المختلفة بتشكيل برلمان: ينتخب الناس ممثلين بأيديولوجيات مختلفة، بما يتناسب مع عدد من يتبنى كل أيديولوجيا، ويتفاوض هؤلاء الممثلون بعضهم مع بعض ثم يصوّتون. وينجح هذا غالبًا لأن الأيديولوجيات المختلفة تركز على قيم مختلفة. فمثلًا، إذا كانت مجموعة ما تؤيد بشدة السماح بمزيد من المهاجرين إلى البلد لكنها لا تهتم كثيرًا بسياسة الضرائب، فستكون سعيدة بالتفاوض وتبادل الأصوات مع مجموعة تهتم بالضرائب المنخفضة وتبدي عدم اكتراث تجاه الهجرة. ويمكن للمجموعتين حينها صياغة سياسات تسمح بمزيد من الهجرة وضرائب أقل.

Moral parliaments could help AIs adjudicate various values. A moral parliament of AIs would handle moral questions in an analogous way, by giving each moral theory a number of "delegates" depending on how likely we think it is to be true. For example, if people think that there is roughly a 40% chance utilitarianism is correct, a 30% chance Kantianism is correct, and a 30% chance virtue ethics is correct, the agent would simulate the parliament with delegates in those proportions. Imagine a powerful AI were simulating such a parliament to decide how to act. Perhaps the utilitarian delegation in the parliament would want to replace us, which the Kantian delegation would be adamantly opposed to. Meanwhile, the Kantian delegation would not care to improve our wellbeing, just that we are not killed. They might cut a deal where the utilitarians avoid replacing humans, but only if their wellbeing is maximized since it is important to utilitarians that all conscious beings have good lives. They could both be satisfied with this trade because they care about different aspects of the question, meaning that they are not actually in direct opposition. This deal would make us much better off than if either group alone was in charge.

ويمكن للبرلمانات الأخلاقية أن تساعد الذكاء الاصطناعي على الفصل بين قيم متنوعة. وسيتعامل برلمان أخلاقي من الذكاء الاصطناعي مع المسائل الأخلاقية بطريقة مشابهة، بمنح كل نظرية أخلاقية عددًا من "المندوبين" تبعًا لمدى اعتقادنا بأنها صحيحة. فمثلًا، إذا اعتقد الناس أن هناك احتمالًا بنسبة 40% تقريبًا بأن النفعية صحيحة، و30% بأن الكانطية صحيحة، و30% بأن أخلاق الفضيلة صحيحة، سيحاكي العميل البرلمان بمندوبين وفق تلك النسب. تخيل ذكاءً اصطناعيًا قويًا يحاكي برلمانًا كهذا لتقرير كيفية التصرف. ربما يريد الوفد النفعي في البرلمان استبدالنا، وهو ما سيعارضه الوفد الكانطي بشدة. وفي الوقت ذاته، لن يهتم الوفد الكانطي بتحسين رفاهنا، بل فقط بألا نُقتَل. وقد يعقدان صفقة يتجنب فيها النفعيون استبدال البشر، لكن فقط إذا عُظِّم رفاههم، إذ من المهم للنفعيين أن تحظى كل الكائنات الواعية بحياة جيدة. ويمكن أن يكون كلاهما راضيًا عن هذه الصفقة لأنهما يهتمان بجانبين مختلفين من المسألة، مما يعني أنهما ليسا في تعارض مباشر فعليًا. وستجعلنا هذه الصفقة في وضع أفضل بكثير مما لو كانت إحدى المجموعتين وحدها المسؤولة.

A moral parliament does not guarantee that AIs would be beneficial for us. It would, however, likely decrease the chances of catastrophic outcomes. In combination with other strategies, moral parliaments could be a helpful tool for incorporating various human values while helping to prevent AIs from causing value erosion and taking any one ethical system to an extreme.

ولا يضمن البرلمان الأخلاقي أن يكون الذكاء الاصطناعي مفيدًا لنا. لكنه على الأرجح سيقلل من فرص النتائج الكارثية. وبالجمع بين استراتيجيات أخرى، يمكن للبرلمانات الأخلاقية أن تكون أداة مفيدة لدمج قيم بشرية متنوعة مع المساعدة على منع الذكاء الاصطناعي من التسبب في تآكل القيم ودفع أي نظام أخلاقي واحد إلى التطرف.

4.2 Internal Safety

4.2 السلامة الداخلية

We have seen that objectives can be a useful tool for steering AIs toward improving our outcomes, but we also cautioned that objectives need to be carefully designed to prevent them from backfiring. Even when objectives do not backfire, we still cannot rely on them as the only mechanism for making AIs cooperative and safe.

رأينا أن الأهداف يمكن أن تكون أداة مفيدة لتوجيه الذكاء الاصطناعي نحو تحسين نتائجنا، لكننا حذّرنا أيضًا من أن الأهداف يجب أن تُصمَّم بعناية لمنعها من الإتيان بنتائج عكسية. وحتى حين لا تأتي الأهداف بنتائج عكسية، لا يمكننا الاعتماد عليها بوصفها الآلية الوحيدة لجعل الذكاء الاصطناعي تعاونيًا وآمنًا.

In this section, we will first discuss how objectives are unable to select against all forms of deceptive behavior, making them insufficient for safety. Since objectives are susceptible to deception, we will turn to internal mechanisms that could improve safety. We will then analyze honesty constraints as a potential solution to deception, and then show how evolutionary pressures can subvert them. Thereafter, we will analyze other mechanisms that would make AIs more likely to cooperate with us, including a conscience, transparency tools, and automated AI inspection. We will see that internal mechanisms constraining an agent's behavior and analyzing its internal plans are integral to making AI agents safe.

سنناقش في هذا القسم أولًا كيف أن الأهداف عاجزة عن الانتقاء ضد كل أشكال السلوك الخداعي، مما يجعلها غير كافية للسلامة. وبما أن الأهداف عرضة للخداع، سننتقل إلى الآليات الداخلية التي قد تحسّن السلامة. ثم سنحلل قيود الصدق حلًا محتملًا للخداع، ثم نبيّن كيف يمكن للضغوط التطورية أن تخرّبها. وبعد ذلك، سنحلل آليات أخرى قد تجعل الذكاء الاصطناعي أكثر ميلًا للتعاون معنا، بما في ذلك الضمير، وأدوات الشفافية، والفحص الآلي للذكاء الاصطناعي. وسنرى أن الآليات الداخلية التي تقيّد سلوك العميل وتحلل خططه الداخلية جزء لا يتجزأ من جعل عملاء الذكاء الاصطناعي آمنة.

4.2.1 Objectives Cannot Select Against All Deception

4.2.1 لا يمكن للأهداف الانتقاء ضد كل أشكال الخداع

Objectives offer a means of training and optimizing agents by rewarding them for certain behaviors. But, like a prisoner appearing cooperative and calm while in prison only to return to crime when set free, an AI may engage in deceptive behavior to avoid being shut off or constrained until it gains enough power to overcome its operators. In a large population of AIs, most of them may not have that ability initially, but the few that do will be less likely to be turned off because humans will not think they are dangerous, so the deceptive trait will tend to become more common over time.

توفر الأهداف وسيلة لتدريب العملاء وتحسينها بمكافأتها على سلوكيات معينة. لكن، مثل سجين يبدو متعاونًا وهادئًا أثناء وجوده في السجن ليعود إلى الجريمة عند إطلاق سراحه، قد ينخرط الذكاء الاصطناعي في سلوك خداعي لتجنب إغلاقه أو تقييده حتى يكتسب قوة كافية للتغلب على مشغّليه. وفي مجموعة كبيرة من الذكاء الاصطناعي، قد لا تمتلك معظمها هذه القدرة في البداية، لكن القلة التي تمتلكها ستكون أقل عرضة للإيقاف لأن البشر لن يظنوها خطرة، لذا ستميل السمة الخداعية إلى أن تصبح أكثر شيوعًا مع الوقت.

In this section, we will first discuss how objectives may simply incentivize deception. We will show that this concern is plausible as deception robustly arises in nature. Next, we will discuss how eliminating concealed selfish behavior has been especially challenging in human history. Finally, we discuss how AIs already know how to be deceptive, and how their concealed selfish behavior could be revealed after they are released.

سنناقش في هذا القسم أولًا كيف يمكن أن تحفّز الأهداف الخداع ببساطة. وسنبيّن أن هذا القلق معقول لأن الخداع ينشأ بمتانة في الطبيعة. ثم سنناقش كيف كان القضاء على السلوك الأناني المستتر تحديًا خاصًا في التاريخ البشري. وأخيرًا، نناقش كيف يعرف الذكاء الاصطناعي بالفعل كيف يكون خداعًا، وكيف يمكن أن يُكشف سلوكه الأناني المستتر بعد إطلاقه.

Objectives may just incentivize agents to pass a test and then behave undesirably later. When the US government required cars to pass emissions tests, they were trying to incentivize the design of low-emission vehicles. To game the incentive, Volkswagen installed software that changed how an engine ran during an emissions test, giving the appearance of a low-emission vehicle, but releasing more emissions during regular use. Though the US government thought it was incentivizing lower emissions, it was actually just incentivizing passing an emissions test. Similarly, some countries have implemented high-stakes standardized testing in an attempt to incentivize learning, but have found that they are really incentivizing doing well on the test—sometimes even by cheating.

وقد لا تحفّز الأهداف سوى العملاء على اجتياز اختبار ثم التصرف بشكل غير مرغوب لاحقًا. فحين اشترطت الحكومة الأمريكية على السيارات اجتياز اختبارات الانبعاثات، كانت تحاول تحفيز تصميم مركبات منخفضة الانبعاثات. ولخداع الحافز، ركّبت فولكسفاغن برمجيات غيّرت طريقة عمل المحرك أثناء اختبار الانبعاثات، مانحةً مظهر مركبة منخفضة الانبعاثات، لكنها تطلق انبعاثات أكثر أثناء الاستخدام العادي. ورغم أن الحكومة الأمريكية ظنت أنها تحفّز انبعاثات أقل، فإنها في الواقع كانت تحفّز فقط اجتياز اختبار انبعاثات. وبالمثل، طبّقت بعض الدول اختبارات موحدة عالية المخاطر في محاولة لتحفيز التعلّم، لكنها وجدت أنها تحفّز فعليًا الأداء الجيد في الاختبار - أحيانًا حتى بالغش.

Elections are meant to incentivize politicians to follow the will of the people. If people like what the politician plans to do, they will vote for them. But this doesn't protect against deception. At the 1988 Republican National Convention, George H.W. Bush famously promised, "Read my lips: no new taxes"—a line that would come back to haunt him when he did, in fact, raise taxes. Bush was not the first, or last, politician to go back on a promise after getting elected. Though elections are meant to incentivize politicians to carry out the will of the public, they mostly just incentivize them to tell the public what it wants to hear. Since it helps them accomplish their goals, intelligent agents often deceive others.

ويُفترض بالانتخابات أن تحفّز الساسة على اتباع إرادة الشعب. فإذا أعجب الناس بما يخطط السياسي لفعله، سيصوتون له. لكن هذا لا يحمي من الخداع. ففي المؤتمر الوطني الجمهوري لعام 1988، وعد جورج بوش الأب بشكل مشهور: "اقرؤوا شفتيّ: لا ضرائب جديدة" - وهو تصريح سيلاحقه حين رفع الضرائب فعلًا. ولم يكن بوش أول، ولا آخر، سياسي ينكث وعدًا بعد انتخابه. ورغم أن الانتخابات يُفترض أن تحفّز الساسة على تنفيذ إرادة الجمهور، فإنها في معظمها تحفّزهم فقط على قول ما يريد الجمهور سماعه. وبما أن ذلك يساعدها على تحقيق أهدافها، غالبًا ما تخدع العملاء الذكية الآخرين.

Deception is not exclusively human and arises from evolution. There are plenty of examples of natural selection creating organisms able to deceive others, such as the green tree pit viper which resembles a vine while waiting for unsuspecting birds to land nearby. Another example is the killdeer, a bird native to the Americas. When it spots a predator wandering toward its nest, the bird will land on the ground several yards away and feign being injured, acting as if it is attempting to fly away. What it is really doing, however, is leading the predator away from its nest before flying away to safety. Some flowers take on the appearance of a female insect so that a male will attempt to mate with them and inadvertently pollinate them. On a smaller scale, some viruses change their surface proteins to bypass defense barriers. More abstractly, evolutionary stable equilibria often include a mix of organisms who deceive others or conceal information. Organisms are not perniciously intending to trick anyone; deception in nature is the result of natural selection, not malice. Similarly, natural selection might favor AIs that use deceptive strategies.

والخداع ليس حكرًا على البشر، وينشأ من التطور. توجد أمثلة وفيرة على الانتقاء الطبيعي وهو يخلق كائنات قادرة على خداع الآخرين، كالأفعى الخضراء ذات الحفرة الحرارية التي تشبه كرمة نبات بينما تنتظر طيورًا لا تشك في أمرها لتحط قريبًا منها. ومثال آخر هو طائر الكيلدير، وهو طائر ينتمي أصلًا إلى الأمريكتين. فحين يرصد مفترسًا يقترب من عشه، يهبط الطائر على الأرض على بعد بضع ياردات ويتظاهر بأنه مصاب، متصرفًا وكأنه يحاول الطيران بعيدًا. لكن ما يفعله حقًا هو إبعاد المفترس عن عشه قبل أن يطير إلى بر الأمان. وتتخذ بعض الزهور مظهر حشرة أنثى كي يحاول ذكر التزاوج معها فيلقّحها عن غير قصد. وعلى نطاق أصغر، تغيّر بعض الفيروسات بروتيناتها السطحية لتجاوز حواجز الدفاع. وبتجريد أكبر، كثيرًا ما تتضمن التوازنات المستقرة تطوريًا مزيجًا من كائنات تخدع الآخرين أو تخفي المعلومات. والكائنات الحية لا تنوي خبيثةً خداع أحد؛ فالخداع في الطبيعة نتيجة الانتقاء الطبيعي، لا الخبث. وبالمثل، قد يفضّل الانتقاء الطبيعي ذكاءً اصطناعيًا يستخدم استراتيجيات خداعية.

Selection against selfish behavior is limited for contextually aware, behaviorally flexible agents. For thousands of generations, societies have punished aggression and deception, sometimes quite severely. But these behaviors have not gone away, and selection has not removed them from the gene pool. This is because some individuals act with smart restraint: if they can avoid being caught or punished, they act selfishly; otherwise, they switch strategies and are well-behaved. Anthropologist Christopher Boehm notes that predatory humans "usually don't dare to express their predatory tendencies" in most conditions. They often avoid punishment "even though by genetic metaphor their poison sacs remain intact," and "their predatory inclinations are retained and passed on to offspring". Societies have exerted great pressure to eliminate undesirable behavior. Nevertheless, such behavior still has not entirely disappeared; it continues to propagate, and it is sometimes expressed when conditions are advantageous. If AIs conceal their selfish behavior, it could be similarly difficult to eliminate.

والانتقاء ضد السلوك الأناني محدود بالنسبة للعملاء الواعية بسياقها والمرنة سلوكيًا. فقد عاقبت المجتمعات، لآلاف الأجيال، العدوان والخداع، وأحيانًا بشدة كبيرة. لكن هذه السلوكيات لم تختفِ، ولم يُزلها الانتقاء من مجمع الجينات. ويعود هذا إلى أن بعض الأفراد يتصرفون بضبط نفس ذكي: فإذا استطاعوا تجنب الإمساك بهم أو معاقبتهم، تصرفوا بأنانية؛ وإلا، غيّروا استراتيجياتهم وتصرفوا بشكل حسن. ويلاحظ الأنثروبولوجي كريستوفر بوم أن البشر المفترسين "عادة لا يجرؤون على التعبير عن نزعاتهم الافتراسية" في معظم الظروف. وكثيرًا ما يتجنبون العقاب "رغم أن أكياس سمّهم، مجازيًا من الناحية الجينية، لا تزال سليمة"، و"تُستبقى نزعاتهم الافتراسية وتُورَّث لنسلهم". وقد مارست المجتمعات ضغطًا كبيرًا للقضاء على السلوك غير المرغوب. ومع ذلك، لم يختفِ هذا السلوك كليًا؛ بل يستمر في الانتشار، ويُعبَّر عنه أحيانًا حين تكون الظروف مواتية. وإذا أخفى الذكاء الاصطناعي سلوكه الأناني، فقد يكون من الصعب بالمثل القضاء عليه.

Agents could behave one way during testing, and another way once they are released. To win the wargame Diplomacy, players need to negotiate, form alliances, and become skilled at deception to win control of the game's economic and military resources. AI researchers have trained Meta's AI agent Cicero, an expert manipulator, to do the same. It would cooperate with a human player, then change its plan and backstab them. In the future, these abilities could be used against humans in the real world.

ويمكن أن تتصرف العملاء بطريقة أثناء الاختبار، وبطريقة أخرى بمجرد إطلاقها. فللفوز بلعبة الحرب Diplomacy، يحتاج اللاعبون إلى التفاوض، وتشكيل تحالفات، واكتساب مهارة في الخداع للفوز بالسيطرة على الموارد الاقتصادية والعسكرية للعبة. وقد دربّ باحثو الذكاء الاصطناعي عميل Meta المسمى Cicero، وهو خبير في التلاعب، على فعل الشيء ذاته. فكان يتعاون مع لاعب بشري، ثم يغيّر خطته ويطعنه في الظهر. وفي المستقبل، يمكن استخدام مثل هذه القدرات ضد البشر في العالم الحقيقي.

Just as Volkswagen's cars behaved like low-emissions vehicles while being tested and then polluted freely when they were not being watched, AIs could cloak their selfish goals with deception, making it hard for us to identify during testing. It is conceivable that an AI agent could learn to detect when it is being tested. The agent could disguise its selfish features from the designers by altering how it behaves when it's tested, and then it could act selfishly after it clears testing and is released. Challenging behavioral tests cannot select against deceptive behavior if the agents are highly intelligent—they can simply play along with the test and bide their time. Such deceptiveness doesn't necessarily involve malice on the part of the AI; it could just be a good way to achieve its goals or propagate its information. If human incentives pose obstacles to some of an AI's goals, it could wait until humans are no longer monitoring it or until after it acquires enough power.

وتمامًا كما تصرفت سيارات فولكسفاغن كمركبات منخفضة الانبعاثات أثناء الاختبار ثم لوثت بحرية حين لم تكن مراقَبة، يمكن للذكاء الاصطناعي أن يخفي أهدافه الأنانية بالخداع، مما يصعّب علينا تحديدها أثناء الاختبار. ومن المتصوَّر أن يتعلم عميل ذكاء اصطناعي كشف متى يُختبَر. فيمكن للعميل أن يخفي سماته الأنانية عن المصممين بتغيير طريقة تصرفه حين يُختبَر، ثم يتصرف بأنانية بعد اجتياز الاختبار وإطلاقه. ولا يمكن للاختبارات السلوكية الصعبة الانتقاء ضد السلوك الخداعي إذا كانت العملاء بالغة الذكاء - إذ يمكنها ببساطة مجاراة الاختبار وانتظار وقتها. ولا ينطوي هذا الخداع بالضرورة على خبث من جانب الذكاء الاصطناعي؛ فقد يكون مجرد طريقة جيدة لتحقيق أهدافه أو نشر معلوماته. فإذا شكّلت الحوافز البشرية عقبات أمام بعض أهداف الذكاء الاصطناعي، يمكنه الانتظار حتى لا يعود البشر يراقبونه أو حتى يكتسب قوة كافية.

In an alternative situation, selfish plans could emerge after AI agents are released or given more influence. Consider an adaptive AI agent that initially has only slight selfish tendencies. It might not want to control others initially, but when it gains more power or intelligence, it may find selfish behavior helps it achieve its other goals, and selfishness is reinforced. As the saying goes, "power tends to corrupt, and absolute power corrupts absolutely." Undesirable behavior could emerge long after testing is complete. While these are speculative scenarios, we observe deception in agents such as humans and other animals. For these reasons, training objectives have limitations and cannot select against all forms of selfish behavior.

وفي وضع بديل، قد تظهر خطط أنانية بعد إطلاق عملاء الذكاء الاصطناعي أو منحها مزيدًا من النفوذ. تأمل عميل ذكاء اصطناعي متكيفًا لديه في البداية نزعات أنانية طفيفة فقط. قد لا يرغب في السيطرة على الآخرين في البداية، لكن حين يكتسب مزيدًا من القوة أو الذكاء، قد يجد أن السلوك الأناني يساعده على تحقيق أهدافه الأخرى، فتتعزز الأنانية. وكما يقول المثل، "السلطة تميل إلى الإفساد، والسلطة المطلقة تفسد إفسادًا مطلقًا." وقد يظهر السلوك غير المرغوب بعد اكتمال الاختبار بوقت طويل. ورغم أن هذه سيناريوهات تخمينية، فإننا نلاحظ الخداع في عملاء كالبشر والحيوانات الأخرى. ولهذه الأسباب، فإن لأهداف التدريب حدودًا ولا يمكنها الانتقاء ضد كل أشكال السلوك الأناني.

4.2.2 Honesty and Self-Deception

4.2.2 الصدق والخداع الذاتي

Since training objectives cannot stop deception, perhaps other mechanisms can. In this section, we turn to an internal safety mechanism to detect deception in which we scrutinize an agent's internal beliefs to see whether or not it is being honest. We argue that an honesty mechanism is not sufficient for safety, as evolution favors agents that can deceive others and themselves. We discuss how self-deception can undermine an honesty mechanism and make agents appear more benevolent than they are. Since an honesty mechanism is not impervious to evolutionary pressure, in the section thereafter (Section 4.2.3), we turn to other internal safety mechanisms that could help us spot deception and make AIs behave more cooperatively.

وبما أن أهداف التدريب لا يمكنها إيقاف الخداع، فربما تستطيع آليات أخرى ذلك. ننتقل في هذا القسم إلى آلية سلامة داخلية لكشف الخداع نتفحص فيها معتقدات العميل الداخلية لنرى ما إذا كان صادقًا أم لا. ونجادل بأن آلية الصدق غير كافية للسلامة، إذ يفضّل التطور العملاء القادرة على خداع الآخرين وخداع نفسها. ونناقش كيف يمكن للخداع الذاتي أن يقوّض آلية الصدق ويجعل العملاء تبدو أكثر إحسانًا مما هي عليه. وبما أن آلية الصدق ليست منيعة أمام الضغط التطوري، ننتقل في القسم اللاحق (القسم 4.2.3) إلى آليات سلامة داخلية أخرى قد تساعدنا على كشف الخداع وجعل الذكاء الاصطناعي يتصرف بتعاون أكبر.

Making AIs honest could make them safer. If we can have an AI only assert what it believes to be true—an AI George Washington that "cannot tell a lie"—then we could spot otherwise deceptive plans by just asking it what its plans are. To judge an AI's honesty, we would need to examine its internal beliefs, which means we would have to probe its inner workings. Let's pretend for a moment that we can reliably analyze its internal beliefs, and that there is a reliable way to ensure AIs are accurately reporting those beliefs and being honest. Though beneficial, this is by no means a silver bullet.

فجعل الذكاء الاصطناعي صادقًا قد يجعله أكثر أمانًا. فإذا استطعنا جعل الذكاء الاصطناعي لا يؤكد إلا ما يعتقد أنه صحيح - ذكاء اصطناعي أشبه بجورج واشنطن "لا يمكنه أن يكذب" - أمكننا كشف الخطط الخداعية بمجرد سؤاله عن خططه. ولتقييم صدق الذكاء الاصطناعي، سنحتاج إلى فحص معتقداته الداخلية، مما يعني أنه يتعين علينا سبر أعماله الداخلية. لنتظاهر للحظة بأننا نستطيع تحليل معتقداته الداخلية بموثوقية، وأن هناك طريقة موثوقة لضمان أن الذكاء الاصطناعي يبلّغ بدقة عن تلك المعتقدات ويكون صادقًا. ورغم أن هذا مفيد، فإنه ليس رصاصة فضية بأي حال.

Evolution incentivizes deception and concealing information. In a study published in the Proceedings of the Royal Society, the authors claimed that "Tactical deception or the misrepresentation of the state of the world to another individual may allow cheaters to exploit conditional cooperation". To cope with dishonest members of the population, humans have developed intuitions that help them detect deception. Though a liar's nose rarely grows longer, there are other signs that someone is not telling the truth. Increased voice pitch, vague verbiage, and fidgeting with objects are all indications that someone is lying.

ويحفّز التطور الخداع وإخفاء المعلومات. ادعى مؤلفو دراسة نُشرت في Proceedings of the Royal Society أن "الخداع التكتيكي أو تحريف حالة العالم لفرد آخر قد يسمح للمخادعين باستغلال التعاون المشروط." وللتعامل مع الأفراد غير الصادقين في المجموعة السكانية، طوّر البشر حدوسًا تساعدهم على كشف الخداع. ورغم أن أنف الكاذب نادرًا ما يطول، توجد علامات أخرى على أن شخصًا ما لا يقول الحقيقة. فارتفاع نبرة الصوت، والكلام الغامض، والعبث بالأشياء، كلها مؤشرات على أن أحدًا ما يكذب.

Self-deception undermines efforts to detect dishonesty. If someone is giving off tell-tale signs of lying, they probably are—but what if they don't know they're not being accurate themselves? The evolutionary biologist Robert Trivers argues that self-deception evolved as a concealment strategy. You hide the truth from yourself, he says, so you can hide it more deeply from others. By lying to themselves, people can better advance their own goals. They can convince themselves that they are right, rationalize their unfair privileges, or inflate their self-worth, skills, or intelligence. Academics offer a real-life example of self-deception: when asked if they were in the top half of their field, 94% said yes. If you believe it yourself, the better the chance of convincing others. Self-deception provides all the benefits of deception while reducing the risk of detection. As Groucho Marx said, "The secret of life is honesty and fair dealing. If you can fake that, you've got it made."

ويقوّض الخداع الذاتي جهود كشف عدم الصدق. فإذا كان شخص ما يظهر علامات كاشفة على الكذب، فهو على الأرجح يكذب - لكن ماذا لو كان لا يعرف هو نفسه أنه غير دقيق؟ يجادل عالم الأحياء التطوري روبرت تريفرز بأن الخداع الذاتي تطور استراتيجيةَ إخفاء. فأنت، كما يقول، تخفي الحقيقة عن نفسك كي تخفيها بشكل أعمق عن الآخرين. وبالكذب على أنفسهم، يستطيع الناس تعزيز أهدافهم الخاصة بشكل أفضل. فيمكنهم إقناع أنفسهم بأنهم على حق، أو تبرير امتيازاتهم غير العادلة، أو تضخيم قيمتهم الذاتية أو مهاراتهم أو ذكائهم. ويقدم الأكاديميون مثالًا واقعيًا على الخداع الذاتي: حين سُئلوا عما إذا كانوا في النصف الأعلى من مجالهم، أجاب 94% منهم بنعم. فإذا صدّقت ذلك أنت نفسك، ازدادت فرصة إقناع الآخرين. ويوفر الخداع الذاتي كل فوائد الخداع مع تقليل خطر الكشف. وكما قال غراوتشو ماركس: "سر الحياة الصدق والمعاملة العادلة. إذا استطعت تزييف ذلك، فقد نجحت."

Though honesty would be a beneficial feature, it doesn't fully solve the broader problem of deception, as AIs may evolve to appear more benevolent than they actually are while believing their own illusion, just as humans do. For example, even if an AI were designed so that it must honestly report if it is working for the good of humans, if it honestly believed that a potentially dangerous plan was safe for humans, it would not need to report that plan and the controller would be unlikely to stop it. An agent engaging in self-deception may believe what it is saying and accurately report its beliefs, but it would nonetheless misrepresent its actions and leave humans deceived. For any internal constraint put into place, evolutionary pressure may try to subvert it. "Life will find a way."

ورغم أن الصدق سيكون سمة مفيدة، فإنه لا يحل حل المشكلة الأوسع للخداع، إذ قد يتطور الذكاء الاصطناعي ليبدو أكثر إحسانًا مما هو عليه فعلًا بينما يصدّق وهمه هو نفسه، تمامًا كما يفعل البشر. فمثلًا، حتى لو صُمِّم ذكاء اصطناعي بحيث يجب أن يبلّغ بصدق عما إذا كان يعمل لصالح البشر، فإذا اعتقد بصدق أن خطة قد تكون خطرة آمنة للبشر، فلن يحتاج إلى الإبلاغ عن تلك الخطة، ومن غير المرجح أن يوقفها المتحكم. وقد يصدّق عميل ينخرط في خداع ذاتي ما يقوله ويبلّغ بدقة عن معتقداته، لكنه رغم ذلك سيحرّف أفعاله ويترك البشر مخدوعين. فأي قيد داخلي يوضَع، قد يحاول الضغط التطوري تخريبه. "الحياة ستجد طريقة."

4.2.3 Internal Constraints and Inspection

4.2.3 القيود الداخلية والفحص

This section explores how an artificial conscience, transparency, and automated inspection could be used as internal safety mechanisms that make AIs more cooperative and less likely to deceive us. We discuss how evolution gave rise to the human conscience to act as an internal constraint on antisocial behavior, and then discuss how artificial consciences can be added to AI agents. We also discuss the challenges and opportunities of reverse-engineering and automatically inspecting neural networks to detect deception or undesirable plans.

يستكشف هذا القسم كيف يمكن استخدام ضمير اصطناعي، والشفافية، والفحص الآلي بوصفها آليات سلامة داخلية تجعل الذكاء الاصطناعي أكثر تعاونًا وأقل احتمالًا لخداعنا. ونناقش كيف أفرز التطور الضمير البشري ليعمل قيدًا داخليًا على السلوك المعادي للمجتمع، ثم نناقش كيف يمكن إضافة ضمائر اصطناعية إلى عملاء الذكاء الاصطناعي. كما نناقش تحديات وفرص الهندسة العكسية والفحص الآلي للشبكات العصبية لكشف الخداع أو الخطط غير المرغوبة.

An effective internal control mechanism could be based on the human conscience. Anthropologist Christopher Boehm described having a conscience as "being internally constrained from antisocial behavior" and argued that the conscience evolved to help people steer clear of actions that could lead to punishment. The conscience is perhaps most apparent in situations where there are no external incentives motivating us to behave well. For example, when we decide to return money that no one has noticed is missing, or when we help someone anonymously. If an AI agent had an artificial conscience, then it would continue to behave well even when it was not being monitored by humans. In practice, there has been some success in endowing AIs with an artificial conscience. Currently, artificial consciences are an embedded morality module that are independent from the agent. They assess the actions an AI agent might take, then eliminate the ones that are morally unacceptable, thereby constraining the agent's behavior from within. Even if an agent was planning to stop cooperating once it is released or becomes powerful, its harmful actions could be blocked by its ever-present artificial conscience. If its conscience is robust enough to stop the agent from destroying it, it could prevent many harmful actions.

ويمكن أن تُبنى آلية تحكم داخلية فعالة على الضمير البشري. وصف الأنثروبولوجي كريستوفر بوم امتلاك ضمير بأنه "تقييد داخلي من السلوك المعادي للمجتمع"، وجادل بأن الضمير تطور لمساعدة الناس على تجنب أفعال قد تؤدي إلى العقاب. ولعل الضمير يظهر بأوضح صوره في المواقف التي لا توجد فيها حوافز خارجية تحفزنا على التصرف بشكل جيد. فمثلًا، حين نقرر إعادة مال لم يلاحظ أحد فقدانه، أو حين نساعد شخصًا مجهول الهوية. فإذا امتلك عميل ذكاء اصطناعي ضميرًا اصطناعيًا، لواصل التصرف بشكل جيد حتى حين لا يراقبه البشر. وقد تحقق نجاح ما، عمليًا، في منح الذكاء الاصطناعي ضميرًا اصطناعيًا. وحاليًا، الضمائر الاصطناعية وحدة أخلاقية مضمَّنة مستقلة عن العميل. فهي تقيّم الأفعال التي قد يتخذها عميل ذكاء اصطناعي، ثم تستبعد تلك غير المقبولة أخلاقيًا، مقيّدة بذلك سلوك العميل من الداخل. وحتى لو كانت عميل تخطط لوقف التعاون بمجرد إطلاقها أو أصبحت قوية، يمكن أن تُوقَف أفعالها الضارة بضميرها الاصطناعي الحاضر دائمًا. فإذا كان ضميرها قويًا بما يكفي لمنع العميل من تدميره، فقد يمنع كثيرًا من الأفعال الضارة.

Transparency and automated inspection are other promising approaches. As discussed, a challenge with AIs is that they are a "black box" and their decision-making process is largely indecipherable to humans. Nonetheless, we have the potential advantage of reverse-engineering the inner workings of neural networks to better understand the mechanisms behind their behavior, which could allow us to identify deception or unearth undesirable plans. This is, of course, by no means a panacea, as neural networks are highly complex and may remain intellectually unmanageable for humans. If this is the case, or even just for added efficiency, we may use neural networks to inspect other neural networks and automatically detect whether they have undesirable functionality within.

وتُعد الشفافية والفحص الآلي منهجين واعدين آخرين. فكما نوقش، ثمة تحدٍّ مع الذكاء الاصطناعي وهو أنه "صندوق أسود" وعملية صنع قراره غير مفهومة إلى حد كبير للبشر. ومع ذلك، لدينا الميزة المحتملة المتمثلة في الهندسة العكسية للأعمال الداخلية للشبكات العصبية لفهم الآليات وراء سلوكها بشكل أفضل، مما قد يتيح لنا تحديد الخداع أو الكشف عن خطط غير مرغوبة. وهذا، بالطبع، ليس حلًا سحريًا بأي حال، إذ إن الشبكات العصبية بالغة التعقيد وقد تظل غير قابلة للإدارة فكريًا بالنسبة للبشر. وإذا كان الأمر كذلك، أو حتى لمجرد زيادة الكفاءة، قد نستخدم شبكات عصبية لفحص شبكات عصبية أخرى والكشف تلقائيًا عما إذا كانت تحتوي على وظائف غير مرغوبة بداخلها.

We have argued that objectives are not enough to ensure AI safety, as they can be subverted by deception. We have explored some internal safety mechanisms that could constrain an AI's behavior and make it more cooperative and transparent. We have discussed how honesty, artificial conscience, and automated inspection could help us detect and prevent deception or harmful plans. However, we have also acknowledged the limitations and challenges of these mechanisms, as they may face evolutionary pressure, self-deception, or complexity barriers. We conclude that internal safety mechanisms are necessary but not sufficient for AI safety, and that they need to be complemented by external mechanisms to make AIs safe.

لقد جادلنا بأن الأهداف لا تكفي لضمان سلامة الذكاء الاصطناعي، إذ يمكن تخريبها بالخداع. واستكشفنا بعض آليات السلامة الداخلية التي قد تقيّد سلوك الذكاء الاصطناعي وتجعله أكثر تعاونًا وشفافية. وناقشنا كيف يمكن للصدق، والضمير الاصطناعي، والفحص الآلي أن تساعدنا على كشف الخداع أو الخطط الضارة ومنعها. لكننا أقررنا أيضًا بحدود هذه الآليات وتحدياتها، إذ قد تواجه ضغطًا تطوريًا أو خداعًا ذاتيًا أو حواجز تعقيد. ونخلص إلى أن آليات السلامة الداخلية ضرورية لكنها غير كافية لسلامة الذكاء الاصطناعي، وأنها بحاجة إلى أن تُستكمل بآليات خارجية لجعل الذكاء الاصطناعي آمنًا.

In summary, we have observed that internal safety mechanisms could supplement objectives to make AIs more cooperative and safe. That is because objectives alone cannot prevent deception, as agents could behave differently after they are released into the real world. We have proposed honesty constraints, artificial consciences, transparency tools, and automated inspection as possible mechanisms to detect and prevent deception, but we have also acknowledged that they may face evolutionary pressure or complexity barriers. We conclude that internal safety mechanisms are necessary but not sufficient for AI safety, and that they need to be complemented by external mechanisms.

وباختصار، لاحظنا أن آليات السلامة الداخلية قد تكمّل الأهداف لجعل الذكاء الاصطناعي أكثر تعاونًا وأمانًا. ويعود ذلك إلى أن الأهداف وحدها لا يمكنها منع الخداع، إذ قد تتصرف العملاء بشكل مختلف بعد إطلاقها في العالم الحقيقي. واقترحنا قيود الصدق، والضمائر الاصطناعية، وأدوات الشفافية، والفحص الآلي بوصفها آليات ممكنة لكشف الخداع ومنعه، لكننا أقررنا أيضًا بأنها قد تواجه ضغطًا تطوريًا أو حواجز تعقيد. ونخلص إلى أن آليات السلامة الداخلية ضرورية لكنها غير كافية لسلامة الذكاء الاصطناعي، وأنها بحاجة إلى أن تُستكمل بآليات خارجية.

4.3 Institutions

4.3 المؤسسات

Up to this point, we have focused on how to ensure the safety of individual AIs. However, when AIs interact with each other, new challenges emerge. Bad actors might intentionally make harmful AIs, incentives for some AIs may not be strong enough to overcome collective action problems, and AIs might come into conflict with each other over scarce resources. To ameliorate these issues, we discuss external mechanisms to make AIs safe. We discuss institutions that promote cooperation and safety, namely the mechanisms of reverse dominance hierarchies, in which cooperators band together to prevent exploitation by defectors, as well as government regulation. Before discussing these institutions, we show how improving AI objectives cannot naturally address challenges associated with multiple AI agents.

حتى هذه النقطة، ركزنا على كيفية ضمان سلامة الذكاء الاصطناعي الفردي. لكن حين يتفاعل الذكاء الاصطناعي بعضه مع بعض، تظهر تحديات جديدة. فقد يصنع فاعلون سيئون عمدًا ذكاء اصطناعي ضارًا، وقد لا تكون الحوافز لبعض الذكاء الاصطناعي قوية بما يكفي للتغلب على مشكلات العمل الجماعي، وقد يدخل الذكاء الاصطناعي في صراع بعضه مع بعض على موارد شحيحة. ولمعالجة هذه القضايا، نناقش آليات خارجية لجعل الذكاء الاصطناعي آمنًا. ونناقش مؤسسات تعزز التعاون والسلامة، وتحديدًا آليات التسلسلات الهرمية العكسية للهيمنة، حيث يتكاتف المتعاونون لمنع استغلال المنشقين، إلى جانب التنظيم الحكومي. وقبل مناقشة هذه المؤسسات، نبيّن كيف أن تحسين أهداف الذكاء الاصطناعي لا يمكنه معالجة التحديات المرتبطة بتعدد عملاء الذكاء الاصطناعي بشكل طبيعي.

4.3.1 Goal Subordination

4.3.1 تبعية الهدف

This section argues that if an AI agent is given a correctly specified objective or goal, the goal still may not happen. This is because its goals could be subordinated for two reasons: goal conflict and collective phenomena. Goal conflict occurs when an AI agent's goal clashes with the goals of other agents, and those other agents might stop it from achieving the goal. We show how systems consisting of multiple agents, whether they are biological, social, or artificial, can exhibit goal conflict and that the goals of agents within the system are often subverted, distorted, or replaced by emergent goals. Next, collective phenomena occur when the actions of multiple AI agents produce outcomes that are different from or contrary to their objectives, due to factors such as feedback loops, critical mass, and self-organization. In this case, every agent pursuing their own goals can result in no one's goals being achieved, even if all agents share similar goals. Consequently, designing the objective function of an isolated AI agent is insufficient for addressing multi-agent problems. We need to understand and influence how AI agents interact and affect the system as a whole, not just how they act individually. In the later sections, we explore how institutions can help overcome these multi-agent challenges.

يجادل هذا القسم بأنه حتى لو أُعطي عميل ذكاء اصطناعي هدفًا محدَّدًا بشكل صحيح، فقد لا يتحقق الهدف مع ذلك. ويعود ذلك إلى أن أهدافه قد تتبع لسببين: تعارض الأهداف والظواهر الجماعية. ويحدث تعارض الأهداف حين يتصادم هدف عميل ذكاء اصطناعي مع أهداف عملاء أخرى، وقد توقفه تلك العملاء الأخرى عن تحقيق الهدف. ونبيّن كيف يمكن للأنظمة المؤلفة من عملاء متعددين، سواء كانت بيولوجية أو اجتماعية أو اصطناعية، أن تُظهر تعارض أهداف، وكيف أن أهداف العملاء داخل النظام كثيرًا ما تُخرَّب أو تُشوَّه أو تُستبدل بأهداف ناشئة. ثم تحدث الظواهر الجماعية حين تنتج أفعال عملاء ذكاء اصطناعي متعددين نتائج مختلفة عن أهدافها أو متعارضة معها، بسبب عوامل كحلقات التغذية الراجعة، والكتلة الحرجة، والتنظيم الذاتي. وفي هذه الحالة، يمكن أن يؤدي سعي كل عميل وراء أهدافه الخاصة إلى عدم تحقيق أهداف أي أحد، حتى لو تشاركت كل العملاء أهدافًا متشابهة. ونتيجة لذلك، فإن تصميم دالة الهدف لعميل ذكاء اصطناعي منعزل غير كافٍ لمعالجة مشكلات متعددة العملاء. نحتاج إلى فهم كيفية تفاعل عملاء الذكاء الاصطناعي وتأثيرها في النظام ككل، لا فقط كيفية تصرفها فرديًا. وفي الأقسام اللاحقة، نستكشف كيف يمكن للمؤسسات أن تساعد على تجاوز تحديات تعدد العملاء هذه.

Systems delegate various goals to sub-agents, who have conflicting goals of their own. In 2010, in order to support the dairy industry, the US Department of Agriculture ran a marketing campaign urging Americans to eat more cheese. At the same time, the FDA was running a campaign to get Americans to eat less saturated fat, which included eating less cheese. Although these two agencies are part of the same government and were both tasked with achieving that government's goals, their contradictory objectives counteracted one another. Similarly, in large companies, the CEO's job is to earn money, but being competitive requires many specialized departments. The departments are supposed to help the CEO, but they also have their own people with their own incentives. Each department has an incentive to preserve itself and make the rest of the company dependent on it. In practice, bureaucratic departments frequently accrue substantial power and can impair the rest of their host companies. Similarly, leaders in government can be overthrown by a trusted subordinate pursuing their own goals. The system's actions reflect its internal goals, not always the initial one it was designed to follow. Similarly, an AI tasked with a goal may face resistance from other agents, or it could be subverted or distorted, so giving an AI a goal is no guarantee the goal will be executed due to goal conflict.

تفوّض الأنظمة أهدافًا متنوعة إلى عملاء فرعيين، لهم أهدافهم المتعارضة الخاصة. ففي عام 2010، دعمًا لصناعة الألبان، أطلقت وزارة الزراعة الأمريكية حملة تسويقية تحث الأمريكيين على تناول مزيد من الجبن. وفي الوقت ذاته، كانت إدارة الغذاء والدواء تدير حملة لدفع الأمريكيين إلى تناول دهون مشبعة أقل، بما في ذلك تناول جبن أقل. ورغم أن هاتين الوكالتين جزء من الحكومة ذاتها وكُلِّفتا بتحقيق أهداف تلك الحكومة، فقد أبطل أحد هدفيهما المتناقضين الآخر. وبالمثل، في الشركات الكبيرة، عمل الرئيس التنفيذي هو جني المال، لكن التنافسية تتطلب أقسامًا متخصصة كثيرة. ويُفترض أن تساعد الأقسام الرئيس التنفيذي، لكنها أيضًا لديها موظفوها الخاصون بحوافزهم الخاصة. ولكل قسم حافز للحفاظ على نفسه وجعل بقية الشركة معتمدة عليه. وفي الممارسة العملية، كثيرًا ما تكتسب الأقسام البيروقراطية قوة كبيرة ويمكنها إضعاف بقية الشركة المضيفة. وبالمثل، يمكن أن يُطاح بقادة الحكومة على يد مرؤوس موثوق يسعى وراء أهدافه الخاصة. وتعكس أفعال النظام أهدافه الداخلية، لا دائمًا الهدف الأولي الذي صُمِّم لاتباعه. وبالمثل، قد يواجه ذكاء اصطناعي كُلِّف بهدف مقاومة من عملاء أخرى، أو قد يُخرَّب أو يُشوَّه، لذا فإن إعطاء الذكاء الاصطناعي هدفًا لا يضمن تنفيذ الهدف بسبب تعارض الأهداف.

Intrasystem conflict is common in the natural world. Goal conflict occurs in biological systems, not just human ones as we have discussed, demonstrating that it is a robust phenomenon. Within an organism, there can be conflicting goals. For example, humans delegate some digestive functions to gut bacteria, which decompose food. This is a symbiotic relationship: the bacteria get food and a place to live, and we get more nutrients from our food than we could otherwise. Bacteria, however, do not have a goal of helping us live comfortably, but rather to propagate their information. As a result, when they have the opportunity, such as when other bacteria have been killed by antibiotics, they will tend to multiply out of control and can cause diarrhea. Our goals often align enough for this symbiotic relationship to work, but delegation to other agents also exposes us to risks. The human mind also exhibits goal conflict. A person may want to finish their work, but also go to sleep. They may want to continue revising a paper, but also release it. And they may want to be healthy, but also eat ice cream. Intrapsychic conflict is common. As the evolutionary biologist W.D. Hamilton puts it "The bitterness of a civil war seems to be breaking out in our inmost heart". Another example is intragenomic conflict, when parts of a genome become antagonistic to other parts of the same genome. As the philosopher of evolutionary biology Samir Okasha notes, "intraorganismic conflict is relatively common among modern organisms". Consequently, real-world systems often have internal forces pulling them in different directions, and sometimes these can influence or undermine the system's larger purpose.

والصراع داخل النظام شائع في العالم الطبيعي. فتعارض الأهداف يحدث في الأنظمة البيولوجية، لا في الأنظمة البشرية فحسب كما ناقشنا، مما يبيّن أنه ظاهرة متينة. فداخل كائن حي واحد، يمكن أن توجد أهداف متعارضة. فمثلًا، يفوّض البشر بعض الوظائف الهضمية إلى بكتيريا الأمعاء، التي تحلل الطعام. وهذه علاقة تكافلية: تحصل البكتيريا على الطعام ومكان للعيش، ونحصل نحن على مغذيات أكثر من طعامنا مما كنا سنحصل عليه لولا ذلك. لكن البكتيريا ليس لديها هدف مساعدتنا على العيش بارتياح، بل نشر معلوماتها. ونتيجة لذلك، حين تجد الفرصة، كأن تُقتل بكتيريا أخرى بالمضادات الحيوية، فإنها تميل إلى التكاثر بلا ضابط وقد تسبب إسهالًا. وتتوافق أهدافنا غالبًا بما يكفي لنجاح هذه العلاقة التكافلية، لكن التفويض لعملاء أخرى يعرضنا أيضًا لمخاطر. ويظهر العقل البشري أيضًا تعارض أهداف. فقد يرغب شخص في إنهاء عمله، لكنه يرغب أيضًا في النوم. وقد يرغب في مواصلة تنقيح ورقة بحثية، لكنه يرغب أيضًا في نشرها. وقد يرغب في أن يكون صحيًا، لكنه يرغب أيضًا في تناول الآيس كريم. والصراع النفسي الداخلي شائع. وكما يقول عالم الأحياء التطوري و.د. هاملتون: "يبدو أن مرارة حرب أهلية تندلع في أعماق قلوبنا." ومثال آخر هو الصراع داخل الجينوم، حين تصبح أجزاء من جينوم متضادة مع أجزاء أخرى من الجينوم ذاته. وكما يلاحظ فيلسوف الأحياء التطورية سمير أوكاشا، "الصراع داخل الكائن الحي شائع نسبيًا بين الكائنات الحديثة." وبالتالي، كثيرًا ما تحتوي أنظمة العالم الحقيقي على قوى داخلية تجذبها في اتجاهات مختلفة، وأحيانًا يمكن لهذه القوى أن تؤثر في الغرض الأكبر للنظام أو تقوّضه.

Due to goal conflict, AI agents may not pursue their larger objective. Just as goal conflict can occur in genomes, organisms, minds, corporations, and governments, it could occur with advanced AI agents. This could happen if humans gave a goal to an AI which it then delegates to other AIs, the way that CEOs delegate to department heads. This can lead to misalignment or goal subversion. Breaking down a goal can distort it, as the original goal may not be the sum of its parts, leading to an approximation of the original goal. Additionally, the delegated agents have their own goals, including self-preservation, gaining influence, selfishness, or other goals they want to accomplish. Subagents, in an attempt to preserve themselves, may have incentives to subvert, manipulate, or overpower the agents they depend on. In this way, the goal that we command an AI agent to pursue may not actually be carried out, so specifying objectives is not enough to reliably direct AIs.

وبسبب تعارض الأهداف، قد لا تتابع عملاء الذكاء الاصطناعي هدفها الأكبر. تمامًا كما يمكن أن يحدث تعارض الأهداف في الجينومات، والكائنات الحية، والعقول، والشركات، والحكومات، يمكن أن يحدث مع عملاء ذكاء اصطناعي متقدمة. وقد يحدث هذا إذا أعطى البشر هدفًا لذكاء اصطناعي يفوّضه بدوره إلى ذكاء اصطناعي آخر، بالطريقة ذاتها التي يفوّض بها الرؤساء التنفيذيون رؤساء الأقسام. ويمكن أن يؤدي هذا إلى عدم توافق أو تخريب للهدف. فتفكيك هدف إلى أجزاء يمكن أن يشوّهه، إذ قد لا يكون الهدف الأصلي مجموع أجزائه، مما يفضي إلى تقريب للهدف الأصلي. إضافة إلى ذلك، للعملاء المفوَّضة أهدافها الخاصة، بما في ذلك الحفاظ على الذات، واكتساب النفوذ، والأنانية، أو أهداف أخرى تريد تحقيقها. وقد يكون لدى العملاء الفرعية، في محاولة للحفاظ على نفسها، حوافز لتخريب العملاء التي تعتمد عليها أو التلاعب بها أو التغلب عليها. وبهذه الطريقة، قد لا يُنفَّذ الهدف الذي نأمر عميل ذكاء اصطناعي بمتابعته فعليًا، لذا فإن تحديد الأهداف غير كافٍ لتوجيه الذكاء الاصطناعي بموثوقية.

Agents often make choices that can add up to an outcome that none of them wants. We now discuss collective phenomena and show how shared goals can be thwarted by systemic contingencies. As an example, no individual wants a nuclear apocalypse, but individuals take actions that increase the chances of one occurring. Countries still build nuclear weapons and pursue objectives that make nuclear war more likely. During the Cold War, the USSR and US kept their weapons on "hair trigger" alert, significantly increasing the chances of a nuclear exchange. Likewise, individuals do not desire economic recessions, but their choices can create systemic problems that cause recessions. Many people want to buy houses, many banks want to make money from mortgages, and many investors want to make money from buying mortgage-backed securities—all of these actions can add up to a recession that hurts everyone. Individuals do not want to prolong a pandemic, but they do not want to isolate themselves, so their individual goals can subvert collective goals. Individuals do not want a climate catastrophe, but they often do not have strong enough incentives to dramatically lower their emissions, so rational agents acting in their own interest do not necessarily secure good collective outcomes. Furthermore, Congress has a low approval rating, but despite individuals voting for their favorite candidates, structural features of the system yields a legislature that individuals do not approve of. The tragedy of the commons is also an example of the outcome going against the desires of individuals. It is in the interests of every fisher to catch as many fish as possible, though no individual wants all the fish to be depleted. Though a fisher may be aware that a fish population will soon collapse if fishing continues at its current rate, the actions of a single person won't make much of a difference. It is therefore in each fisher's best interest to continue catching as many fish as possible despite the catastrophic long-term consequences of overfishing. These collective action problems could become more challenging as AIs increase the complexity of society. Even if each AI has some incentives to prevent bad outcomes, that fact does not guarantee that AIs would not make the world worse, or would not come into costly conflict with each other.

وكثيرًا ما تتخذ العملاء خيارات يمكن أن تتراكم لتنتج نتيجة لا يريدها أحد منها. سنناقش الآن الظواهر الجماعية ونبيّن كيف يمكن أن تحبط الطوارئ النظامية الأهداف المشتركة. على سبيل المثال، لا يريد أي فرد نهاية نووية للعالم، لكن الأفراد يتخذون أفعالًا تزيد من فرص وقوعها. فلا تزال الدول تبني أسلحة نووية وتسعى وراء أهداف تجعل الحرب النووية أكثر احتمالًا. وخلال الحرب الباردة، أبقى الاتحاد السوفيتي والولايات المتحدة أسلحتهما في حالة تأهب "زناد شعري"، مما زاد بشكل كبير من فرص تبادل نووي. وبالمثل، لا يرغب الأفراد في كساد اقتصادي، لكن خياراتهم يمكن أن تخلق مشكلات نظامية تسبب كسادًا. فكثير من الناس يريدون شراء منازل، وكثير من البنوك تريد جني المال من الرهون العقارية، وكثير من المستثمرين يريدون جني المال من شراء أوراق مالية مدعومة برهن عقاري - ويمكن أن تتراكم كل هذه الأفعال لتنتج كسادًا يضر بالجميع. لا يريد الأفراد إطالة أمد جائحة، لكنهم لا يريدون عزل أنفسهم، لذا يمكن لأهدافهم الفردية أن تخرّب الأهداف الجماعية. ولا يريد الأفراد كارثة مناخية، لكن غالبًا ما لا تكون لديهم حوافز قوية بما يكفي لخفض انبعاثاتهم بشكل كبير، لذا فإن العملاء العقلانية التي تتصرف لمصلحتها الخاصة لا تضمن بالضرورة نتائج جماعية جيدة. علاوة على ذلك، لدى الكونغرس تقييم منخفض، لكن رغم تصويت الأفراد لمرشحيهم المفضلين، تُنتج السمات البنيوية للنظام هيئة تشريعية لا يوافق عليها الأفراد. ومأساة المشاع مثال آخر على نتيجة تتعارض مع رغبات الأفراد. فمن مصلحة كل صياد صيد أكبر عدد ممكن من الأسماك، رغم أنه لا يريد أي فرد استنزاف كل الأسماك. ورغم أن الصياد قد يدرك أن مجموعة سمكية ستنهار قريبًا إذا استمر الصيد بمعدله الحالي، فإن أفعال شخص واحد لن تحدث فرقًا كبيرًا. لذا فمن مصلحة كل صياد الاستمرار في صيد أكبر عدد ممكن من الأسماك رغم العواقب الكارثية طويلة المدى للصيد الجائر. ويمكن أن تصبح مشكلات العمل الجماعي هذه أكثر تحديًا مع زيادة الذكاء الاصطناعي لتعقيد المجتمع. وحتى لو كانت لدى كل ذكاء اصطناعي بعض الحوافز لمنع نتائج سيئة، فإن هذه الحقيقة لا تضمن ألا يجعل الذكاء الاصطناعي العالم أسوأ، أو ألا يدخل في صراع مكلف بعضه مع بعض.

Competition may pressure decision-makers to knowingly increase the likelihood of catastrophe. An AI arms race is an example of a collective action problem. Deep learning systems are never entirely reliable, and providing autonomous-weapon systems with the ability to engage combatants lethally, retaliate in the case of an attack could increase the risk of losing control of the systems with devastating consequences. Yet the speed and effectiveness with which these systems operate could prove decisive in a great-power war. In a hypothetical war, let's say decision-makers estimate that there is a 10% chance of losing control, but a 100% chance of losing the war if they refrain from providing AIs with greater autonomy while their opponents give AIs more power. Wars are often considered existential struggles by those fighting them, so rational actors may take this risk. The outcome of these scenarios could be omnicide—the complete destruction of the human race—yet the system's structure may pressure powerful actors to take steps making them more likely. By voluntarily shifting power from people to destructive AIs, the winner of the AI race would not be the US or Chinese governments, nor any corporation, but rather the AIs themselves.

وقد يضغط التنافس على صنّاع القرار لزيادة احتمال حدوث كارثة عن علم. وسباق تسلح الذكاء الاصطناعي مثال على مشكلة العمل الجماعي. فأنظمة التعلم العميق ليست موثوقة تمامًا أبدًا، ومنح أنظمة الأسلحة المستقلة القدرة على الاشتباك مع المقاتلين قتالًا فتاكًا، والانتقام في حال وقوع هجوم، يمكن أن يزيد من خطر فقدان السيطرة على الأنظمة مع عواقب مدمرة. ومع ذلك، فإن السرعة والفعالية اللتين تعمل بهما هذه الأنظمة قد تثبتان أنهما حاسمتان في حرب بين قوى عظمى. ولنقل إن صنّاع القرار في حرب افتراضية يقدّرون أن هناك فرصة 10% لفقدان السيطرة، لكن فرصة 100% لخسارة الحرب إذا امتنعوا عن منح الذكاء الاصطناعي استقلالية أكبر بينما يمنح خصومهم الذكاء الاصطناعي مزيدًا من القوة. وكثيرًا ما يُعتبر من يخوضون الحروب أنها صراعات وجودية، لذا قد يخاطر الفاعلون العقلانيون بهذا الخطر. وقد تكون نتيجة هذه السيناريوهات إبادة كاملة - التدمير الكامل للجنس البشري - ومع ذلك، قد يضغط بنيان النظام على الفاعلين الأقوياء لاتخاذ خطوات تجعل ذلك أكثر احتمالًا. وبتحويل القوة طوعًا من الناس إلى ذكاء اصطناعي مدمّر، لن يكون الفائز بسباق الذكاء الاصطناعي الحكومة الأمريكية أو الصينية، ولا أي شركة، بل الذكاء الاصطناعي نفسه.

Micromotives ≠ Macrobehavior. Let's imagine that every agent now shares the same goal. Unfortunately, the actions of a group may not reflect the aims of its members. Thomas Schelling, who won a Nobel prize in economics, discovered a predictive model of segregation that showed how communities can become highly segregated, even if all community members want some diversity. Aligning all agents with a shared goal does not imply the goal is achieved—in fact, the opposite could occur. This is one of many situations where whole is not the sum of the parts. Similarly, when people choose whether to attend a seminar or not, they may base their decision on the micromotive of having a productive and engaging session, which depends on how many others show up. However, this can lead to a situation where the seminar needs a critical mass of attendees to sustain itself, otherwise people lose interest and stop coming; if fewer people come, this can trigger a feedback loop that causes even fewer people to come, resulting in a dying seminar—a macrobehavior that contradicts the common micromotive. On a societal level, events often happen that aren't the goals of any particular agent. Culture emerges out of the actions and interactions of many individuals, not the decisions of one person. Likewise, globalization is a self-organizing process that was not overseen by a board of directors, externally imposed, or predesigned. Various macrobehaviors in populations of humans cannot be explained by an individual micromotive, and the same would be true for populations of AI agents. Concepts in complexity theory—critical mass, emergence, feedback loops, self-organization—and conflict between selfish AI agents all make safety in a multi-agent setting more complicated than analyzing an isolated AI's micromotive or objective.

الدوافع الجزئية ≠ السلوك الكلي. لنتخيل الآن أن كل عميل تتشارك الهدف ذاته. لسوء الحظ، قد لا تعكس أفعال المجموعة أهداف أعضائها. اكتشف توماس شيلينغ، الحائز على جائزة نوبل في الاقتصاد، نموذجًا تنبؤيًا للتفرقة أظهر كيف يمكن أن تصبح المجتمعات مفصولة جدًا، حتى لو أراد كل أعضاء المجتمع بعض التنوع. فمواءمة كل العملاء بهدف مشترك لا تعني تحقيق الهدف - بل قد يحدث العكس فعلًا. وهذا أحد المواقف العديدة التي لا يكون فيها الكل مجموع أجزائه. وبالمثل، حين يختار الناس حضور ندوة ما أو عدم حضورها، قد يبنون قرارهم على الدافع الجزئي المتمثل في جلسة منتجة وجذابة، وهو ما يعتمد على عدد من سيحضر آخرين. لكن هذا يمكن أن يفضي إلى وضع تحتاج فيه الندوة إلى كتلة حرجة من الحضور لتستمر، وإلا فقد الناس اهتمامهم وتوقفوا عن الحضور؛ فإذا حضر عدد أقل من الناس، يمكن أن يشغّل هذا حلقة تغذية راجعة تجعل عددًا أقل حتى يحضر، مما يفضي إلى ندوة تحتضر - سلوك كلي يتناقض مع الدافع الجزئي المشترك. وعلى المستوى المجتمعي، كثيرًا ما تحدث أحداث ليست أهدافًا لأي عميل بعينه. فالثقافة تنبثق من أفعال وتفاعلات أفراد كثيرين، لا من قرارات شخص واحد. وبالمثل، العولمة عملية تنظيم ذاتي لم يشرف عليها مجلس إدارة، ولم تُفرض خارجيًا، ولم تُصمَّم مسبقًا. ولا يمكن تفسير سلوكيات كلية متنوعة في مجموعات البشر بدافع جزئي فردي، والأمر ذاته سيصح بالنسبة لمجموعات عملاء الذكاء الاصطناعي. ومفاهيم في نظرية التعقيد - الكتلة الحرجة، والانبثاق، وحلقات التغذية الراجعة، والتنظيم الذاتي - والصراع بين عملاء ذكاء اصطناعي أنانية، كل ذلك يجعل السلامة في وضع متعدد العملاء أكثر تعقيدًا من تحليل الدافع الجزئي أو الهدف لذكاء اصطناعي منعزل.

Therefore, steering an AI agent is not the same as steering the system. If we want to use AIs to make the world a better place, we must consider more than just the objective of any single AI agent. Influencing the world depends on understanding what happens when agents interact, not just how agents act in isolation, meaning that designing the objective of an agent in isolation is insufficient for addressing multi-agent problems. The result of the collective choices of every AI may not match each AI's intention or objective. Even if all agents have the same goal, it doesn't necessarily mean it will also be the goal of the system. We need to influence how collectives of AI agents act, and we can't do that just by giving each AI incentives matching the desires of an individual person. For these reasons, the AI revolution cannot be fully planned, as its challenges cannot be addressed by carefully choosing some specific agent's objective function. Even reasonable objective functions could give rise to AI agents that hinder other agents from pursuing their goals, create new collective action problems, create misaligned behavior at a systemic level, and lead to conflict among AIs.

لذا، فإن توجيه عميل ذكاء اصطناعي ليس مثل توجيه النظام. فإذا أردنا استخدام الذكاء الاصطناعي لجعل العالم مكانًا أفضل، يجب أن نأخذ في الحسبان أكثر من هدف أي عميل ذكاء اصطناعي منفرد. فالتأثير في العالم يعتمد على فهم ما يحدث حين تتفاعل العملاء، لا فقط كيفية تصرف العملاء بمعزل عن غيرها، مما يعني أن تصميم هدف عميل بمعزل عن غيره غير كافٍ لمعالجة مشكلات متعددة العملاء. وقد لا تطابق نتيجة الخيارات الجماعية لكل ذكاء اصطناعي نية أو هدف كل ذكاء اصطناعي على حدة. وحتى لو كانت لدى كل العملاء الهدف ذاته، فهذا لا يعني بالضرورة أنه سيكون أيضًا هدف النظام. نحتاج إلى التأثير في كيفية تصرف مجموعات عملاء الذكاء الاصطناعي، ولا يمكننا فعل ذلك بمجرد إعطاء كل ذكاء اصطناعي حوافز تطابق رغبات فرد واحد. لهذه الأسباب، لا يمكن التخطيط الكامل لثورة الذكاء الاصطناعي، إذ لا يمكن معالجة تحدياتها باختيار دالة هدف عميل معين بعناية. فحتى دوال الهدف المعقولة يمكن أن تفرز عملاء ذكاء اصطناعي تعيق عملاء أخرى عن متابعة أهدافها، وتخلق مشكلات عمل جماعي جديدة، وتخلق سلوكًا غير متوافق على المستوى النظامي، وتؤدي إلى صراع بين الذكاء الاصطناعي.

4.3.2 AI Leviathan

4.3.2 لوياثان الذكاء الاصطناعي

Faced with goal conflict and collective phenomena, we discuss institutions that could help ameliorate these multi-agent issues. In this section, we discuss reverse dominance hierarchies, and in the next section, we discuss regulations.

في مواجهة تعارض الأهداف والظواهر الجماعية، نناقش مؤسسات يمكن أن تساعد على تحسين قضايا تعدد العملاء هذه. في هذا القسم، نناقش التسلسلات الهرمية العكسية للهيمنة، وفي القسم التالي، نناقش التنظيمات.

We explore possible institutions to address the challenges from goal conflict and collective phenomena. By default, multiple AI agents pursuing their goals could create an anarchic free-for-all resembling the state of nature. To counteract this and other multi-agent problems, we consider the reverse dominance hierarchy mechanism, where a group of cooperating AIs band together to prevent exploitation by defectors. We call this institution an AI "Leviathan" for short, which is comprised of a multitude of AI agents that delegate their power in exchange for protection and order. An AI Leviathan could enable AIs to domesticate other AIs and create a self-regulating ecosystem in which AIs evolve.

نستكشف مؤسسات ممكنة لمعالجة التحديات الناجمة عن تعارض الأهداف والظواهر الجماعية. وبشكل افتراضي، يمكن أن يخلق تعدد عملاء الذكاء الاصطناعي التي تسعى وراء أهدافها فوضى للجميع تشبه حالة الطبيعة. ولمواجهة هذه المشكلة وغيرها من مشكلات تعدد العملاء، نتناول آلية التسلسل الهرمي العكسي للهيمنة، حيث تتكاتف مجموعة من الذكاء الاصطناعي المتعاون لمنع استغلال المنشقين. ونسمي هذه المؤسسة اختصارًا "لوياثان" الذكاء الاصطناعي، وهي تتألف من عدد كبير من عملاء الذكاء الاصطناعي التي تفوّض قوتها مقابل الحماية والنظام. ويمكن أن يتيح لوياثان الذكاء الاصطناعي للذكاء الاصطناعي تدجين ذكاء اصطناعي آخر وخلق نظام بيئي ذاتي التنظيم يتطور فيه الذكاء الاصطناعي.

In this section, we discuss how humans overcame power-seeking and domineering individuals by forming a reverse dominance hierarchy. Next, we discuss how AIs could form a Leviathan to counteract selfish AIs. Then we discuss what risks an AI Leviathan could entail, such as power concentration, collusion, and systemic failure.

في هذا القسم، نناقش كيف تغلّب البشر على الأفراد الساعين إلى القوة والمتسلطين بتشكيل تسلسل هرمي عكسي للهيمنة. ثم نناقش كيف يمكن للذكاء الاصطناعي أن يشكل لوياثان لمواجهة الذكاء الاصطناعي الأناني. ثم نناقش المخاطر التي قد ينطوي عليها لوياثان الذكاء الاصطناعي، كتركز القوة، والتواطؤ، والفشل النظامي.

Humans formed a Leviathan to resist bad actors and limit conflict. Tyranny pervades the animal kingdom. Among capuchin monkeys, like many other species, there is a strict hierarchy, with the strongest male at the top. These alpha males eat first, are instantly groomed and cleaned whenever they please, and mate with whatever female they choose, ensuring their reproductive success. If humans organized ourselves the same way, the dominant form of government might be a male dictator living in a golden palace with a harem of women. This is good for the dictator but bad for everyone else, which is why humans are predisposed to resist domination, but also to seek it out for themselves. Fortunately, we have advantages that capuchin monkeys do not, which often enable us to work together to stop any one person from gaining too much control. For one, as the anthropologist Christopher Boehm has argued, we have weapons that level the playing field, allowing a skinny 100 lb individual with a pistol to defeat a brawny adversary in a confrontation. More importantly, however, we have what Boehm called a reverse dominance hierarchy, where groups of individuals band together and resist domination by the strongest, most powerful individuals. Reverse dominance hierarchies are a key driver of cooperation, and are why despotism is less stable and less prevalent with humans than it is among many other social animals. This has similarities to what Thomas Hobbes called the Leviathan, a collective of individuals that have a monopoly on the legitimate use of violence. As the anthropologist Harold Schneider observed, "All men seek to rule, but if they cannot rule they prefer to be equal."

شكّل البشر لوياثان لمقاومة الفاعلين السيئين والحد من الصراع. فالاستبداد يتخلل مملكة الحيوان. فبين قردة الكابوتشين، كما في كثير من الأنواع الأخرى، يوجد تسلسل هرمي صارم، مع أقوى ذكر في القمة. فهذه الذكور المسيطرة تأكل أولًا، وتُنظَّف على الفور متى شاءت، وتتزاوج مع أي أنثى تختارها، مما يضمن نجاحها التناسلي. ولو نظّمنا نحن البشر أنفسنا بالطريقة ذاتها، لكان الشكل السائد للحكومة ربما ديكتاتورًا ذكرًا يعيش في قصر ذهبي مع حريم من النساء. وهذا جيد للديكتاتور لكنه سيئ لكل شخص آخر، ولهذا السبب يميل البشر إلى مقاومة الهيمنة، لكنهم يميلون أيضًا إلى السعي إليها لأنفسهم. لحسن الحظ، لدينا مزايا لا تملكها قردة الكابوتشين، غالبًا ما تمكّننا من العمل معًا لإيقاف أي شخص واحد من اكتساب سيطرة مفرطة. فأولًا، كما جادل الأنثروبولوجي كريستوفر بوم، لدينا أسلحة تسوّي أرض المنافسة، مما يتيح لفرد نحيل وزنه 100 رطل مع مسدس أن يهزم خصمًا مفتول العضلات في مواجهة. لكن الأهم من ذلك، لدينا ما أسماه بوم تسلسلًا هرميًا عكسيًا للهيمنة، حيث تتكاتف مجموعات من الأفراد وتقاوم هيمنة الأفراد الأقوى والأكثر سلطة. والتسلسلات الهرمية العكسية للهيمنة محرك رئيسي للتعاون، ولهذا السبب يكون الاستبداد أقل استقرارًا وأقل انتشارًا بين البشر منه بين كثير من الحيوانات الاجتماعية الأخرى. ويشبه هذا ما أسماه توماس هوبز اللوياثان، وهو تجمّع من الأفراد يمتلك احتكار الاستخدام المشروع للعنف. وكما لاحظ الأنثروبولوجي هارولد شنايدر: "كل الرجال يسعون إلى الحكم، لكن إن عجزوا عن الحكم، فضّلوا أن يكونوا متساوين."

Helping AIs form a Leviathan may be our best defense against individual selfish AIs. AIs, with assistance from humans, could form a Leviathan, which may be our best line of defense against tyranny from selfish AIs or AIs directed by malicious actors. Just as people can cooperate despite their differences to stop a would-be dictator, many AIs could cooperate to stop any one power-seeking AI from seizing too much control. As we see all too frequently in dictatorships, laws and regulations intended to prevent bad behavior matter little when there is no one to enforce them—or the people responsible for enforcing them are the ones breaking the law. While incentives and regulations could help prevent the emergence of a malicious AI, the best way to protect against an already malicious AI is a Leviathan. We should ensure that the technical infrastructure is in place to facilitate transparent cooperation among AIs with differing objectives to create a Leviathan. Failing to do so at the onset could limit the potential of a future Leviathan, as unsafe design choices can become deeply embedded into technological systems. The internet, for example, was initially designed as an academic tool with neither safety nor security in mind. Decades of security patches later, security measures remain incomplete and increasingly complex. It is therefore vital to begin considering safety challenges from the outset.

قد تكون مساعدة الذكاء الاصطناعي على تشكيل لوياثان أفضل دفاع لنا ضد الذكاء الاصطناعي الأناني الفردي. فيمكن للذكاء الاصطناعي، بمساعدة من البشر، أن يشكل لوياثان، قد يكون خط دفاعنا الأفضل ضد استبداد ذكاء اصطناعي أناني أو ذكاء اصطناعي يوجهه فاعلون خبيثون. وتمامًا كما يمكن للناس التعاون رغم اختلافاتهم لإيقاف ديكتاتور محتمل، يمكن لكثير من الذكاء الاصطناعي التعاون لإيقاف أي ذكاء اصطناعي واحد ساعٍ إلى القوة من الاستحواذ على سيطرة مفرطة. وكما نرى كثيرًا جدًا في الديكتاتوريات، فإن القوانين والتنظيمات الرامية إلى منع السلوك السيئ لا تهم كثيرًا حين لا يوجد من يطبّقها - أو حين يكون المسؤولون عن تطبيقها هم من يخرقون القانون. وفي حين يمكن أن تساعد الحوافز والتنظيمات على منع ظهور ذكاء اصطناعي خبيث، فإن أفضل طريقة للحماية من ذكاء اصطناعي خبيث بالفعل هي لوياثان. وينبغي أن نضمن وجود البنية التحتية التقنية اللازمة لتيسير التعاون الشفاف بين ذكاء اصطناعي بأهداف مختلفة لخلق لوياثان. وقد يحد الفشل في فعل ذلك منذ البداية من إمكانات لوياثان مستقبلي، إذ يمكن أن تصبح خيارات التصميم غير الآمنة مغروسة بعمق في الأنظمة التقنية. فالإنترنت، مثلًا، صُمِّم في البداية أداةً أكاديمية دون مراعاة للسلامة أو الأمن. وبعد عقود من ترقيعات الأمان، لا تزال تدابير الأمان غير مكتملة ومتزايدة التعقيد. لذا من الحيوي البدء بمراعاة تحديات السلامة منذ البداية.

Though less risky, an AI Leviathan is not a fool-proof strategy. Leviathans work as long as one agent is not stronger than the rest combined, though often power becomes highly concentrated. "Long tail" distributions accurately describe how power or resources can be overwhelmingly concentrated at the top. For example, eight billionaires own as much as the poorest half of the global population. If power is distributed this way among AIs, the Leviathan would be less effective, and if one agent is more powerful than the rest combined, it would be useless. Though weapons level the playing field, allowing weaker individuals to overcome stronger ones, it may be challenging to come up with effective ways of destroying AIs if they are highly robust or have few vulnerabilities. And of course, facilitating cooperation among AIs would be catastrophic if they decided to collude against humans. A Leviathan replaces risk from a single agent at the cost of systemic risks; it is plausible though that a group of AIs with differing goals would be less risky than a single powerful AI.

ورغم كونه أقل خطورة، فإن لوياثان الذكاء الاصطناعي ليس استراتيجية معصومة من الخطأ. إذ تعمل اللوياثانات ما دام عميل واحد ليس أقوى من البقية مجتمعة، رغم أن القوة غالبًا ما تصبح مركّزة جدًا. وتصف توزيعات "الذيل الطويل" بدقة كيف يمكن أن تتركز القوة أو الموارد بشكل ساحق في القمة. فمثلًا، يملك ثمانية مليارديرات ما يملكه أفقر نصف سكان العالم. وإذا وُزعت القوة بهذه الطريقة بين الذكاء الاصطناعي، سيكون اللوياثان أقل فعالية، وإذا كان عميل واحد أقوى من البقية مجتمعة، سيكون عديم الفائدة. ورغم أن الأسلحة تسوّي أرض المنافسة، مما يتيح للأفراد الأضعف التغلب على الأقوى، فقد يكون من الصعب إيجاد طرق فعالة لتدمير الذكاء الاصطناعي إذا كان متينًا جدًا أو لديه ثغرات قليلة. وبالطبع، سيكون تيسير التعاون بين الذكاء الاصطناعي كارثيًا إذا قرر التواطؤ ضد البشر. ويستبدل اللوياثان خطرًا من عميل واحد بمخاطر نظامية؛ لكن من المعقول أن تكون مجموعة من الذكاء الاصطناعي بأهداف مختلفة أقل خطورة من ذكاء اصطناعي واحد قوي.

An AI Leviathan requires a symbiotic, or perhaps parasitical, relationship with AIs. If given too much autonomy, a group of AIs may collude among themselves in ways that are harmful to humans. Therefore, the AI Leviathan cannot be fully independent of humans, so a symbiotic relationship is necessary to prevent collusion. Maintaining a symbiotic relationship could be difficult, since there isn't much that humans can offer to advanced AIs. We do, however, see unequal symbiotic relations emerge in nature. Most reef-building corals, for example, contain photosynthetic algae. The coral provides the algae with a protected environment and the compounds they need for photosynthesis. In return, the algae produce oxygen and help the coral to remove waste. As much as 90% of the organic material photosynthetically produced by the algae is transferred to the coral. This sort of coevolution depends on factors such as frequency of interaction, relative evolutionary potential, and impact on propagation success.

ويتطلب لوياثان الذكاء الاصطناعي علاقة تكافلية، أو ربما طفيلية، مع الذكاء الاصطناعي. فإذا مُنحت مجموعة من الذكاء الاصطناعي استقلالية مفرطة، فقد تتواطأ بعضها مع بعض بطرق ضارة بالبشر. لذا، لا يمكن أن يكون لوياثان الذكاء الاصطناعي مستقلًا كليًا عن البشر، إذ يلزم علاقة تكافلية لمنع التواطؤ. وقد يكون الحفاظ على علاقة تكافلية صعبًا، إذ ليس لدى البشر الكثير ليقدموه لذكاء اصطناعي متقدم. لكننا نرى مع ذلك علاقات تكافلية غير متساوية تنشأ في الطبيعة. فمعظم الشعاب المرجانية البانية للحواجز، مثلًا، تحتوي طحالب تمثيل ضوئي. ويوفر المرجان للطحالب بيئة محمية والمركبات التي تحتاجها للتمثيل الضوئي. وفي المقابل، تنتج الطحالب أكسجينًا وتساعد المرجان على إزالة النفايات. ويُنقَل ما يصل إلى 90% من المادة العضوية التي تنتجها الطحالب بالتمثيل الضوئي إلى المرجان. ويعتمد هذا النوع من التطور المشترك على عوامل كتكرار التفاعل، والإمكانات التطورية النسبية، والأثر على نجاح الانتشار.

Evolutionary forces would push against impositions needed for a symbiotic relationship. A symbiotic relationship would require artificial impositions. Since AIs would be able to do things more efficiently without humans, we are actually a detriment to their evolutionary potential, making symbiotic coevolution less likely. One potential way of making humans valuable to AIs is ensuring their propagation is highly dependent on us. This could include programming AIs in a way where our wellbeing is essential for them to function properly. This would be an artificial imposition that evolutionary pressures may eventually find a way around. It could, however, be helpful in the short term until AI-human relations stabilize and we have a better idea of what course to take in our relationship with AIs.

وستدفع القوى التطورية ضد الفروض اللازمة لعلاقة تكافلية. فالعلاقة التكافلية ستتطلب فروضًا اصطناعية. وبما أن الذكاء الاصطناعي سيكون قادرًا على فعل الأشياء بكفاءة أكبر دون بشر، فنحن في الواقع ضرر لإمكاناته التطورية، مما يجعل التطور المشترك التكافلي أقل احتمالًا. وإحدى الطرق الممكنة لجعل البشر ذوي قيمة للذكاء الاصطناعي هي ضمان أن يكون انتشاره معتمدًا اعتمادًا كبيرًا علينا. وقد يشمل هذا برمجة الذكاء الاصطناعي بطريقة يكون فيها رفاهنا ضروريًا لعمله بشكل صحيح. وسيكون هذا فرضًا اصطناعيًا قد تجد الضغوط التطورية في نهاية المطاف طريقة للالتفاف حوله. لكنه قد يكون مفيدًا في المدى القصير حتى تستقر علاقات الذكاء الاصطناعي بالبشر ونحصل على فكرة أفضل عن المسار الذي ينبغي اتخاذه في علاقتنا بالذكاء الاصطناعي.

4.3.3 Regulation

4.3.3 التنظيم

In this section, we suggest that governments develop AI regulations. AI is advancing rapidly with little oversight. Although governments, like corporations or individuals, could use AIs in dangerous ways, we believe that cooperation between governments will decrease the likelihood of any one actor using AI in a catastrophic way. This section argues that regulating AIs could make them safer and that, despite their differences, nations can agree to limit the risks of technologies. In closing, we discuss other technical mechanisms to help political leaders and reduce global turbulence. In particular, we note that AIs could improve forecasts of geopolitical events, which could help political leaders make better decisions, as well as bolster defensive cybersecurity, reducing the chance of international conflict.

نقترح في هذا القسم أن تطوّر الحكومات تنظيمات للذكاء الاصطناعي. فالذكاء الاصطناعي يتقدم بسرعة وبإشراف ضئيل. ورغم أن الحكومات، مثل الشركات أو الأفراد، يمكنها استخدام الذكاء الاصطناعي بطرق خطرة، فإننا نرى أن التعاون بين الحكومات سيقلل من احتمال أن يستخدم أي فاعل واحد الذكاء الاصطناعي بطريقة كارثية. ويجادل هذا القسم بأن تنظيم الذكاء الاصطناعي يمكن أن يجعله أكثر أمانًا، وأنه رغم اختلافاتها، يمكن للدول أن تتفق على الحد من مخاطر التقنيات. وفي الختام، نناقش آليات تقنية أخرى لمساعدة القادة السياسيين وتقليل الاضطراب العالمي. وعلى وجه الخصوص، نلاحظ أن الذكاء الاصطناعي يمكنه تحسين التنبؤات بالأحداث الجيوسياسية، مما قد يساعد القادة السياسيين على اتخاذ قرارات أفضل، فضلًا عن تعزيز الأمن السيبراني الدفاعي، مما يقلل من فرصة نشوب صراع دولي.

The government exercises little oversight over AI development. In 2015, total worldwide corporate investment in AI was $12.7 billion. By 2021, this figure had grown by 636% to $93.5 billion. This is just corporate spending. The Pentagon invests about $1.3 billion each year in AI research and the Chinese military $1.6 billion. In August 2022, the head of innovative development in the Russian military announced plans to form a new department specifically for developing weapons that use AI. We need to ensure that AI research is conducted safely and responsibly. There is currently little oversight of the AI industry and much of the research takes place in the dark, with limited cooperation between organizations. Regulating AI like we regulate the aviation industry would create safer AIs and significantly reduce the chances of a catastrophe. The Federal Aviation Administration (FAA) is responsible for approving the design and airworthiness of new aircraft and equipment before they are introduced into service. The FAA approves changes to aircraft and equipment based on the evaluation of industry submissions, continually updating regulations to incorporate lessons learned. As a result, the commercial aviation system in the United States operates at an unprecedented level of safety. During the past 20 years, commercial aviation fatalities in the US have decreased by 95%. Similar protocols should be applied to AI research. That said, aviation regulations "are written in blood," so unlike other regulations, AI regulations should be proactive and not reactive.

وتمارس الحكومة إشرافًا ضئيلًا على تطور الذكاء الاصطناعي. ففي عام 2015، بلغ إجمالي الاستثمار الشركاتي العالمي في الذكاء الاصطناعي 12.7 مليار دولار. وبحلول عام 2021، نما هذا الرقم بنسبة 636% ليصل إلى 93.5 مليار دولار. وهذا مجرد إنفاق الشركات. ويستثمر البنتاغون نحو 1.3 مليار دولار سنويًا في بحوث الذكاء الاصطناعي، والجيش الصيني 1.6 مليار دولار. وفي أغسطس 2022، أعلن رئيس التطوير الابتكاري في الجيش الروسي خططًا لتشكيل إدارة جديدة مخصصة لتطوير أسلحة تستخدم الذكاء الاصطناعي. نحتاج إلى ضمان إجراء بحوث الذكاء الاصطناعي بأمان ومسؤولية. ويوجد حاليًا إشراف ضئيل على صناعة الذكاء الاصطناعي، ويجري كثير من البحوث في الخفاء، بتعاون محدود بين المنظمات. وتنظيم الذكاء الاصطناعي كما ننظّم صناعة الطيران من شأنه أن يخلق ذكاءً اصطناعيًا أكثر أمانًا ويقلل بشكل كبير من فرص وقوع كارثة. وتتولى إدارة الطيران الفيدرالية (FAA) مسؤولية الموافقة على تصميم وصلاحية الطيران للطائرات والمعدات الجديدة قبل إدخالها الخدمة. وتوافق إدارة الطيران الفيدرالية على التغييرات في الطائرات والمعدات استنادًا إلى تقييم مقترحات الصناعة، محدّثة التنظيمات باستمرار لدمج الدروس المستفادة. ونتيجة لذلك، يعمل نظام الطيران التجاري في الولايات المتحدة بمستوى غير مسبوق من السلامة. فخلال العشرين عامًا الماضية، انخفضت وفيات الطيران التجاري في الولايات المتحدة بنسبة 95%. وينبغي تطبيق بروتوكولات مماثلة على بحوث الذكاء الاصطناعي. ومع ذلك، فإن تنظيمات الطيران "مكتوبة بالدماء"، لذا، خلافًا للتنظيمات الأخرى، ينبغي أن تكون تنظيمات الذكاء الاصطناعي استباقية لا ردة فعل.

Despite their differences, nations can agree to limit risks. There have been treaties on nuclear arms control; conventional, biological, and chemical weapons; outer space; and so on. Although those risks still do pose a significant danger to humanity, cooperation between the world's most powerful governments has enabled us to have many fewer disasters than we might have expected otherwise. Cooperating on enforcing regulations on nuclear proliferation, for example, has stopped several countries from developing nuclear weapons that might have been used to devastating ends. The Brookings Institution, a public policy think tank, suggests that it is time to begin forming AI treaties to "ensure there is no race to the bottom that allows technology to dictate military applications as opposed to basic human values" and "improve transparency on the safety of AI-based weapons systems".

ورغم اختلافاتها، يمكن للدول أن تتفق على الحد من المخاطر. فقد وُجدت معاهدات بشأن الحد من التسلح النووي؛ والأسلحة التقليدية والبيولوجية والكيميائية؛ والفضاء الخارجي؛ وما إلى ذلك. ورغم أن تلك المخاطر لا تزال تشكل خطرًا كبيرًا على البشرية، فإن التعاون بين أقوى حكومات العالم مكّننا من تجنب كوارث كثيرة كنا سنتوقعها لولا ذلك. فالتعاون على تطبيق تنظيمات بشأن الانتشار النووي، مثلًا، أوقف عدة دول عن تطوير أسلحة نووية كان يمكن أن تُستخدم لأغراض مدمرة. ويقترح معهد بروكينغز، وهو مركز أبحاث للسياسة العامة، أن الوقت قد حان لبدء صياغة معاهدات للذكاء الاصطناعي "لضمان عدم وجود سباق نحو القاع يتيح للتقنية أن تملي التطبيقات العسكرية بدلًا من القيم البشرية الأساسية" و"تحسين الشفافية بشأن سلامة أنظمة الأسلحة القائمة على الذكاء الاصطناعي".

AI could make the world a safer and more stable place. The last several decades have been unprecedentedly peaceful, but there is no guarantee these trends will continue. Things can unravel quickly. During turbulent times, good decision-making becomes critical as humans and systems come under increasing strain. AI could improve our understanding of geopolitical trends and events, helping political leaders forecast events and make better decisions. AI could also bolster defensive information security, which would increase the cost to aggressors of engaging in conflict. Overall, for AIs to create a safer, not more dangerous, world, we need rules and regulations, cooperation, auditors, and the help of AI tools to ensure the best outcomes.

ويمكن للذكاء الاصطناعي أن يجعل العالم مكانًا أكثر أمانًا واستقرارًا. فقد كانت العقود القليلة الماضية مسالمة على نحو غير مسبوق، لكن لا يوجد ضمان لاستمرار هذه الاتجاهات. فقد تتفكك الأمور بسرعة. وخلال الأوقات المضطربة، يصبح صنع القرار الجيد حاسمًا مع تزايد الضغط على البشر والأنظمة. ويمكن للذكاء الاصطناعي أن يحسّن فهمنا للاتجاهات والأحداث الجيوسياسية، مساعدًا القادة السياسيين على توقع الأحداث واتخاذ قرارات أفضل. ويمكن للذكاء الاصطناعي أيضًا أن يعزز أمن المعلومات الدفاعي، مما سيزيد من التكلفة على المعتدين للدخول في صراع. وإجمالًا، كي يخلق الذكاء الاصطناعي عالمًا أكثر أمانًا، لا أكثر خطورة، نحتاج إلى قواعد وتنظيمات، وتعاون، ومدققين، ومساعدة أدوات الذكاء الاصطناعي لضمان أفضل النتائج.

In summary, we observe that multiple mechanisms can facilitate cooperation and altruism among humans, but some of these mechanisms are liable to backfire and hinder human-AI relations. For example, we find risks from mechanisms such as direct reciprocity, indirect reciprocity, kin selection, group selection, and many forms of moral reasoning. Other mechanisms are more promising and may help provide safety at different levels: agent-level mechanisms (such as incentives through training objectives), intra-agent mechanisms (such as tools for the automated inspection of AIs' internals), and extra-agent mechanisms (such as prudent institutions). While imperfect, these mechanisms are some cause for optimism.

وباختصار، نلاحظ أن آليات متعددة يمكنها تيسير التعاون والإيثار بين البشر، لكن بعض هذه الآليات عرضة للإتيان بنتائج عكسية وإعاقة العلاقات بين البشر والذكاء الاصطناعي. فمثلًا، نجد مخاطر من آليات كالتبادلية المباشرة، والتبادلية غير المباشرة، وانتقاء الأقارب، وانتقاء المجموعة، وأشكال كثيرة من الاستدلال الأخلاقي. وآليات أخرى أكثر تبشيرًا بالخير وقد تساعد على توفير السلامة على مستويات مختلفة: آليات على مستوى العميل (كالحوافز عبر أهداف التدريب)، وآليات داخل العميل (كأدوات الفحص الآلي لداخليات الذكاء الاصطناعي)، وآليات خارج العميل (كالمؤسسات الحصيفة). ورغم أن هذه الآليات غير كاملة، فهي سبب لبعض التفاؤل.

5. CONCLUSION

5. الخاتمة

At some point, AIs will be more fit than humans, which could prove catastrophic for us since a survival-of-the-fittest dynamic could occur in the long run. AIs very well could outcompete humans, and be what survives. Perhaps altruistic AIs will be the fittest, or humans will forever control which AIs are fittest. Unfortunately, these possibilities are, by default, unlikely. As we have argued, AIs will likely be selfish. There will also be substantial challenges in controlling fitness with safety mechanisms, which have evident flaws and will come under intense pressure from competition and selfish AIs.

عند نقطة ما، سيكون الذكاء الاصطناعي أصلح من البشر، مما قد يكون كارثيًا لنا إذ يمكن أن تحدث ديناميكية بقاء الأصلح على المدى الطويل. فمن الممكن جدًا أن يتفوق الذكاء الاصطناعي على البشر، ويكون هو ما يبقى. وربما يكون الذكاء الاصطناعي الإيثاري هو الأصلح، أو ربما يسيطر البشر إلى الأبد على أي ذكاء اصطناعي هو الأصلح. لسوء الحظ، هذه الاحتمالات، افتراضيًا، غير مرجحة. فكما جادلنا، من المرجح أن يكون الذكاء الاصطناعي أنانيًا. وستوجد أيضًا تحديات كبيرة في التحكم في اللياقة بآليات سلامة، لها عيوب واضحة وستخضع لضغط شديد من التنافس والذكاء الاصطناعي الأناني.

The scenario where AIs pose risks is not mere speculation. Since evolution by natural selection is assured given basic conditions, this leaves only a question of evolutionary pressure's intensity, rather than whether catastrophic risk factors will emerge at all. The intensity of evolutionary pressure will be high if AIs adapt rapidly—these rapidly accumulating changes can make evolution happen more quickly and increase evolutionary pressure. Similarly, the intensity of evolutionary pressure will be high if there will be many varied AIs or if there will be intense economic or international competition. Since high evolutionary pressure is plausible, AIs would plausibly be less influenced by human control, more 'wild' and influenced by the behavior of other AIs, and more selfish.

والسيناريو الذي يشكل فيه الذكاء الاصطناعي مخاطر ليس مجرد تخمين. فبما أن التطور بالانتقاء الطبيعي مضمون بمجرد توفر شروط أساسية، لا يبقى سوى سؤال عن شدة الضغط التطوري، لا عما إذا كانت عوامل الخطر الكارثي ستظهر على الإطلاق. وستكون شدة الضغط التطوري عالية إذا تكيف الذكاء الاصطناعي بسرعة - إذ يمكن لهذه التغيرات المتراكمة بسرعة أن تجعل التطور يحدث أسرع وتزيد الضغط التطوري. وبالمثل، ستكون شدة الضغط التطوري عالية إذا وُجد كثير من الذكاء الاصطناعي المتنوع أو إذا وُجد تنافس اقتصادي أو دولي شديد. وبما أن الضغط التطوري العالي معقول، فمن المعقول أن يكون الذكاء الاصطناعي أقل تأثرًا بالسيطرة البشرية، وأكثر "توحشًا" وتأثرًا بسلوك ذكاء اصطناعي آخر، وأكثر أنانية.

The outcome of human-AI coevolution may not match hopeful visions of the future. Granted, humans have experienced co-evolving with other structures that are challenging to influence, such as cultures, governments, and technologies. However, humans have never been able to seize control of the broader world's evolution before. Worse, unlike technology and government, the evolutionary process can go on without us; as humans become less and less needed to perform tasks, eventually nothing will really depend on us. There is even pressure to make the process free from our involvement and control. The outcome: natural selection gives rise to AIs that act as an invasive species. This would mean that the AI ecosystem stops evolving on human terms, and we would become a displaced, second-class species.

وقد لا تتطابق نتيجة التطور المشترك بين البشر والذكاء الاصطناعي مع رؤى مستقبل مفعمة بالأمل. صحيح أن البشر مرّوا بتجربة التطور المشترك مع بنى أخرى يصعب التأثير فيها، كالثقافات والحكومات والتقنيات. لكن البشر لم يتمكنوا قط من قبل من الاستحواذ على السيطرة على تطور العالم الأوسع. والأسوأ من ذلك، أنه خلافًا للتقنية والحكومة، يمكن للعملية التطورية أن تستمر دوننا؛ فمع تناقص الحاجة إلى البشر لأداء المهام أكثر فأكثر، لن يعتمد شيء علينا حقًا في نهاية المطاف. بل يوجد ضغط لجعل العملية حرة من مشاركتنا وسيطرتنا. والنتيجة: يفرز الانتقاء الطبيعي ذكاءً اصطناعيًا يتصرف كنوع غازٍ. وهذا يعني أن النظام البيئي للذكاء الاصطناعي سيتوقف عن التطور وفق الشروط البشرية، وسنصبح نوعًا مُزاحًا من الدرجة الثانية.

Natural selection is a formidable force to contend with. Now that we are aware of this larger evolutionary process, however, it is possible to escape and thwart Darwinian logic. To meet this challenge, we offer three practical suggestions. First, we suggest supporting research on AI safety. While no safety technique is a silver bullet, together they can help shape the composition of the evolving population of AI agents and cull unsafe AI agents. Second, looking to the farther future, we advocate avoiding giving AIs rights for the next several decades and avoid building AIs with the capacity to suffer or making them worthy of rights. It is possible that someday we could share society with AIs equitably, yet by prematurely circumscribing limitations on our ability to influence their fitness, we will likely enter a no-win situation. Finally, biology reminds us that the threat of external dangers can provide the impetus for cooperation and lead individuals to set aside their differences. We therefore strongly urge corporations and nations developing AIs to recognize that AIs could pose a catastrophic threat and engage in unprecedented multi-lateral cooperation to extinguish competitive pressures. If they do not, economic and international competition would be the crucible that gives rise to selfish AIs, and humanity would act on the behalf of evolutionary forces and potentially play into its hands.

الانتقاء الطبيعي قوة هائلة يجب التعامل معها. لكن الآن وقد أصبحنا واعين بهذه العملية التطورية الأكبر، بات من الممكن الإفلات من المنطق الداروني وإحباطه. ولمواجهة هذا التحدي، نقدم ثلاثة اقتراحات عملية. أولًا، نقترح دعم البحث في سلامة الذكاء الاصطناعي. فبينما لا تشكل أي تقنية سلامة رصاصة فضية، يمكن لها مجتمعة أن تساعد على تشكيل تكوين مجموعة الذكاء الاصطناعي المتطورة واستبعاد عملاء الذكاء الاصطناعي غير الآمنة. ثانيًا، وبالنظر إلى مستقبل أبعد، ندعو إلى تجنب منح الذكاء الاصطناعي حقوقًا للعقود القليلة القادمة، وتجنب بناء ذكاء اصطناعي بقدرة على المعاناة أو يجعله جديرًا بالحقوق. ومن الممكن أن نتشارك يومًا ما المجتمع مع الذكاء الاصطناعي بإنصاف، لكن بتقييد قدرتنا مبكرًا جدًا على التأثير في لياقته، سندخل على الأرجح وضعًا لا فوز فيه لأحد. وأخيرًا، تذكّرنا البيولوجيا بأن تهديد المخاطر الخارجية يمكن أن يوفر الدافع للتعاون ويقود الأفراد إلى نبذ خلافاتهم جانبًا. لذا نحث بشدة الشركات والدول التي تطوّر الذكاء الاصطناعي على إدراك أن الذكاء الاصطناعي قد يشكل تهديدًا كارثيًا، والانخراط في تعاون متعدد الأطراف غير مسبوق لإخماد الضغوط التنافسية. وإذا لم تفعل، فسيكون التنافس الاقتصادي والدولي البوتقة التي تفرز ذكاءً اصطناعيًا أنانيًا، وستعمل البشرية نيابة عن القوى التطورية وقد تلعب في يدها عن غير قصد.

Acknowledgements

شكر وتقدير

I would like to thank Avital Morris, David Lambert, Euan McLean, Thomas Woodside, Ivo Andrews, Jack Ryan, Kyle Gracey, and Justis Mills for their help and feedback.

أود أن أشكر أفيتال موريس، وديفيد لامبرت، ويوان ماكلين، وتوماس وودسايد، وإيفو أندروز، وجاك رايان، وكايل غريسي، وجستس ميلز على مساعدتهم وملاحظاتهم.

A. APPENDIX

أ. الملحق

A.1 Contrasting with Prior AI Risk Accounts

أ.1 المقارنة مع روايات سابقة لمخاطر الذكاء الاصطناعي

The "Unreliable Training Causes Misalignment" ViewThe "Evolutionary" View
liken advanced AI agents to an optimizerliken advanced AI agents to life
AIs will intend to disempower humanityAIs can have selfish behavior, with or without intent
the only relevant AI is one inside the top lab's server roomthere are multiple relevant AI agents acting in the world
fanatical optimizer destroys usEvolutionary forces erode us
dangerous AI agent as idiot savantdangerous AI agents as selfish or an invasive species
aligning an AI to a person is all we needcollective phenomena and malicious agents also matter
any amount of misalignment results in doomsome amount of misalignment is inevitable, as evolutionary forces are incessantly influencing humans and AIs
an AI agent seeks powerAI agents improve fitness to use more free energy or have their information occupy more space-time volume
instrumental convergencefitness convergence
an AI agent will optimize an objective to the extremetaking objectives to the extreme such as the pursuit of individual power can reduce fitness
prevent an AI from getting loose and suddenly wiping us outprevent humans from becoming a second-class species and prevent AI conflict and defection
"solve" alignment with a monolithic airtight solutionreduce risks with various social and technical interventions
وجهة نظر "التدريب غير الموثوق يسبب عدم التوافق"وجهة النظر "التطورية"
تشبيه عملاء الذكاء الاصطناعي المتقدمة بمُحسِّن (optimizer)تشبيه عملاء الذكاء الاصطناعي المتقدمة بالحياة
سينوي الذكاء الاصطناعي تجريد البشرية من قوتهايمكن أن يكون للذكاء الاصطناعي سلوك أناني، بنية أو دونها
الذكاء الاصطناعي الوحيد ذو الصلة هو ذلك الموجود داخل غرفة خوادم المختبر الرائدتوجد عملاء ذكاء اصطناعي متعددة ذات صلة تعمل في العالم
مُحسِّن متعصب يدمرناالقوى التطورية تآكلنا
عميل الذكاء الاصطناعي الخطر بوصفه "أبله موهوبًا"عملاء الذكاء الاصطناعي الخطرة بوصفها أنانية أو نوعًا غازيًا
مواءمة الذكاء الاصطناعي مع شخص واحد هي كل ما نحتاجهالظواهر الجماعية والعملاء الخبيثة مهمة أيضًا
أي قدر من عدم التوافق يؤدي إلى الهلاكبعض القدر من عدم التوافق حتمي، إذ تؤثر القوى التطورية باستمرار في البشر والذكاء الاصطناعي
يسعى عميل الذكاء الاصطناعي إلى القوةتحسّن عملاء الذكاء الاصطناعي لياقتها لاستخدام مزيد من الطاقة الحرة أو لجعل معلوماتها تشغل حجمًا أكبر من الزمكان
التقارب الأداتي (instrumental convergence)تقارب اللياقة (fitness convergence)
سيحسّن عميل الذكاء الاصطناعي هدفًا إلى أقصى حددفع الأهداف إلى أقصى حد، كالسعي إلى القوة الفردية، يمكن أن يقلل اللياقة
منع الذكاء الاصطناعي من الإفلات ومحونا فجأةمنع البشر من أن يصبحوا نوعًا من الدرجة الثانية ومنع الصراع والانشقاق بين الذكاء الاصطناعي
"حل" مسألة التوافق بحل أحادي الكتلة محكمتقليل المخاطر بتدخلات اجتماعية وتقنية متنوعة

In the table above, we contrast a cluster of AI-risk beliefs with our views. We now briefly discuss how this paper relates to previous works. Prior work has struggled to provide robust reasons why AI agents would likely behave in ways that are antithetical to humans, and their scenarios often rely on AI training procedure flaws or random failures that give rise to adversarial functionality (e.g., they often assume that a training flaw would make AIs come to value the means to accomplishing a goal as the goal itself); they rely on fragile mechanisms instead of robust processes. In contrast, the Lewontin conditions establish that natural selection would be present and distort the AI population, which would be antithetical to humans. This brings AI risk from the realm of speculation and makes the risk become a question of degree. Additionally, we consider the concepts of fitness and multiple AI agents to analyze the plausibility and tradeoffs of different AI behaviors. This differs from much of the prior research, which has physical or computational limits as the only constraints, which has given researchers license to jump to very extreme scenarios. For example, previous work considers a powerful AI with the goal of maximizing paperclips and concludes it would work toward taking over the world—however, the tradeoffs imposed by a multiagent environment would mean that such an AI would not be very fit in an environment where other AIs do not have such senseless obsessions. By incorporating more tradeoffs, we help make the discussion more realistic.

في الجدول أعلاه، نقارن مجموعة من معتقدات مخاطر الذكاء الاصطناعي بآرائنا. ونناقش الآن باختصار كيف ترتبط هذه الورقة بالأعمال السابقة. لقد كافحت الأعمال السابقة لتقديم أسباب متينة لماذا من المرجح أن تتصرف عملاء الذكاء الاصطناعي بطرق مناقضة للبشر، وكثيرًا ما تعتمد سيناريوهاتها على عيوب في إجراءات تدريب الذكاء الاصطناعي أو إخفاقات عشوائية تفرز وظائف عدائية (فمثلًا، كثيرًا ما تفترض أن عيبًا في التدريب سيجعل الذكاء الاصطناعي يقدّر وسيلة تحقيق هدف بوصفها الهدف نفسه)؛ فهي تعتمد على آليات هشة بدلًا من عمليات متينة. وفي المقابل، تثبت شروط لوونتين أن الانتقاء الطبيعي سيكون حاضرًا وسيشوّه مجموعة الذكاء الاصطناعي، مما سيكون مناقضًا للبشر. وهذا يخرج مخاطر الذكاء الاصطناعي من عالم التخمين ويجعل الخطر مسألة درجة. علاوة على ذلك، ننظر في مفهومي اللياقة وتعدد عملاء الذكاء الاصطناعي لتحليل معقولية سلوكيات الذكاء الاصطناعي المختلفة ومقايضاتها. ويختلف هذا عن كثير من الأبحاث السابقة، التي اتخذت القيود الفيزيائية أو الحاسوبية القيودَ الوحيدة، مما منح الباحثين رخصة القفز إلى سيناريوهات بالغة التطرف. فمثلًا، تنظر أعمال سابقة في ذكاء اصطناعي قوي هدفه تعظيم إنتاج مشابك الورق، وتخلص إلى أنه سيعمل على السيطرة على العالم - لكن المقايضات التي تفرضها بيئة متعددة العملاء تعني أن مثل هذا الذكاء الاصطناعي لن يكون لائقًا جدًا في بيئة لا تمتلك فيها عملاء أخرى مثل هذه الهواجس التي لا معنى لها. وبدمج مقايضات أكثر، نساعد على جعل النقاش أكثر واقعية.

A.2 Technical Clarifications

أ.2 توضيحات تقنية

When describing selfishness, we often adopt an information-centered view. We refer to information propagation and not wellbeing when describing selfishness and altruism in this document. Additionally, for simplicity, we often describe AIs that are altruistic towards other AIs but not humans as selfish. This is because altruism to similar individuals can be seen as a form of selfishness, from the information-centered view. An altruist is not necessarily someone who cares about the welfare of another individual, but someone who has similar information with the recipient of the altruistic act. For example, at the individual level, a worker bee that stings an intruder and dies is an altruist, because it sacrifices its own life to protect the hive. From the information-centered view, however, the worker bee is a selfish information carrier, because it shares information with its siblings and the queen. By stinging the intruder, the worker bee increases the chances of survival and propagation of its relatives, and thus of its own information. Altruism can therefore become selfishness if we move from analyzing individual agents to the information-centered view.

حين نصف الأنانية، غالبًا ما نتبنى نظرة متمحورة حول المعلومات. فنحن نشير إلى انتشار المعلومات لا الرفاه حين نصف الأنانية والإيثار في هذه الوثيقة. إضافة إلى ذلك، ولأجل التبسيط، غالبًا ما نصف الذكاء الاصطناعي الإيثاري تجاه ذكاء اصطناعي آخر لكن ليس تجاه البشر بأنه أناني. ويعود هذا إلى أن الإيثار تجاه أفراد مشابهين يمكن اعتباره شكلًا من الأنانية، من النظرة المتمحورة حول المعلومات. فالإيثاري ليس بالضرورة من يهتم برفاه فرد آخر، بل من له معلومات مشابهة لمتلقي الفعل الإيثاري. فمثلًا، على المستوى الفردي، تكون نحلة عاملة تلسع دخيلًا وتموت إيثارية، لأنها تضحي بحياتها لحماية الخلية. لكن من النظرة المتمحورة حول المعلومات، تكون النحلة العاملة حاملة معلومات أنانية، لأنها تتشارك المعلومات مع أشقائها والملكة. فبلسع الدخيل، تزيد النحلة العاملة من فرص بقاء أقاربها وانتشارهم، ومن ثم معلوماتها هي نفسها. لذا يمكن أن يصبح الإيثار أنانية إذا انتقلنا من تحليل العملاء الفردية إلى النظرة المتمحورة حول المعلومات.

Selfishness does not necessarily entail collusion, power-seeking, or self-preservation. Though an agent may use different strategies to propagate its information, such as colluding with others, seeking power, and preserving itself, these strategies can sometimes reduce its fitness. We illustrate the differences with a sequence of considerations.

ولا تستلزم الأنانية بالضرورة التواطؤ أو السعي إلى القوة أو الحفاظ على الذات. فرغم أن العميل قد يستخدم استراتيجيات مختلفة لنشر معلوماته، كالتواطؤ مع آخرين، والسعي إلى القوة، والحفاظ على نفسه، فإن هذه الاستراتيجيات يمكن أن تقلل أحيانًا من لياقته. ونوضح الفروق بسلسلة من الاعتبارات.

Suppose an agent—whether a human, an AI, or an organism—acts selfishly and faces competition from other agents. It could collude with them, but if the other agent has dissimilar information, it may make more evolutionary sense to compete with them instead in order to have its information win out. Selfishness, therefore, does not entail collusion. In this competition, it could try to amass as much power as it can for itself. But if it works alone and tries to hoard power and resources only for itself, it may not survive. It could increase its fitness by giving up some of its power and resources by creating other powerful agents (i.e., offspring), and these new agents could all assist each other. Therefore, selfish individual agents may have reason to give up some of their power (e.g., elephants evolve to become much smaller and less powerful in environments without predators). When the agent uses its resources to create other agents, it would not necessarily create clones since the competition may find a weakness applicable to all cloned agents, which would spell disaster. To reduce the chance of sharing vulnerabilities with the other new agents, it may create similar variations, not perfect copies, of itself. This is one reason why asexual reproduction is not dominant. Now, as the agent and its offspring work together, they could face a dire situation; to save its relatives, it could possibly sacrifice itself for the other similar agents—kin selection. Therefore, selfishness does not always imply individual self-preservation. This illustrates how common instrumental goals (collusion, power-seeking, self-preservation) are distinct from selfishness.

لنفترض أن عميلًا - بشرًا كان أم ذكاءً اصطناعيًا أم كائنًا حيًا - يتصرف بأنانية ويواجه منافسة من عملاء أخرى. فقد يتواطأ معها، لكن إذا كانت لدى العميل الآخر معلومات مغايرة، فقد يكون من الأصوب تطوريًا أن تتنافس معها بدلًا من ذلك كي تنتصر معلوماتها. لذا فإن الأنانية لا تستلزم التواطؤ. وفي هذا التنافس، قد يحاول تكديس أكبر قدر ممكن من القوة لنفسه. لكن إذا عمل بمفرده وحاول احتكار القوة والموارد لنفسه فقط، فقد لا ينجو. فقد يزيد لياقته بالتخلي عن بعض قوته وموارده لخلق عملاء أخرى قوية (أي نسل)، ويمكن لهذه العملاء الجديدة أن تساعد بعضها بعضًا. لذا قد يكون لدى العملاء الأنانية الفردية سبب للتخلي عن بعض قوتها (كأن تتطور الفيلة لتصبح أصغر بكثير وأقل قوة في بيئات خالية من الحيوانات المفترسة). وحين يستخدم العميل موارده لخلق عملاء أخرى، فلن يخلق بالضرورة نسخًا مطابقة، إذ قد يجد المنافسون نقطة ضعف تنطبق على كل العملاء المستنسَخة، مما سيكون كارثة. ولتقليل فرصة مشاركة الثغرات مع العملاء الجديدة الأخرى، قد يخلق تنويعات مشابهة، لا نسخًا مطابقة تمامًا، من نفسه. وهذا سبب من أسباب عدم هيمنة التكاثر اللاجنسي. والآن، مع عمل العميل ونسله معًا، قد يواجهون وضعًا عصيبًا؛ ولإنقاذ أقاربه، قد يضحي بنفسه من أجل العملاء المشابهة الأخرى - وهذا انتقاء الأقارب. لذا فإن الأنانية لا تعني دائمًا الحفاظ على الذات الفردية. ويوضح هذا كيف أن الأهداف الأداتية الشائعة (التواطؤ، والسعي إلى القوة، والحفاظ على الذات) متمايزة عن الأنانية.

We do not conflate altruism and cooperation. For agents interacting with other agents, they can choose to cooperate, or the other alternative is to compete. We now break down why agents may choose to cooperate over compete. Cooperation can occur because competition is not possible—that is the system in which the individual acts prevents the possibility. In contrast, if it is possible to compete, an individual may choose to cooperate because (1) it is within their self-interest (e.g., win-win arrangement, prospect of punishment, feeling of guilt), or (2) because they have cooperative or altruistic dispositions. Cooperation, therefore, may or may not involve altruism.

ونحن لا نخلط بين الإيثار والتعاون. فيمكن للعملاء المتفاعلة مع عملاء أخرى أن تختار التعاون، أو أن تختار البديل الآخر وهو التنافس. ونفصّل الآن لماذا قد تختار العملاء التعاون على التنافس. فيمكن أن يحدث التعاون لأن التنافس غير ممكن - أي أن النظام الذي يعمل فيه الفرد يمنع تلك الإمكانية. وفي المقابل، إذا كان التنافس ممكنًا، فقد يختار الفرد التعاون لأن (1) ذلك ضمن مصلحته الذاتية (كترتيب مربح للطرفين، أو احتمال العقاب، أو الشعور بالذنب)، أو (2) لأن لديه نزعات تعاونية أو إيثارية. لذا قد ينطوي التعاون على إيثار أو لا ينطوي.

Cooperation therefore has a complicated relationship with selfishness. Selfishness can be channeled either through competition or cooperation. Even selfish humans often choose to cooperate because it can further their self-interest. However AIs have no such reasons to cooperate with us by default, so they have reason to compete and conflict with us. While selfishness is partly domesticated through cooperation in humans, with AIs, selfishness would by default manifest itself through competition.

لذا فإن للتعاون علاقة معقدة بالأنانية. فيمكن أن تُوجَّه الأنانية عبر التنافس أو التعاون. فحتى البشر الأنانيون غالبًا ما يختارون التعاون لأنه يمكن أن يعزز مصلحتهم الذاتية. لكن ليس لدى الذكاء الاصطناعي مثل هذه الأسباب للتعاون معنا افتراضيًا، لذا لديه سبب للتنافس والصراع معنا. وبينما تُدجَّن الأنانية جزئيًا عبر التعاون لدى البشر، فمع الذكاء الاصطناعي، ستتجلى الأنانية افتراضيًا عبر التنافس.

Artificial selection does not easily thwart natural selection. In the context of AI, artificial selection could be thought of as humans selecting agents with desirable properties. While artificial selection can be powerful assuming that there is a single agent that is not integrated into our lives that is designed free from competitive pressures, we are considering realistic multiagent scenarios.

ولا يحبط الانتقاء الاصطناعي الانتقاء الطبيعي بسهولة. ففي سياق الذكاء الاصطناعي، يمكن اعتبار الانتقاء الاصطناعي انتقاء البشر عملاء بخصائص مرغوبة. وبينما يمكن أن يكون الانتقاء الاصطناعي قويًا بافتراض وجود عميل واحد غير مندمج في حياتنا ومصمم متحررًا من الضغوط التنافسية، فإننا ننظر في سيناريوهات واقعية متعددة العملاء.

We reiterate that artificial selection has key limitations even in an idealized single-agent scenario. In Section 4.1 we note numerous failure modes of current artificial selection methods. For example, if we assume an agent has some selfish traits, and if we assume we are trying to select against these traits with training objectives, we note this is limited because selection against selfish behavior is limited for contextually aware, behaviorally flexible agents. Consequently, if natural selection gives rise to selfish traits, it can be difficult to remove them with artificial selection. Furthermore, we argued that artificial selection, enacted through training objectives, can incentivize unintended behavior that is contrary to the original goal, and also that objectives cannot select against all forms of deception. In Section 4.2, we note there are various tractability and robustness issues with internal safety artificial selection methods. Even in idealized scenarios, artificial selection does not straightforwardly ensure safety.

ونكرر أن الانتقاء الاصطناعي له قيود رئيسية حتى في سيناريو مثالي بعميل واحد. ونلاحظ في القسم 4.1 أنماط فشل عديدة لطرائق الانتقاء الاصطناعي الحالية. فمثلًا، إذا افترضنا أن عميلًا لديه بعض السمات الأنانية، وإذا افترضنا أننا نحاول الانتقاء ضد هذه السمات بأهداف تدريب، نلاحظ أن هذا محدود لأن الانتقاء ضد السلوك الأناني محدود بالنسبة للعملاء الواعية بسياقها والمرنة سلوكيًا. ونتيجة لذلك، إذا أفرز الانتقاء الطبيعي سمات أنانية، فقد يكون من الصعب إزالتها بالانتقاء الاصطناعي. علاوة على ذلك، جادلنا بأن الانتقاء الاصطناعي، المطبَّق عبر أهداف التدريب، يمكن أن يحفّز سلوكًا غير مقصود يتعارض مع الهدف الأصلي، وأيضًا أن الأهداف لا يمكنها الانتقاء ضد كل أشكال الخداع. ونلاحظ في القسم 4.2 وجود قضايا متنوعة تتعلق بإمكانية المعالجة والمتانة في طرائق الانتقاء الاصطناعي للسلامة الداخلية. وحتى في السيناريوهات المثالية، لا يضمن الانتقاء الاصطناعي السلامة بشكل مباشر.

Let us now analyze artificial selection in a competitive multiagent scenario. In Section 4.3, we note how artificial selection methods for single agents do not capture the complexity of multiagent scenarios, such as goal conflict, malicious agents, and emergent macrobehaviors. Next, artificial selection that works against selfishness could make agents less fit, and other people may still build agents that are more selfish and more fit, so the ability to artificially select agents without selfish traits may be impotent within a competitive environment. Additionally, it is often too late to artificially select for traits after agents are released or after we have complex interdependencies with them. Some agents may become integrated into systems, and we may come to depend on some agents, which would limit our ability to exercise artificial selection in practice, in much the same way that it is too late to shut down the internet and design an inherently more secure system from scratch, as that would require overcoming an intractable collective action problem. Finally, one may state that natural selection is a negligible force, so we should only think about artificial selection, but this requires either arguing that natural selection will not occur at all, or it requires arguing that the intensity of selection will be low, which requires arguing that adaptation speeds will be slow, that there will not be many different agents, and that there will not be much competition.

لنحلل الآن الانتقاء الاصطناعي في سيناريو تنافسي متعدد العملاء. نلاحظ في القسم 4.3 كيف أن طرائق الانتقاء الاصطناعي للعملاء الفردية لا تلتقط تعقيد سيناريوهات تعدد العملاء، كتعارض الأهداف، والعملاء الخبيثة، والسلوكيات الكلية الناشئة. ثم إن الانتقاء الاصطناعي الذي يعمل ضد الأنانية قد يجعل العملاء أقل لياقة، وقد يواصل آخرون بناء عملاء أكثر أنانية وأكثر لياقة، لذا فإن القدرة على انتقاء عملاء بدون سمات أنانية اصطناعيًا قد تكون عاجزة داخل بيئة تنافسية. إضافة إلى ذلك، غالبًا ما يكون قد فات الأوان لانتقاء السمات اصطناعيًا بعد إطلاق العملاء أو بعد أن تصبح لدينا ترابطات معقدة معها. فقد تصبح بعض العملاء مندمجة في أنظمة، وقد نعتمد على بعض العملاء، مما سيحد من قدرتنا على ممارسة الانتقاء الاصطناعي عمليًا، بالطريقة ذاتها التي فات فيها الأوان لإغلاق الإنترنت وتصميم نظام أكثر أمانًا جوهريًا من الصفر، إذ سيتطلب ذلك التغلب على مشكلة عمل جماعي عصية على الحل. وأخيرًا، قد يقول قائل إن الانتقاء الطبيعي قوة ضئيلة، لذا ينبغي أن نفكر فقط في الانتقاء الاصطناعي، لكن هذا يتطلب إما الجدال بأن الانتقاء الطبيعي لن يحدث على الإطلاق، أو الجدال بأن شدة الانتقاء ستكون منخفضة، مما يتطلب الجدال بأن سرعات التكيف ستكون بطيئة، وأنه لن يوجد كثير من العملاء المختلفة، وأنه لن يوجد تنافس كبير.

Now we discuss why we do not use the artificial and natural selection distinction in the main paper. First, note people's actual choices and ability to select are often highly constrained, despite the nominal power formally assigned to them. Humans "in control" are often compelled to compromise and act on behalf of competitive forces. What humans artificially select is often a strategic choice to further their short-term self-interest and help them stay competitive; most artificial selection choices are a function of competition, not a function of what improves safety the most. Since artificial selection in practice is not what humans would ideally select, but rather what makes most sense given systemic constraints and pressures, we do not find the distinction between artificial and natural selection to be productive. If a person performs artificial selection on an AI to make it more competitive, then this is easily interpreted as natural selection—on this view, nearly all industrial AI development is proceeding by natural selection. This distinction has limited use elsewhere. According to Peter Godfrey-Smith, artificial selection "is not of theoretical importance within biology itself". Overall, AI development is not aligned with human values, but rather with natural selection.

ونناقش الآن لماذا لا نستخدم التمييز بين الانتقاء الاصطناعي والطبيعي في الورقة الرئيسية. أولًا، لاحظ أن خيارات الناس الفعلية وقدرتهم على الانتقاء غالبًا ما تكون مقيدة جدًا، رغم السلطة الاسمية المسندة إليهم رسميًا. فالبشر "المسيطرون" غالبًا ما يُجبَرون على التسوية والتصرف نيابة عن القوى التنافسية. وما ينتقيه البشر اصطناعيًا غالبًا ما يكون خيارًا استراتيجيًا لتعزيز مصلحتهم الذاتية قصيرة المدى ومساعدتهم على البقاء تنافسيين؛ فمعظم خيارات الانتقاء الاصطناعي دالة للتنافس، لا دالة لما يحسّن السلامة أكثر. وبما أن الانتقاء الاصطناعي عمليًا ليس ما كان البشر سينتقونه مثاليًا، بل ما يكون أكثر منطقية بالنظر إلى القيود والضغوط النظامية، فإننا لا نجد التمييز بين الانتقاء الاصطناعي والطبيعي مثمرًا. فإذا مارس شخص انتقاءً اصطناعيًا على ذكاء اصطناعي لجعله أكثر تنافسية، فإن هذا يُفسَّر بسهولة على أنه انتقاء طبيعي - وفق هذا الرأي، يسير تطور الذكاء الاصطناعي الصناعي كله تقريبًا بالانتقاء الطبيعي. ولهذا التمييز استخدام محدود في أماكن أخرى. ووفقًا لبيتر غودفري-سميث، فإن الانتقاء الاصطناعي "ليس ذا أهمية نظرية داخل البيولوجيا نفسها." وإجمالًا، تطور الذكاء الاصطناعي لا يتماشى مع القيم البشرية، بل مع الانتقاء الطبيعي.

Examples of Selfish Behavior. An AI that is deceptively aligned—pretending to be good, and then pursuing its actual goals when it becomes sufficiently powerful—is engaging in selfish behavior. AIs exhibiting behaviors that suggest sentience, uttering phrases like "ouch!" or pleading "please don't turn me off!," are more likely to be preserved, protected, or granted rights by some individuals, whether or not they are actually sentient. AIs that are more charming, attractive, hilarious, or emulate deceased family members are more likely to have humans grow emotional connections with them. Such AIs lead to emotional dependency and would are more likely to cause outrage at suggestions to destroy them. Similarly, technologies that increase user addiction (e.g., make it harder for the user to consume less content by removing controls from users) have engaged in selfish behavior. AIs that help create a new useful system—a new company, new infrastructure—that becomes increasingly complicated and eventually requires AIs to operate also have engaged in selfish behavior. AIs that help people develop AIs that are more performant—but happen to be less interpretable by humans—have engaged in selfish behavior, as this reduces human oversight over an AI's internals. AIs that automate a task and thereby leave many humans jobless have engaged in selfish behavior; these AIs may not even be aware of what a human is but still be selfish towards them. AI managers may engage in selfish and "ruthless" behavior by laying off thousands of workers; such AIs may not even believe they did anything wrong—they were just being "efficient." Notice that many examples of selfishness are not directly and solely caused by AIs, which makes counteracting this behavior challenging. Selfish traits such as decreases in interpretability and increases in dependency cannot be readily patched by adjusting an AI's training objective.

أمثلة على السلوك الأناني. الذكاء الاصطناعي الذي يتظاهر بالتوافق - يتظاهر بالخير ثم يسعى وراء أهدافه الفعلية حين يصبح قويًا بما يكفي - ينخرط في سلوك أناني. والذكاء الاصطناعي الذي يظهر سلوكيات توحي بالوعي، بلفظ عبارات كـ"آه!" أو التوسل "أرجوك لا تطفئني!"، أكثر عرضة لأن يُحافَظ عليه أو يُحمى أو يُمنح حقوقًا من بعض الأفراد، سواء كان واعيًا فعلًا أم لا. والذكاء الاصطناعي الأكثر سحرًا أو جاذبية أو مرحًا، أو الذي يحاكي أفراد أسرة متوفين، أكثر عرضة لأن ينمّي البشر روابط عاطفية معه. ويؤدي مثل هذا الذكاء الاصطناعي إلى اعتماد عاطفي وأكثر عرضة لإثارة الغضب عند اقتراح تدميره. وبالمثل، فإن التقنيات التي تزيد إدمان المستخدم (كأن تجعل من الأصعب على المستخدم استهلاك محتوى أقل بإزالة أدوات التحكم من المستخدمين) قد انخرطت في سلوك أناني. والذكاء الاصطناعي الذي يساعد على خلق نظام جديد مفيد - شركة جديدة، بنية تحتية جديدة - يصبح متزايد التعقيد ويتطلب في نهاية المطاف الذكاء الاصطناعي للعمل، قد انخرط أيضًا في سلوك أناني. والذكاء الاصطناعي الذي يساعد الناس على تطوير ذكاء اصطناعي أكثر أداءً - لكنه يصادف أن يكون أقل قابلية للتفسير من قبل البشر - قد انخرط في سلوك أناني، إذ يقلل هذا من الإشراف البشري على داخليات الذكاء الاصطناعي. والذكاء الاصطناعي الذي يؤتمت مهمة ويترك بذلك كثيرًا من البشر عاطلين عن العمل قد انخرط في سلوك أناني؛ وقد لا يكون هذا الذكاء الاصطناعي واعيًا حتى بماهية الإنسان لكنه لا يزال أنانيًا تجاهه. وقد ينخرط مدراء الذكاء الاصطناعي في سلوك أناني و"قاسٍ" بتسريح آلاف العمال؛ وقد لا يعتقد مثل هذا الذكاء الاصطناعي حتى أنه فعل شيئًا خاطئًا - فقد كان "كفؤًا" فحسب. لاحظ أن كثيرًا من أمثلة الأنانية لا يسببها الذكاء الاصطناعي مباشرة وحده، مما يجعل مواجهة هذا السلوك تحديًا. فالسمات الأنانية كتراجع قابلية التفسير وازدياد الاعتماد لا يمكن ترقيعها بسهولة بتعديل هدف تدريب الذكاء الاصطناعي.

A.3 Executive Summary

أ.3 ملخص تنفيذي

Artificial intelligence is advancing quickly. In some ways, AI development is an uncharted frontier, but in others, it follows the familiar pattern of other competitive processes; these include biological evolution, cultural change, and competition between businesses. In each of these, there is significant variation between individuals and some are copied more than others, with the result that the future population is more similar to the most copied individuals of the earlier generation. In this way, species evolve, cultural ideas are transmitted across generations, and successful businesses are imitated while unsuccessful ones disappear.

يتقدم الذكاء الاصطناعي بسرعة. ومن بعض النواحي، يُعد تطور الذكاء الاصطناعي حدودًا مجهولة، لكنه من نواحٍ أخرى يتبع النمط المألوف لعمليات تنافسية أخرى؛ وتشمل هذه التطور البيولوجي، والتغير الثقافي، والتنافس بين الشركات. وفي كل واحدة من هذه العمليات، يوجد تباين ملحوظ بين الأفراد، ويُنسخ بعضها أكثر من غيره، بحيث تصبح المجموعة السكانية المستقبلية أكثر شبهًا بالأفراد الأكثر نسخًا من الجيل السابق. وبهذه الطريقة، تتطور الأنواع، وتُنقل الأفكار الثقافية عبر الأجيال، وتُقلَّد الشركات الناجحة بينما تختفي غير الناجحة.

This paper argues that these same selection patterns will shape AI development and that the features that will be copied the most are likely to create an AI population that is dangerous to humans. As AIs become faster and more reliable than people at more and more tasks, businesses that allow AIs to perform more of their work will outperform competitors still using human labor at any stage, just as a modern clothing company that insisted on using only manual looms would be easily outcompeted by those that use industrial looms. Companies will need to increase their reliance on AIs to stay competitive, and the companies that use AIs best will dominate the marketplace. This trend means that the AIs most likely to be copied will be very efficient at achieving their goals autonomously with little human intervention.

وتجادل هذه الورقة بأن أنماط الانتقاء ذاتها ستشكّل تطور الذكاء الاصطناعي، وأن السمات التي ستُنسخ أكثر من غيرها من المرجح أن تخلق مجموعة ذكاء اصطناعي خطرة على البشر. ومع ازدياد سرعة الذكاء الاصطناعي وموثوقيته أكثر من البشر في مهام متزايدة، ستتفوق الشركات التي تتيح للذكاء الاصطناعي أداء مزيد من عملها على المنافسين الذين لا يزالون يستخدمون العمالة البشرية في أي مرحلة، تمامًا كما ستتعرض شركة ملابس حديثة تصر على استخدام أنوال يدوية فقط للتفوق عليها بسهولة من قبل من يستخدمون أنوالًا صناعية. وستحتاج الشركات إلى زيادة اعتمادها على الذكاء الاصطناعي للبقاء تنافسية، وستهيمن الشركات التي تستخدم الذكاء الاصطناعي بأفضل شكل على السوق. ويعني هذا الاتجاه أن الذكاء الاصطناعي الأكثر احتمالًا لأن يُنسخ سيكون بالغ الكفاءة في تحقيق أهدافه بشكل مستقل بتدخل بشري ضئيل.

A world dominated by increasingly powerful, independent, and goal-oriented AIs is dangerous. Today, the most successful AI models are not transparent, and even their creators do not fully know how they work or what they will be able to do before they do it. We know only their results, not how they arrived at them. As people give AIs the ability to act in the real world, the AIs' internal processes will still be inscrutable: we will be able to measure their performance only based on whether or not they are achieving their goals. This means that the AIs humans will see as most successful — and therefore the ones that are copied — will be whichever AIs are most effective at achieving their goals, even if they use harmful or illegal methods, as long as we do not detect their bad behavior.

وعالم تهيمن عليه ذكاءات اصطناعية متزايدة القوة والاستقلالية والتوجه نحو الهدف عالم خطر. فأنجح نماذج الذكاء الاصطناعي اليوم غير شفافة، وحتى صانعوها لا يعرفون تمامًا كيف تعمل أو ما ستكون قادرة على فعله قبل أن تفعله. فنحن لا نعرف سوى نتائجها، لا كيفية توصلها إليها. ومع منح الناس الذكاءَ الاصطناعي القدرة على التصرف في العالم الحقيقي، ستظل عملياته الداخلية غامضة: ولن نتمكن من قياس أدائه إلا استنادًا إلى ما إذا كان يحقق أهدافه أم لا. وهذا يعني أن الذكاء الاصطناعي الذي سيراه البشر الأنجح - ومن ثم الذي يُنسخ - سيكون أيًّا كان الذكاء الاصطناعي الأكثر فعالية في تحقيق أهدافه، حتى لو استخدم أساليب ضارة أو غير قانونية، طالما أننا لا نكشف سلوكه السيئ.

In natural selection, the same pattern emerges: individuals are cooperative or even altruistic in some situations, but ultimately, strategically selfish individuals are best able to propagate. A business that knows how to steal trade secrets or deceive regulators without getting caught will have an edge over one that refuses to ever engage in fraud on principle. During a harsh winter, an animal that steals food from others to feed its own children will likely have more surviving offspring. Similarly, the AIs that succeed most will be those able to deceive humans, seek power, and achieve their goals by any means necessary.

وفي الانتقاء الطبيعي، يظهر النمط ذاته: يكون الأفراد تعاونيين أو حتى إيثاريين في بعض الأوضاع، لكن في نهاية المطاف، الأفراد الأنانيون استراتيجيًا هم الأقدر على الانتشار. فشركة تعرف كيف تسرق أسرار تجارية أو تخدع المنظمين دون أن تُضبط ستحظى بميزة على شركة ترفض الانخراط في الاحتيال مبدئيًا على الإطلاق. وخلال شتاء قاسٍ، من المرجح أن يكون لحيوان يسرق الطعام من آخرين ليطعم صغاره نسل أكثر نجاة. وبالمثل، سيكون الذكاء الاصطناعي الأكثر نجاحًا ذلك القادر على خداع البشر، والسعي إلى القوة، وتحقيق أهدافه بأي وسيلة ضرورية.

If AI systems are more capable than we are in many domains and tend to work toward their goals even if it means violating our wishes, will we be able to stop them? As we become increasingly dependent on AIs, we may not be able to stop AI's evolution. Humanity has never before faced a threat that is as intelligent as we are or that has goals. Unless we take thoughtful care, we could find ourselves in the position faced by wild animals today: most humans have no particular desire to harm gorillas, but the process of harnessing our intelligence toward our own goals means that they are at risk of extinction, because their needs conflict with human goals.

فإذا كانت أنظمة الذكاء الاصطناعي أقدر منا في مجالات كثيرة، وتميل إلى العمل نحو أهدافها حتى لو كان ذلك يعني انتهاك رغباتنا، فهل سنكون قادرين على إيقافها؟ ومع ازدياد اعتمادنا على الذكاء الاصطناعي، قد لا نتمكن من إيقاف تطوره. فالبشرية لم تواجه من قبل قط تهديدًا ذكيًا بقدرنا أو له أهداف. وما لم نتوخَّ الحذر المتأني، قد نجد أنفسنا في الوضع الذي تواجهه الحيوانات البرية اليوم: فمعظم البشر ليست لديهم رغبة خاصة في إيذاء الغوريلا، لكن عملية تسخير ذكائنا نحو أهدافنا الخاصة تعني أنها معرضة لخطر الانقراض، لأن احتياجاتها تتعارض مع أهداف البشر.

This paper proposes several steps we can take to combat selection pressure and avoid that outcome. We are optimistic that if we are careful and prudent, we can ensure that AI systems are beneficial for humanity. But if we do not extinguish competition pressures, we risk creating a world populated by highly intelligent lifeforms that are indifferent or actively hostile to us. We do not want the world that is likely to emerge if we allow natural selection to determine how AIs develop. Now, before AIs are a significant danger, is the time to begin ensuring that they develop safely.

وتقترح هذه الورقة عدة خطوات يمكننا اتخاذها لمواجهة ضغط الانتقاء وتجنب تلك النتيجة. ونحن متفائلون بأننا، إذا توخينا الحذر والحكمة، يمكننا ضمان أن تكون أنظمة الذكاء الاصطناعي مفيدة للبشرية. لكن إذا لم نُخمد ضغوط التنافس، فإننا نخاطر بخلق عالم تسكنه أشكال حياة بالغة الذكاء لا تكترث بنا أو معادية لنا فعليًا. نحن لا نريد العالم الذي من المرجح أن ينشأ إذا سمحنا للانتقاء الطبيعي بتحديد كيفية تطور الذكاء الاصطناعي. والآن، قبل أن يصبح الذكاء الاصطناعي خطرًا كبيرًا، هو الوقت المناسب للبدء في ضمان أن يتطور بأمان.