Welcome to the Era of Experience
David Silver, Richard S. Sutton | ديفيد سيلفر، ريتشارد س. ساتون
Google DeepMind | جوجل ديب مايند
Preprint chapter, "Designing an Intelligence" (MIT Press), 2025 | فصل أولي من كتاب "تصميم الذكاء" (منشورات معهد ماساتشوستس للتقنية)، 2025
Welcome to the Era of Experience
مرحبًا بكم في عصر التجربة
ABSTRACT
الملخص
We stand on the threshold of a new era in artificial intelligence that promises to achieve an unprecedented level of ability. A new generation of agents will acquire superhuman capabilities by learning predominantly from experience. This note explores the key characteristics that will define this upcoming era.
نقف على عتبة عصر جديد في الذكاء الاصطناعي يعِد بتحقيق مستوى غير مسبوق من القدرة. سيكتسب جيل جديد من العملاء قدرات تفوق قدرات البشر من خلال التعلّم بشكل أساسي من التجربة. تستكشف هذه المذكرة السمات الرئيسية التي ستحدد ملامح هذا العصر المقبل.
The Era of Human Data
عصر البيانات البشرية
Artificial intelligence (AI) has made remarkable strides over recent years by training on massive amounts of human-generated data and fine-tuning with expert human examples and preferences. This approach is exemplified by large language models (LLMs) that have achieved a sweeping level of generality. A single LLM can now perform tasks spanning from writing poetry and solving physics problems to diagnosing medical issues and summarising legal documents.
حقق الذكاء الاصطناعي خطوات لافتة خلال السنوات الأخيرة من خلال التدريب على كميات هائلة من البيانات التي ينتجها البشر، وضبطه الدقيق باستخدام أمثلة وتفضيلات بشرية متخصصة. ويجسّد هذا النهج نماذج اللغة الكبيرة (LLMs) التي بلغت مستوى واسعًا من التعميم. فبات بإمكان نموذج لغة كبير واحد أداء مهام تمتد من كتابة الشعر وحل مسائل الفيزياء إلى تشخيص الحالات الطبية وتلخيص الوثائق القانونية.
However, while imitating humans is enough to reproduce many human capabilities to a competent level, this approach in isolation has not and likely cannot achieve superhuman intelligence across many important topics and tasks. In key domains such as mathematics, coding, and science, the knowledge extracted from human data is rapidly approaching a limit. The majority of high-quality data sources - those that can actually improve a strong agent's performance - have either already been, or soon will be consumed. The pace of progress driven solely by supervised learning from human data is demonstrably slowing, signalling the need for a new approach. Furthermore, valuable new insights, such as new theorems, technologies or scientific breakthroughs, lie beyond the current boundaries of human understanding and cannot be captured by existing human data.
غير أنه رغم أن محاكاة البشر تكفي لإعادة إنتاج كثير من القدرات البشرية بمستوى كفء، فإن هذا النهج بمفرده لم يحقق، ومن غير المرجح أن يحقق، ذكاءً يفوق البشر في كثير من الموضوعات والمهام المهمة. ففي مجالات رئيسية مثل الرياضيات والبرمجة والعلوم، تقترب المعرفة المستخلصة من البيانات البشرية بسرعة من حد أقصى. ومعظم مصادر البيانات عالية الجودة - تلك القادرة فعلًا على تحسين أداء عميل قوي - إما استُهلكت بالفعل أو ستُستهلك قريبًا. كما أن وتيرة التقدم المدفوعة فقط بالتعلّم الخاضع للإشراف من البيانات البشرية تتباطأ بوضوح، مما يشير إلى الحاجة إلى نهج جديد. علاوة على ذلك، فإن الرؤى الجديدة القيّمة، كالنظريات الجديدة أو التقنيات أو الاختراقات العلمية، تقع خارج الحدود الحالية للفهم البشري، ولا يمكن التقاطها من البيانات البشرية المتاحة.
The Era of Experience
عصر التجربة
To progress significantly further, a new source of data is required. This data must be generated in a way that continually improves as the agent becomes stronger; any static procedure for synthetically generating data will quickly become outstripped. This can be achieved by allowing agents to learn continually from their own experience, i.e., data that is generated by the agent interacting with its environment. AI is at the cusp of a new period in which experience will become the dominant medium of improvement and ultimately dwarf the scale of human data used in today's systems.
لإحراز تقدم أكبر بكثير، لا بد من مصدر جديد للبيانات. ويجب أن تُولَّد هذه البيانات بطريقة تتحسن باستمرار مع ازدياد قوة العميل؛ فأي إجراء ثابت لتوليد البيانات اصطناعيًا سرعان ما يصبح متجاوَزًا. ويمكن تحقيق ذلك بالسماح للعملاء بالتعلّم المستمر من تجربتهم الخاصة، أي البيانات التي يولّدها العميل من خلال تفاعله مع بيئته. ويقف الذكاء الاصطناعي على أعتاب حقبة جديدة ستصبح فيها التجربة الوسيط المهيمن للتحسين، وستتجاوز في نهاية المطاف حجم البيانات البشرية المستخدمة في أنظمة اليوم.
This transition may have already started, even for the large language models that epitomise human-centric AI. One example is in the capability of mathematics. AlphaProof recently became the first program to achieve a medal in the International Mathematical Olympiad, eclipsing the performance of human-centric approaches. Initially exposed to around a hundred thousand formal proofs, created over many years by human mathematicians, AlphaProof's reinforcement learning (RL) algorithm subsequently generated a hundred million more through continual interaction with a formal proving system. This focus on interactive experience allowed AlphaProof to explore mathematical possibilities beyond the confines of pre-existing formal proofs, so as to discover solutions to novel and challenging problems. Informal mathematics has also achieved success by replacing expert generated data with self-generated data; for example, recent work from DeepSeek "underscores the power and beauty of reinforcement learning: rather than explicitly teaching the model on how to solve a problem, we simply provide it with the right incentives, and it autonomously develops advanced problem-solving strategies."
وربما يكون هذا التحول قد بدأ بالفعل، حتى في نماذج اللغة الكبيرة التي تجسّد الذكاء الاصطناعي المتمحور حول الإنسان. ومن الأمثلة على ذلك قدرات الرياضيات. فقد أصبح برنامج AlphaProof مؤخرًا أول برنامج يحرز ميدالية في الأولمبياد الدولي للرياضيات، متجاوزًا أداء المناهج المتمحورة حول الإنسان. وبعد أن تعرّض في البداية لنحو مئة ألف برهان رسمي، أنتجها رياضيون بشر على مدى سنوات عديدة، ولّدت خوارزمية التعلّم المعزَّز (RL) الخاصة بـ AlphaProof مئة مليون برهان إضافي عبر التفاعل المستمر مع نظام إثبات رسمي. وقد سمح هذا التركيز على التجربة التفاعلية لـ AlphaProof باستكشاف إمكانات رياضية تتجاوز حدود البراهين الرسمية القائمة سلفًا، بحيث اكتشف حلولًا لمسائل جديدة وصعبة. كما حققت الرياضيات غير الرسمية نجاحًا باستبدال البيانات التي يولّدها الخبراء ببيانات يولّدها النظام نفسه؛ فعلى سبيل المثال، يؤكد عمل حديث من DeepSeek أن "قوة التعلّم المعزَّز وجماله يكمنان في أننا، بدلًا من تعليم النموذج صراحةً كيفية حل مسألة ما، نكتفي بمنحه الحوافز الصحيحة، فيطوّر من تلقاء نفسه استراتيجيات متقدمة لحل المسائل."
Our contention is that incredible new capabilities will arise once the full potential of experiential learning is harnessed. This era of experience will likely be characterised by agents and environments that, in addition to learning from vast quantities of experiential data, will break through the limitations of human-centric AI systems in several further dimensions:
ونحن نرى أن قدرات جديدة مذهلة ستظهر بمجرد تسخير الإمكانات الكاملة للتعلّم من التجربة. ومن المرجح أن يتسم عصر التجربة هذا بعملاء وبيئات لن تكتفي بالتعلّم من كميات هائلة من بيانات التجربة، بل ستتجاوز أيضًا حدود أنظمة الذكاء الاصطناعي المتمحورة حول الإنسان في أبعاد إضافية عدة:
Agents will inhabit streams of experience, rather than short snippets of interaction.
سيعيش العملاء داخل تدفقات من التجربة، بدلًا من مقتطفات قصيرة من التفاعل.
Their actions and observations will be richly grounded in the environment, rather than interacting via human dialogue alone.
ستكون أفعالهم ومشاهداتهم مترسخة بعمق في البيئة، بدلًا من الاقتصار على التفاعل عبر الحوار البشري وحده.
Their rewards will be grounded in their experience of the environment, rather than coming from human prejudgement.
ستكون مكافآتهم متجذرة في تجربتهم للبيئة، بدلًا من أن تأتي من حكم بشري مسبق.
They will plan and/or reason about experience, rather than reasoning solely in human terms.
وسيخططون و/أو يفكرون بناءً على التجربة، بدلًا من الاقتصار على التفكير بمصطلحات بشرية بحتة.
We believe that today's technology, with appropriately chosen algorithms, already provides a sufficiently powerful foundation to achieve these breakthroughs. Furthermore, the pursuit of this agenda by the AI community will spur new innovations in these directions that rapidly progress AI towards truly superhuman agents.
ونعتقد أن تقنية اليوم، مع اختيار الخوارزميات المناسبة، توفر بالفعل أساسًا قويًا بما يكفي لتحقيق هذه الاختراقات. وعلاوة على ذلك، فإن سعي مجتمع الذكاء الاصطناعي لتحقيق هذا البرنامج سيحفّز ابتكارات جديدة في هذه الاتجاهات، تدفع الذكاء الاصطناعي بسرعة نحو عملاء يفوقون البشر فعلًا.
Streams
التدفقات
An experiential agent can continue to learn throughout a lifetime. In the era of human data, language-based AI has largely focused on short interaction episodes: e.g., a user asks a question and (perhaps after a few thinking steps or tool-use actions) the agent responds. Typically, little or no information carries over from one episode to the next, precluding any adaptation over time. Furthermore, the agent aims exclusively for outcomes within the current episode, such as directly answering a user's question. In contrast, humans (and other animals) exist in an ongoing stream of actions and observations that continues for many years. Information is carried across the entire stream, and their behaviour adapts from past experiences to self-correct and improve. Furthermore, goals may be specified in terms of actions and observations that stretch far into the future of the stream. For example, humans may select actions to achieve long-term goals like improving their health, learning a language, or achieving a scientific breakthrough.
يمكن للعميل الذي يتعلم من التجربة أن يواصل التعلّم طوال حياته. وفي عصر البيانات البشرية، انصبّ تركيز الذكاء الاصطناعي القائم على اللغة إلى حد كبير على حلقات تفاعل قصيرة: فمثلًا، يطرح المستخدم سؤالًا، ثم يجيب العميل (ربما بعد بضع خطوات من التفكير أو استخدام أدوات). وعادة ما ينتقل القليل من المعلومات أو لا شيء منها من حلقة إلى التي تليها، مما يحول دون أي تكيّف عبر الزمن. كما أن العميل يستهدف حصرًا نتائج داخل الحلقة الحالية، كالإجابة المباشرة عن سؤال المستخدم. وعلى النقيض من ذلك، يعيش البشر (وسائر الحيوانات) في تدفق مستمر من الأفعال والمشاهدات يمتد لسنوات عديدة. فتنتقل المعلومات عبر التدفق بأكمله، ويتكيف سلوكهم بناءً على تجارب الماضي لتصحيح أنفسهم والتحسّن. علاوة على ذلك، قد تُحدَّد الأهداف بأفعال ومشاهدات تمتد بعيدًا في مستقبل التدفق. فقد يختار البشر مثلًا أفعالًا لتحقيق أهداف طويلة المدى، كتحسين صحتهم أو تعلّم لغة أو تحقيق اختراق علمي.
Powerful agents should have their own stream of experience that progresses, like humans, over a long time-scale. This will allow agents to take actions to achieve future goals, and to continuously adapt over time to new patterns of behaviour. For example, a health and wellness agent connected to a user's wearables could monitor sleep patterns, activity levels, and dietary habits over many months. It could then provide personalized recommendations, encouragement, and adjust its guidance based on long-term trends and the user's specific health goals. Similarly, a personalized education agent could track a user's progress in learning a new language, identify knowledge gaps, adapt to their learning style, and adjust its teaching methods over months or even years. Furthermore, a science agent could pursue ambitious goals, such as discovering a new material or reducing carbon dioxide. Such an agent could analyse real-world observations over an extended period, developing and running simulations, and suggesting real-world experiments or interventions.
وينبغي أن يكون للعملاء الأقوياء تدفق تجربة خاص بهم يتقدم، كما هو الحال مع البشر، على نطاق زمني طويل. وسيتيح ذلك للعملاء اتخاذ إجراءات لتحقيق أهداف مستقبلية، والتكيف باستمرار عبر الزمن مع أنماط سلوك جديدة. فعلى سبيل المثال، يمكن لعميل صحي متصل بالأجهزة القابلة للارتداء الخاصة بالمستخدم أن يراقب أنماط النوم ومستويات النشاط والعادات الغذائية على مدى أشهر عديدة. ثم بوسعه تقديم توصيات مخصصة وتشجيع، وتعديل إرشاداته بناءً على الاتجاهات طويلة المدى وأهداف المستخدم الصحية المحددة. وبالمثل، يمكن لعميل تعليمي مخصص أن يتتبع تقدّم المستخدم في تعلّم لغة جديدة، وتحديد الفجوات المعرفية، والتكيف مع أسلوب تعلّمه، وتعديل طرائق تدريسه على مدى أشهر أو حتى سنوات. علاوة على ذلك، يمكن لعميل علمي أن يسعى لتحقيق أهداف طموحة، كاكتشاف مادة جديدة أو خفض ثاني أكسيد الكربون. ويمكن لعميل من هذا القبيل أن يحلل مشاهدات من العالم الحقيقي على مدى فترة ممتدة، مطوّرًا محاكاة ومشغّلًا لها، ومقترحًا تجارب أو تدخلات في العالم الحقيقي.
In each case, the agent takes a sequence of steps so as to maximise long-term success with respect to the specified goal. An individual step may not provide any immediate benefit, or may even be detrimental in the short term, but may nevertheless contribute in aggregate to longer term success. This contrasts strongly with current AI systems that provide immediate responses to requests, without any ability to measure or optimise the future consequences of their actions on the environment.
وفي كل حالة، يتخذ العميل سلسلة من الخطوات بهدف تعظيم النجاح طويل المدى فيما يخص الهدف المحدد. وقد لا توفر خطوة بمفردها أي فائدة فورية، بل قد تكون ضارة في المدى القصير، لكنها مع ذلك قد تسهم إجمالًا في النجاح على المدى الأطول. ويتناقض هذا بشدة مع أنظمة الذكاء الاصطناعي الحالية التي تقدم استجابات فورية للطلبات، دون أي قدرة على قياس أو تحسين العواقب المستقبلية لأفعالها على البيئة.
Actions and Observations
الأفعال والمشاهدات
Agents in the era of experience will act autonomously in the real world. LLMs in the era of human data focused primarily on human-privileged actions and observations that output text to a user, and input text from the user back into the agent. This differs markedly from natural intelligence, in which an animal interacts with its environment through motor control and sensors. While animals, and most notably humans, may communicate with other animals, this occurs through the same interface as other sensorimotor control rather than a privileged channel.
سيتصرف العملاء في عصر التجربة بشكل مستقل في العالم الحقيقي. وقد ركزت نماذج اللغة الكبيرة في عصر البيانات البشرية أساسًا على أفعال ومشاهدات ذات امتياز بشري تُخرج نصًا إلى المستخدم، وتُدخل نصًا من المستخدم إلى العميل. ويختلف هذا اختلافًا واضحًا عن الذكاء الطبيعي، حيث يتفاعل الحيوان مع بيئته من خلال التحكم الحركي والحواس. وفي حين قد تتواصل الحيوانات، ولا سيما البشر، مع حيوانات أخرى، فإن ذلك يحدث عبر الواجهة ذاتها المستخدمة في التحكم الحسي الحركي الآخر، لا عبر قناة ذات امتياز خاص.
It has long been recognised that LLMs may also invoke actions in the digital world, for example by calling APIs. Initially, these capabilities came largely from human examples of tool-use, rather than from the experience of the agent. However, coding and tool-use capabilities have built increasingly upon execution feedback, where the agent actually runs code and observes what happens. Recently, a new wave of prototype agents have started to interact with computers in an even more general manner, by using the same interface that humans use to operate a computer. These changes herald a transition from exclusively human-privileged communication, to much more autonomous interactions where the agent is able to act independently in the world. Such agents will be able to actively explore the world, adapt to changing environments, and discover strategies that might never occur to a human.
ومن المعترف به منذ زمن طويل أن بإمكان نماذج اللغة الكبيرة أيضًا استدعاء أفعال في العالم الرقمي، كاستدعاء واجهات برمجة التطبيقات (APIs) على سبيل المثال. وفي البداية، كانت هذه القدرات مستمدة أساسًا من أمثلة بشرية على استخدام الأدوات، لا من تجربة العميل نفسه. غير أن قدرات البرمجة واستخدام الأدوات باتت تعتمد على نحو متزايد على تغذية راجعة من التنفيذ، حيث يشغّل العميل فعليًا شفرة برمجية ويلاحظ ما يحدث. ومؤخرًا، بدأت موجة جديدة من العملاء التجريبيين بالتفاعل مع الحواسيب بطريقة أعمّ حتى من ذلك، باستخدام الواجهة ذاتها التي يستخدمها البشر لتشغيل الحاسوب. وتُنبئ هذه التغيرات بانتقال من التواصل الحصري ذي الامتياز البشري إلى تفاعلات أكثر استقلالية بكثير، يكون فيها العميل قادرًا على التصرف بشكل مستقل في العالم. وستكون مثل هذه العملاء قادرة على استكشاف العالم بنشاط، والتكيف مع البيئات المتغيرة، واكتشاف استراتيجيات قد لا تخطر ببال إنسان قط.
These richer interactions will provide a means to autonomously understand and control the digital world. The agent may use 'human-friendly' actions and observations such as user interfaces, that naturally facilitate communication and collaboration with the user. The agent may also take 'machine-friendly' actions that execute code and call APIs, allowing the agent to act autonomously in service of its goals. In the era of experience, agents will also interact with the real world via digital interfaces. For example, a scientific agent could monitor environmental sensors, remotely operate a telescope, or control a robotic arm in a laboratory to autonomously conduct experiments.
وستوفر هذه التفاعلات الأغنى وسيلةً لفهم العالم الرقمي والتحكم فيه بشكل مستقل. وقد يستخدم العميل أفعالًا ومشاهدات "صديقة للإنسان" كواجهات المستخدم، التي تُيسّر بطبيعتها التواصل والتعاون مع المستخدم. وقد يتخذ العميل أيضًا أفعالًا "صديقة للآلة" تُنفّذ شفرة برمجية وتستدعي واجهات برمجة التطبيقات، مما يتيح له التصرف باستقلالية خدمةً لأهدافه. وفي عصر التجربة، سيتفاعل العملاء أيضًا مع العالم الحقيقي عبر واجهات رقمية. فعلى سبيل المثال، يمكن لعميل علمي مراقبة مستشعرات بيئية، أو تشغيل تلسكوب عن بُعد، أو التحكم بذراع آلية في مختبر لإجراء تجارب بشكل مستقل.
Rewards
المكافآت
What if experiential agents could learn from external events and signals, and not just human preferences?
ماذا لو استطاعت العملاء التي تتعلم من التجربة أن تتعلم من أحداث وإشارات خارجية، لا من تفضيلات البشر فحسب؟
Human-centric LLMs typically optimise for rewards based on human prejudgement: an expert observes the agent's action and decides whether it is a good action, or picks the best agent action among multiple alternatives. For example, an expert may judge a health agent's advice, an educational assistant's teaching, or a scientist agent's suggested experiment. The fact that these rewards or preferences are determined by humans in absence of their consequences, rather than measuring the effect of those actions on the environment, means that they are not directly grounded in the reality of the world. Relying on human prejudgement in this manner usually leads to an impenetrable ceiling on the agent's performance: the agent cannot discover better strategies that are underappreciated by the human rater. To discover new ideas that go far beyond existing human knowledge, it is instead necessary to use grounded rewards: signals that arise from the environment itself. For example, a health assistant could ground the user's health goals into a reward based on a combination of signals such as their resting heart rate, sleep duration, and activity levels, while an educational assistant could use exam results to provide a grounded reward for language learning. Similarly, a science agent with a goal to reduce global warming might use a reward based on empirical observations of carbon dioxide levels, while a goal to discover a stronger material might be grounded in a combination of measurements from a materials simulator, such as tensile strength or Young's modulus.
تُحسِّن نماذج اللغة الكبيرة المتمحورة حول الإنسان عادةً مكافآت مبنية على حكم بشري مسبق: إذ يلاحظ خبير فعل العميل ويقرر ما إذا كان فعلًا جيدًا، أو يختار أفضل فعل من بين بدائل متعددة. فقد يحكم خبير مثلًا على نصيحة عميل صحي، أو تدريس مساعد تعليمي، أو تجربة يقترحها عميل عالِم. وكون هذه المكافآت أو التفضيلات يحددها البشر في غياب معرفتهم بعواقبها الفعلية، بدلًا من قياس أثر تلك الأفعال على البيئة، يعني أنها غير متجذرة مباشرة في واقع العالم. وعادة ما يؤدي الاعتماد على الحكم البشري المسبق بهذه الطريقة إلى سقف لا يمكن تجاوزه لأداء العميل: فلا يستطيع العميل اكتشاف استراتيجيات أفضل لا يقدّرها المقيّم البشري حق قدرها. ولاكتشاف أفكار جديدة تتجاوز المعرفة البشرية القائمة بمراحل، يلزم بدلًا من ذلك استخدام مكافآت متجذرة: إشارات تنبع من البيئة نفسها. فعلى سبيل المثال، يمكن لمساعد صحي أن يُرسّخ أهداف المستخدم الصحية في مكافأة مبنية على مزيج من إشارات كمعدل ضربات القلب أثناء الراحة، ومدة النوم، ومستويات النشاط، بينما يمكن لمساعد تعليمي استخدام نتائج الامتحانات لتقديم مكافأة متجذرة لتعلّم اللغة. وبالمثل، قد يستخدم عميل علمي هدفه خفض الاحتباس الحراري مكافأة مبنية على مشاهدات تجريبية لمستويات ثاني أكسيد الكربون، في حين قد يترسّخ هدف اكتشاف مادة أقوى في مزيج من قياسات محاكي مواد، كمقاومة الشد أو معامل يونغ.
Grounded rewards may arise from humans that are part of the agent's environment. For example, a human user could report whether they found a cake tasty, how fatigued they are after exercising, or the level of pain from a headache, enabling an assistant agent to provide better recipes, refine its fitness suggestions, or improve its recommended medication. Such rewards measure the consequence of the agent's actions within their environment, and should ultimately lead to better assistance than a human expert that prejudges a proposed cake recipe, exercise program, or treatment program.
ويمكن أن تنبع المكافآت المتجذرة من بشر يشكّلون جزءًا من بيئة العميل. فعلى سبيل المثال، يمكن لمستخدم بشري أن يبلغ عما إذا كانت الكعكة لذيذة، أو مدى إرهاقه بعد التمرين، أو مستوى الألم الناجم عن صداع، مما يمكّن عميلًا مساعدًا من تقديم وصفات أفضل، أو تحسين اقتراحاته الرياضية، أو تحسين الدواء الذي يوصي به. وتقيس مثل هذه المكافآت عاقبة أفعال العميل داخل بيئته، وينبغي أن تؤدي في نهاية المطاف إلى مساعدة أفضل من خبير بشري يصدر حكمًا مسبقًا على وصفة كعكة مقترحة أو برنامج تمارين أو برنامج علاجي.
Where do rewards come from, if not from human data? Once agents become connected to the world through rich action and observation spaces, there will be no shortage of grounded signals to provide a basis for reward. In fact, the world abounds with quantities such as cost, error rates, hunger, productivity, health metrics, climate metrics, profit, sales, exam results, success, visits, yields, stocks, likes, income, pleasure/pain, economic indicators, accuracy, power, distance, speed, efficiency, or energy consumption. In addition there are innumerable additional signals arising from the occurrence of specific events, or from features derived from raw sequences of observations and actions.
فمن أين تأتي المكافآت إذن، إن لم يكن من البيانات البشرية؟ بمجرد أن يتصل العملاء بالعالم عبر مساحات غنية من الأفعال والمشاهدات، لن يكون هناك نقص في الإشارات المتجذرة التي توفر أساسًا للمكافأة. والواقع أن العالم يزخر بكميات مثل التكلفة، ومعدلات الخطأ، والجوع، والإنتاجية، والمقاييس الصحية، والمقاييس المناخية، والربح، والمبيعات، ونتائج الامتحانات، والنجاح، والزيارات، والمحاصيل، والأسهم، والإعجابات، والدخل، واللذة/الألم، والمؤشرات الاقتصادية، والدقة، والقدرة، والمسافة، والسرعة، والكفاءة، أو استهلاك الطاقة. وإضافة إلى ذلك، توجد إشارات إضافية لا حصر لها تنشأ من وقوع أحداث محددة، أو من سمات مستخلصة من تسلسلات خام من المشاهدات والأفعال.
One could in principle create a variety of distinct agents, each optimising for one grounded signal as its reward. There is an argument that even a single such reward signal, optimised with great effectiveness, may be sufficient to induce broadly capable intelligence. This is because the achievement of a simple goal in a complex environment may often require a wide variety of skills to be mastered.
ويمكن من حيث المبدأ إنشاء مجموعة متنوعة من العملاء المتمايزين، يُحسِّن كل منها إشارة متجذرة واحدة بوصفها مكافأته. وثمة حجة مفادها أن إشارة مكافأة واحدة من هذا القبيل، إذا حُسِّنت بفعالية كبيرة، قد تكفي لتوليد ذكاء واسع القدرات. ويعود ذلك إلى أن تحقيق هدف بسيط في بيئة معقدة كثيرًا ما يتطلب إتقان طائفة واسعة من المهارات.
However, the pursuit of a single reward signal does not on the surface appear to meet the requirements of a general-purpose AI that can be steered reliably towards arbitrary user-desired behaviours. Is the autonomous optimisation of grounded, non-human reward signals therefore in opposition to the requirements of modern AI systems? We argue that this is not necessarily the case, by sketching one approach that may meet these desiderata; other approaches may also be possible.
غير أن السعي وراء إشارة مكافأة واحدة لا يبدو، للوهلة الأولى، متوافقًا مع متطلبات ذكاء اصطناعي عام الغرض يمكن توجيهه بموثوقية نحو أي سلوك يرغبه المستخدم. فهل يتعارض إذن التحسين المستقل لإشارات مكافأة متجذرة وغير بشرية مع متطلبات أنظمة الذكاء الاصطناعي الحديثة؟ نحن نرى أن هذا ليس بالضرورة صحيحًا، وذلك برسم نهج واحد قد يفي بهذه المتطلبات؛ وقد تكون هناك نهج أخرى ممكنة أيضًا.
The idea is to flexibly adapt the reward, based on grounded signals, in a user-guided manner. For example, the reward function could be defined by a neural network that takes the agent's interactions with both the user and the environment as input, and outputs a scalar reward. This allows the reward to select or combine together signals from the environment in a manner that depends upon the user's goal. For example, a user might specify a broad goal such as 'improve my fitness' and the reward function might return a function of the user's heart rate, sleep duration, and steps taken. Or the user might specify a goal of 'help me learn Spanish' and the reward function could return the user's Spanish exam results.
وتقوم الفكرة على تكييف المكافأة بمرونة، استنادًا إلى إشارات متجذرة، بطريقة يوجهها المستخدم. فعلى سبيل المثال، يمكن تعريف دالة المكافأة بشبكة عصبية تأخذ تفاعلات العميل مع كل من المستخدم والبيئة كمدخلات، وتُخرج مكافأة عددية. ويتيح هذا لدالة المكافأة اختيار إشارات من البيئة أو دمجها معًا بطريقة تعتمد على هدف المستخدم. فقد يحدد المستخدم مثلًا هدفًا عامًا مثل "حسِّن لياقتي"، فتُعيد دالة المكافأة دالةً لمعدل ضربات قلب المستخدم ومدة نومه وعدد خطواته. أو قد يحدد المستخدم هدف "ساعدني على تعلّم الإسبانية"، فتُعيد دالة المكافأة نتائج امتحانات المستخدم في الإسبانية.
Furthermore, users could provide feedback during the learning process, such as their satisfaction level, which could be used to fine-tune the reward function. The reward function can then adapt over time, to improve the way in which it selects or combines signals, and to identify and correct any misalignment. This can also be understood as a bi-level optimisation process that optimises user feedback as the top-level goal, and optimises grounded signals from the environment at the low level. In this way, a small amount of human data may facilitate a large amount of autonomous learning.
علاوة على ذلك، يمكن للمستخدمين تقديم تغذية راجعة أثناء عملية التعلّم، كمستوى رضاهم، يمكن استخدامها لضبط دالة المكافأة بدقة. وحينها يمكن لدالة المكافأة أن تتكيف عبر الزمن، لتحسين الطريقة التي تختار بها الإشارات أو تدمجها، ولتحديد أي اختلال في التوافق وتصحيحه. ويمكن فهم هذا أيضًا بوصفه عملية تحسين ثنائية المستوى، تُحسِّن تغذية المستخدم الراجعة كهدف على المستوى الأعلى، وتُحسِّن الإشارات المتجذرة من البيئة على المستوى الأدنى. وبهذه الطريقة، قد تُيسّر كمية صغيرة من البيانات البشرية قدرًا كبيرًا من التعلّم المستقل.
Planning and Reasoning
التخطيط والاستدلال
Will the era of experience change the way that agents plan and reason? Recently, there has been significant progress using LLMs that can reason, or "think" with language, by following a chain of thought before outputting a response. Conceptually, LLMs can act as a universal computer: an LLM can append tokens into its own context, allowing it to execute arbitrary algorithms before outputting a final result.
فهل سيغيّر عصر التجربة الطريقة التي يخطط بها العملاء ويستدلّون؟ لقد أُحرز مؤخرًا تقدم ملحوظ باستخدام نماذج لغة كبيرة قادرة على الاستدلال، أو "التفكير" باللغة، عبر اتباع سلسلة من الأفكار قبل إخراج استجابة. ومن الناحية المفاهيمية، يمكن لنماذج اللغة الكبيرة أن تعمل بمثابة حاسوب شامل: إذ يمكن للنموذج أن يُلحق رموزًا (tokens) بسياقه الخاص، مما يتيح له تنفيذ خوارزميات عشوائية قبل إخراج نتيجة نهائية.
In the era of human data, these reasoning methods have been explicitly designed to imitate human thought processes. For example, LLMs have been prompted to emit human-like chains of thought, imitate traces of human thinking, or to reinforce steps of thinking that match human examples. The reasoning process may be fine-tuned further to produce thinking traces that match the correct answer, as determined by human experts.
وفي عصر البيانات البشرية، صُمّمت طرائق الاستدلال هذه صراحةً لمحاكاة عمليات التفكير البشري. فعلى سبيل المثال، دُفعت نماذج اللغة الكبيرة لإصدار سلاسل أفكار شبيهة بالبشر، أو محاكاة آثار التفكير البشري، أو لتعزيز خطوات تفكير تطابق أمثلة بشرية. وقد يُضبط مسار الاستدلال بدقة أكبر لإنتاج آثار تفكير تطابق الإجابة الصحيحة، كما يحددها خبراء بشريون.
However, it is highly unlikely that human language provides the optimal instance of a universal computer. More efficient mechanisms of thought surely exist, using non-human languages that may for example utilise symbolic, distributed, continuous, or differentiable computations. A self-learning system can in principle discover or improve such approaches by learning how to think from experience. For example, AlphaProof learned to formally prove complex theorems in a manner quite different to human mathematicians.
غير أنه من المستبعد جدًا أن توفر اللغة البشرية أفضل مثال ممكن لحاسوب شامل. فمن المؤكد أن آليات تفكير أكثر كفاءة موجودة، تستخدم لغات غير بشرية قد تستعين مثلًا بحسابات رمزية أو موزَّعة أو مستمرة أو قابلة للاشتقاق. ويمكن لنظام يتعلم ذاتيًا، من حيث المبدأ، أن يكتشف أو يحسّن مثل هذه النهج بتعلّم كيفية التفكير من التجربة. فعلى سبيل المثال، تعلّم AlphaProof إثبات نظريات معقدة إثباتًا رسميًا بطريقة تختلف تمامًا عن طريقة الرياضيين البشر.
Furthermore, the principle of a universal computer only addresses the internal computation of the agent; it does not connect it to the realities of the external world. An agent trained to imitate human thoughts or even to match human expert answers may inherit fallacious methods of thought deeply embedded within that data, such as flawed assumptions or inherent biases. For example, if an agent had been trained to reason using human thoughts and expert answers from 5,000 years ago it may have reasoned about a physical problem in terms of animism; 1,000 years ago it may have reasoned in theistic terms; 300 years ago it may have reasoned in terms of Newtonian mechanics; and 50 years ago in terms of quantum mechanics. Progressing beyond each method of thought required interaction with the real world: making hypotheses, running experiments, observing results, and updating principles accordingly. Similarly, an agent must be grounded in real-world data in order to overturn fallacious methods of thought. This grounding provides a feedback loop, allowing the agent to test its inherited assumptions against reality and discover new principles that are not limited by current, dominant modes of human thought. Without this grounding, an agent, no matter how sophisticated, will become an echo chamber of existing human knowledge. To move beyond this, agents must actively engage with the world, collect observational data, and use that data to iteratively refine their understanding, mirroring in many ways the process that has driven human scientific progress.
علاوة على ذلك، لا يتناول مبدأ الحاسوب الشامل سوى الحساب الداخلي للعميل؛ فهو لا يربطه بحقائق العالم الخارجي. وقد يرث عميل دُرِّب على محاكاة الأفكار البشرية، أو حتى على مطابقة إجابات خبراء بشر، طرائق تفكير مغلوطة مغروسة بعمق في تلك البيانات، كافتراضات معيبة أو تحيزات كامنة. فعلى سبيل المثال، لو دُرِّب عميل على الاستدلال باستخدام أفكار بشرية وإجابات خبراء من قبل خمسة آلاف عام، لربما استدل على مسألة فيزيائية بمصطلحات إحيائية (أرواحية)؛ ولو كان ذلك قبل ألف عام، لاستدل بمصطلحات لاهوتية؛ ولو كان قبل ثلاثمئة عام، لاستدل بمصطلحات الميكانيكا النيوتونية؛ ولو كان قبل خمسين عامًا، لاستدل بمصطلحات ميكانيكا الكم. وقد تطلّب تجاوز كل طريقة من طرائق التفكير هذه تفاعلًا مع العالم الحقيقي: صياغة الفرضيات، وإجراء التجارب، ومراقبة النتائج، وتحديث المبادئ تبعًا لذلك. وبالمثل، يجب أن يتجذر العميل في بيانات العالم الحقيقي كي يتجاوز طرائق التفكير المغلوطة. ويوفر هذا التجذر حلقة تغذية راجعة، تتيح للعميل اختبار افتراضاته الموروثة في مواجهة الواقع، واكتشاف مبادئ جديدة لا تحدّها الأنماط السائدة الحالية للتفكير البشري. ودون هذا التجذر، سيصبح أي عميل، مهما بلغ من التطور، مجرد صدى لمعرفة بشرية قائمة. وللتغلب على ذلك، يجب أن يتفاعل العملاء بنشاط مع العالم، ويجمعوا بيانات ملاحَظة، ويستخدموا تلك البيانات لتنقيح فهمهم على نحو تكراري، بما يحاكي من نواحٍ عديدة العملية التي دفعت التقدم العلمي البشري.
One possible way to directly ground thinking in the external world is to build a world model that predicts the consequences of the agent's actions upon the world, including predicting reward. For example, a health assistant might consider making a recommendation for a local gym or a health podcast. The agent's world model might predict how a user's heart rate or sleep patterns might subsequently change following this action, as well as predicting future dialogue with the user. This allows the agent to plan directly in terms of its own actions and their causal effect upon the world. As the agent continues to interact with the world throughout its stream of experience, its dynamics model is continually updated to correct any errors in its predictions. Given a world model, an agent may apply scalable planning methods that improve the predicted performance of the agent.
ومن الطرائق الممكنة لتجذير التفكير مباشرة في العالم الخارجي بناء نموذج للعالم يتنبأ بعواقب أفعال العميل على العالم، بما في ذلك التنبؤ بالمكافأة. فعلى سبيل المثال، قد يفكر مساعد صحي في تقديم توصية بنادٍ رياضي محلي أو ببودكاست صحي. وقد يتنبأ نموذج العالم الخاص بالعميل بكيفية تغيّر معدل ضربات قلب المستخدم أو أنماط نومه لاحقًا نتيجة هذا الفعل، إضافة إلى التنبؤ بالحوار المستقبلي مع المستخدم. ويتيح هذا للعميل التخطيط مباشرة بدلالة أفعاله الخاصة وأثرها السببي على العالم. ومع استمرار تفاعل العميل مع العالم طوال تدفق تجربته، يُحدَّث نموذجه الديناميكي باستمرار لتصحيح أي أخطاء في تنبؤاته. وبوجود نموذج للعالم، يمكن للعميل تطبيق طرائق تخطيط قابلة للتوسع تحسّن الأداء المتوقع للعميل.
Planning and reasoning methods are not mutually exclusive: an agent may apply internal LLM computations to select each action during planning, or to simulate and evaluate the consequences of those actions.
وطرائق التخطيط والاستدلال ليست متنافية: فقد يطبّق العميل حسابات داخلية لنموذج اللغة الكبير لاختيار كل فعل أثناء التخطيط، أو لمحاكاة عواقب تلك الأفعال وتقييمها.
Why Now?
لماذا الآن؟
Learning from experience is not new. Reinforcement learning systems have previously mastered a large number of complex tasks that were represented in a simulator with a clear reward signal (c.f., approximately, the "era of simulation"). For example, RL methods equalled or exceeded human performance through self-play in board games such as backgammon, Go, chess, poker and Stratego; video games such as Atari, StarCraft II, Dota 2 and Gran Turismo; dextrous manipulation tasks such as Rubik's cube; and resource management tasks such as data center cooling. Furthermore, powerful RL agents such as AlphaZero exhibited impressive and potentially unlimited scalability with the size of the neural network, the quantity of interactive experience, and the duration of thinking time. However, agents based on this paradigm did not leap the gap between simulation (closed problems with singular, precisely defined rewards) to reality (open-ended problems with a plurality of seemingly ill-defined rewards).
التعلّم من التجربة ليس أمرًا جديدًا. فقد أتقنت أنظمة التعلّم المعزَّز سابقًا عددًا كبيرًا من المهام المعقدة التي مُثّلت في محاكٍ بإشارة مكافأة واضحة (قارن، تقريبًا، بما يُسمى "عصر المحاكاة"). فعلى سبيل المثال، ضاهت طرائق التعلّم المعزَّز الأداء البشري أو تجاوزته عبر اللعب الذاتي في ألعاب اللوح كالطاولة، والجو (Go)، والشطرنج، والبوكر، وستراتيغو؛ وألعاب الفيديو كأتاري، وستاركرافت 2، ودوتا 2، وغران توريزمو؛ ومهام التلاعب اليدوي البارع كمكعب روبيك؛ ومهام إدارة الموارد كتبريد مراكز البيانات. علاوة على ذلك، أظهرت عملاء تعلّم معزَّز قوية كـ AlphaZero قابلية مذهلة للتوسع، ربما غير محدودة، مع حجم الشبكة العصبية، وكمية التجربة التفاعلية، ومدة وقت التفكير. غير أن العملاء المبنية على هذا المذهب لم تتمكن من سد الفجوة بين المحاكاة (المسائل المغلقة ذات المكافآت الوحيدة المحددة بدقة) والواقع (المسائل المفتوحة ذات المكافآت المتعددة التي تبدو غير محدَّدة جيدًا).
The era of human data offered an appealing solution. Massive corpuses of human data contain examples of natural language for a huge diversity of tasks. Agents trained on this data achieved a wide range of competencies compared to the more narrow successes of the era of simulation. Consequently, the methodology of experiential RL was largely discarded in favour of more general-purpose agents, resulting in a widespread transition to human-centric AI.
وقدّم عصر البيانات البشرية حلًا جذابًا. فمدوّنات ضخمة من البيانات البشرية تحتوي على أمثلة من اللغة الطبيعية لتنوع هائل من المهام. وحققت العملاء المدرَّبة على هذه البيانات طائفة واسعة من الكفاءات مقارنة بالنجاحات الأضيق نطاقًا في عصر المحاكاة. ونتيجة لذلك، جرى التخلي إلى حد كبير عن منهجية التعلّم المعزَّز القائم على التجربة لصالح عملاء أعم غرضًا، مما أدى إلى تحول واسع النطاق نحو الذكاء الاصطناعي المتمحور حول الإنسان.
However, something was lost in this transition: an agent's ability to self-discover its own knowledge. For example, AlphaZero discovered fundamentally new strategies for chess and Go, changing the way that humans play these games. The era of experience will reconcile this ability with the level of task-generality achieved in the era of human data. This will become possible, as outlined above, when agents are able to autonomously act and observe in streams of real-world experience, and where the rewards may be flexibly connected to any of an abundance of grounded, real-world signals. The advent of autonomous agents that interact with complex, real-world action spaces, alongside powerful RL methods that can solve open-ended problems in rich reasoning spaces, suggests that the transition to the era of experience is imminent.
غير أن شيئًا ما ضاع في هذا التحول: قدرة العميل على اكتشاف معرفته الخاصة بنفسه. فعلى سبيل المثال، اكتشف AlphaZero استراتيجيات جديدة جذريًا للشطرنج والجو، غيّرت الطريقة التي يلعب بها البشر هاتين اللعبتين. وسيوفّق عصر التجربة بين هذه القدرة ومستوى عمومية المهام الذي تحقق في عصر البيانات البشرية. وسيصبح هذا ممكنًا، كما ورد أعلاه، حين يصبح العملاء قادرين على التصرف والمشاهدة بشكل مستقل في تدفقات من تجربة العالم الحقيقي، وحيث يمكن ربط المكافآت بمرونة بأي من وفرة الإشارات المتجذرة في العالم الحقيقي. ويشير ظهور عملاء مستقلين يتفاعلون مع مساحات أفعال معقدة من العالم الحقيقي، إلى جانب طرائق تعلّم معزَّز قوية قادرة على حل مسائل مفتوحة في مساحات استدلال غنية، إلى أن الانتقال إلى عصر التجربة بات وشيكًا.
Reinforcement Learning Methods
طرائق التعلّم المعزَّز
Reinforcement learning (RL) has a rich history that is deeply rooted in autonomous learning, where agents learn for themselves through direct interaction with their environment. Early RL research yielded a suite of powerful concepts and algorithms. For example, temporal difference learning enabled agents to estimate future rewards, leading to breakthroughs such as superhuman performance in backgammon. Exploration techniques, driven by optimism or curiosity, were developed to help agents discover creative new behaviors and avoid getting stuck in suboptimal routines. Methods like the Dyna algorithm enabled agents to build and learn from models of their world, allowing them to plan and reason about future actions. Concepts like options and inter/intra-option learning facilitated temporal abstraction, enabling agents to reason over longer timescales and break down complex tasks into manageable sub-goals.
للتعلّم المعزَّز تاريخ ثري متجذر بعمق في التعلّم المستقل، حيث يتعلم العملاء بأنفسهم عبر التفاعل المباشر مع بيئتهم. وأثمر البحث المبكر في التعلّم المعزَّز مجموعة من المفاهيم والخوارزميات القوية. فعلى سبيل المثال، مكّن التعلّم بالفارق الزمني العملاء من تقدير المكافآت المستقبلية، مما أفضى إلى اختراقات كأداء يفوق البشر في لعبة الطاولة. وطُوّرت تقنيات استكشاف، مدفوعة بالتفاؤل أو الفضول، لمساعدة العملاء على اكتشاف سلوكيات إبداعية جديدة وتجنّب الانحصار في روتينات دون المستوى الأمثل. كما مكّنت طرائق كخوارزمية Dyna العملاء من بناء نماذج لعالمهم والتعلّم منها، مما أتاح لهم التخطيط والاستدلال بشأن أفعال مستقبلية. وسهّلت مفاهيم كالخيارات والتعلّم بين الخيارات وداخلها التجريدَ الزمني، مما مكّن العملاء من الاستدلال عبر آفاق زمنية أطول وتفكيك المهام المعقدة إلى أهداف فرعية يمكن التعامل معها.
The rise of human-centric LLMs, however, shifted the focus away from autonomous learning and towards leveraging human knowledge. Techniques like RLHF (Reinforcement Learning from Human Feedback) and methods for aligning language models with human reasoning proved incredibly effective, driving rapid progress in AI capabilities. These approaches, while powerful, often bypassed core RL concepts: RLHF side-stepped the need for value functions by invoking human experts in place of machine-estimated values, strong priors from human data reduced the reliance on exploration, and reasoning in human-centric terms lessened the need for world models and temporal abstraction.
غير أن صعود نماذج اللغة الكبيرة المتمحورة حول الإنسان حوّل التركيز بعيدًا عن التعلّم المستقل نحو الاستفادة من المعرفة البشرية. وأثبتت تقنيات كالتعلم المعزَّز من التغذية الراجعة البشرية (RLHF)، وطرائق مواءمة نماذج اللغة مع الاستدلال البشري، فعالية هائلة، دافعةً تقدمًا سريعًا في قدرات الذكاء الاصطناعي. وهذه النهج، رغم قوتها، كثيرًا ما تجاوزت مفاهيم أساسية في التعلّم المعزَّز: فقد تجنب RLHF الحاجة إلى دوال القيمة عبر استدعاء خبراء بشر بدلًا من قيم تقدّرها الآلة، وقلّلت الاستدلالات القوية المسبقة من البيانات البشرية الاعتماد على الاستكشاف، وقلّل الاستدلال بمصطلحات متمحورة حول الإنسان الحاجة إلى نماذج للعالم والتجريد الزمني.
However, it could be argued that the shift in paradigm has thrown out the baby with the bathwater. While human-centric RL has enabled an unprecedented breadth of behaviours, it has also imposed a new ceiling on the agent's performance: agents cannot go beyond existing human knowledge. Furthermore, the era of human data has focused predominantly on RL methods that are designed for short episodes of ungrounded, human interaction, and are not suitable for long streams of grounded, autonomous interaction.
غير أنه يمكن القول إن هذا التحول في المذهب قد رمى بالطفل مع ماء الاستحمام. فبينما مكّن التعلّم المعزَّز المتمحور حول الإنسان من نطاق غير مسبوق من السلوكيات، فإنه فرض أيضًا سقفًا جديدًا على أداء العميل: إذ لا يمكن للعملاء تجاوز المعرفة البشرية القائمة. علاوة على ذلك، ركّز عصر البيانات البشرية أساسًا على طرائق تعلّم معزَّز مصممة لحلقات قصيرة من التفاعل البشري غير المتجذر، وهي غير مناسبة لتدفقات طويلة من التفاعل المستقل المتجذر.
The era of experience presents an opportunity to revisit and improve classic RL concepts. This era will bring new ways to think about reward functions that are flexibly grounded in observational data. It will revisit value functions and methods to estimate them from long streams with as yet incomplete sequences. It will bring principled yet practical methods for real-world exploration that discover new behaviours that are radically different from human priors. Novel approaches to world models will be developed that capture the complexities of grounded interactions. New methods for temporal abstraction will allow agents to reason, in terms of experience, over ever-longer time horizons. By building upon the foundations of RL and adapting its core principles to the challenges of this new era, we can unlock the full potential of autonomous learning and pave the way to truly superhuman intelligence.
ويتيح عصر التجربة فرصة لإعادة النظر في مفاهيم التعلّم المعزَّز الكلاسيكية وتحسينها. وسيجلب هذا العصر طرائق جديدة للتفكير في دوال المكافأة المتجذرة بمرونة في بيانات ملاحَظة. وسيعيد النظر في دوال القيمة وطرائق تقديرها من تدفقات طويلة ذات تسلسلات لم تكتمل بعد. وسيجلب طرائق مبدئية وعملية في آن للاستكشاف في العالم الحقيقي، تكتشف سلوكيات جديدة تختلف اختلافًا جذريًا عن الاستدلالات البشرية المسبقة. وستُطوَّر نهج جديدة لنماذج العالم تلتقط تعقيدات التفاعلات المتجذرة. وستتيح طرائق جديدة للتجريد الزمني للعملاء الاستدلال، بدلالة التجربة، عبر آفاق زمنية أطول فأطول. وبالبناء على أسس التعلّم المعزَّز وتكييف مبادئه الجوهرية مع تحديات هذا العصر الجديد، يمكننا إطلاق الإمكانات الكاملة للتعلّم المستقل، وتمهيد الطريق نحو ذكاء يفوق البشر فعلًا.
Consequences
التبعات
The advent of the era of experience, where AI agents learn from their interactions with the world, promises a future profoundly different from anything we have seen before. This new paradigm, while offering immense potential, also presents important risks and challenges that demand careful consideration, including but not limited to the following points.
يعِد حلول عصر التجربة، حيث تتعلم عملاء الذكاء الاصطناعي من تفاعلاتها مع العالم، بمستقبل مختلف اختلافًا عميقًا عن أي شيء رأيناه من قبل. وهذا المذهب الجديد، رغم ما يوفره من إمكانات هائلة، يطرح أيضًا مخاطر وتحديات مهمة تستدعي تأملًا دقيقًا، منها على سبيل المثال لا الحصر النقاط التالية.
On the positive side, experiential learning will unlock unprecedented capabilities. In everyday life, personalized assistants will leverage continuous streams of experience to adapt to individuals' health, educational, or professional needs towards long-term goals over the course of months or years. Perhaps most transformative will be the acceleration of scientific discovery. AI agents will autonomously design and conduct experiments in fields like materials science, medicine, or hardware design. By continuously learning from the results of their own experiments, these agents could rapidly explore new frontiers of knowledge, leading to the development of novel materials, drugs, and technologies at an unprecedented pace.
وعلى الجانب الإيجابي، سيطلق التعلّم من التجربة قدرات غير مسبوقة. ففي الحياة اليومية، ستستفيد المساعدات الشخصية من تدفقات مستمرة من التجربة للتكيف مع الاحتياجات الصحية أو التعليمية أو المهنية للأفراد سعيًا نحو أهداف طويلة المدى على مدى أشهر أو سنوات. ولعل أكثر ما سيُحدث تحولًا هو تسريع الاكتشاف العلمي. إذ ستصمم عملاء الذكاء الاصطناعي تجارب وتجريها بشكل مستقل في مجالات كعلم المواد أو الطب أو تصميم الأجهزة. وبالتعلّم المستمر من نتائج تجاربها الخاصة، يمكن لهذه العملاء استكشاف آفاق جديدة من المعرفة بسرعة، مما يفضي إلى تطوير مواد وأدوية وتقنيات جديدة بوتيرة غير مسبوقة.
However, this new era also presents significant and novel challenges. While the automation of human capabilities promises to boost productivity, these improvements could also lead to job displacement. Agents may even be able to exhibit capabilities previously considered the exclusive realm of humanity, such as long-term problem-solving, innovation, and a deep understanding of real world consequences.
غير أن هذا العصر الجديد يطرح أيضًا تحديات مهمة وجديدة. فبينما تعِد أتمتة القدرات البشرية بتعزيز الإنتاجية، فإن هذه التحسينات قد تؤدي أيضًا إلى فقدان الوظائف. بل قد تُظهر العملاء قدرات كانت تُعد سابقًا حكرًا على البشرية وحدها، كحل المشكلات طويل المدى، والابتكار، والفهم العميق لعواقب العالم الحقيقي.
Furthermore, whilst general concerns exist around the potential misuse of any AI, heightened risks may arise from agents that can autonomously interact with the world over extended periods of time to achieve long-term goals. By default, this provides fewer opportunities for humans to intervene and mediate the agent's actions, and therefore requires a high bar of trust and responsibility. Moving away from human data and human modes of thinking may also make future AI systems harder to interpret.
علاوة على ذلك، وبينما توجد مخاوف عامة بشأن إمكانية إساءة استخدام أي ذكاء اصطناعي، فقد تنشأ مخاطر متزايدة من عملاء قادرة على التفاعل بشكل مستقل مع العالم على مدى فترات زمنية ممتدة لتحقيق أهداف طويلة المدى. وهذا يتيح، من حيث المبدأ، فرصًا أقل للبشر للتدخل والتوسط في أفعال العميل، مما يستدعي معيارًا رفيعًا من الثقة والمسؤولية. كما أن الابتعاد عن البيانات البشرية وأنماط التفكير البشرية قد يجعل أنظمة الذكاء الاصطناعي المستقبلية أصعب في التفسير.
However, whilst acknowledging that experiential learning will increase certain safety risks, and that further research is surely required to ensure a safe transition into the era of experience, we should also recognise that it may also provide some important safety benefits.
غير أنه مع الإقرار بأن التعلّم من التجربة سيزيد بعض المخاطر الأمنية، وأن مزيدًا من البحث لازم بالتأكيد لضمان انتقال آمن إلى عصر التجربة، ينبغي أيضًا أن ندرك أنه قد يوفر أيضًا بعض المنافع الأمنية المهمة.
Firstly, an experiential agent is aware of the environment it is situated within, and its behaviour can adapt over time to changes in that environment. Any pre-programmed system, including a fixed AI system, can be unaware of its environmental context, and become maladapted to the changing world into which it is deployed. For example, a critical piece of hardware may malfunction, a pandemic might cause rapid societal change, or a new scientific discovery may trigger a cascade of rapid technological developments. By contrast, an experiential agent could observe and learn to circumvent malfunctioning hardware, adjust to rapid societal change, or embrace and build upon new science and technology. Perhaps even more importantly, the agent could recognise when its behaviour is triggering human concern, dissatisfaction, or distress, and adaptively modify its behaviour to avoid these negative consequences.
أولًا، يكون العميل الذي يتعلم من التجربة مدركًا للبيئة التي يوجد فيها، ويمكن لسلوكه أن يتكيف عبر الزمن مع التغيرات في تلك البيئة. أما أي نظام مبرمج سلفًا، بما في ذلك نظام ذكاء اصطناعي ثابت، فقد يكون غافلًا عن سياقه البيئي، ويصبح غير متكيف مع العالم المتغير الذي يُنشر فيه. فعلى سبيل المثال، قد يتعطل جزء بالغ الأهمية من الأجهزة، أو قد يتسبب وباء في تغيّر مجتمعي سريع، أو قد يطلق اكتشاف علمي جديد سلسلة من التطورات التقنية السريعة. وفي المقابل، يمكن لعميل يتعلم من التجربة أن يلاحظ ويتعلم كيفية تجاوز عطل الأجهزة، أو التكيف مع التغير المجتمعي السريع، أو تبنّي علم وتقنية جديدين والبناء عليهما. ولعل الأهم من ذلك، أن العميل قد يدرك متى يثير سلوكه قلق البشر أو استياءهم أو ضيقهم، فيعدّل سلوكه بشكل تكيفي لتجنب هذه العواقب السلبية.
Secondly, the agent's reward function may itself be adapted through experience, for example using the bi-level optimisation described earlier. Importantly, this means that misaligned reward functions can often be incrementally corrected over time by trial and error. For example, rather than blindly optimising a signal, such as the maximisation of paperclips, the reward function could be modified, based upon indications of human concern, before paperclip production consumes all of the Earth's resources. This is analogous to the way that humans set goals for each other, and then adapt those goals if they observe people gaming the system, neglecting long-term well-being, or causing undesired negative consequences; although also like human goal-setting, there is no guarantee of perfect alignment.
ثانيًا، يمكن أن تتكيف دالة مكافأة العميل نفسها عبر التجربة، مثلًا باستخدام التحسين ثنائي المستوى الموصوف آنفًا. والمهم هنا أن هذا يعني أن دوال المكافأة غير المتوائمة يمكن غالبًا تصحيحها تدريجيًا عبر الزمن بالتجربة والخطأ. فعلى سبيل المثال، بدلًا من تحسين إشارة بشكل أعمى، كتعظيم إنتاج مشابك الورق، يمكن تعديل دالة المكافأة، استنادًا إلى مؤشرات قلق بشري، قبل أن يستهلك إنتاج مشابك الورق كل موارد الأرض. ويشبه هذا الطريقة التي يضع بها البشر أهدافًا لبعضهم بعضًا، ثم يكيّفون تلك الأهداف إذا لاحظوا أشخاصًا يتلاعبون بالنظام، أو يهملون الرفاه طويل المدى، أو يتسببون في عواقب سلبية غير مرغوبة؛ رغم أنه، كما هو الحال في وضع الأهداف البشرية، لا يوجد ضمان لتوافق كامل.
Finally, advancements relying on physical experience are inherently constrained by the time it takes to execute actions in the real world and observe their consequences. For example, the development of a new drug, even with AI-assisted design, still requires real-world trials that cannot be completed overnight. This may provide a natural brake on the pace of potential AI self-improvement.
وأخيرًا، فإن التطورات المعتمدة على التجربة الفيزيائية مقيّدة بطبيعتها بالوقت اللازم لتنفيذ الأفعال في العالم الحقيقي ومراقبة عواقبها. فعلى سبيل المثال، لا يزال تطوير دواء جديد، حتى بمساعدة تصميم بالذكاء الاصطناعي، يتطلب تجارب في العالم الحقيقي لا يمكن إتمامها بين ليلة وضحاها. وقد يوفر هذا كابحًا طبيعيًا لوتيرة التحسين الذاتي المحتمل للذكاء الاصطناعي.
Conclusion
الخاتمة
The era of experience marks a pivotal moment in the evolution of AI. Building on today's strong foundations, but moving beyond the limitations of human-derived data, agents will increasingly learn from their own interactions with the world. Agents will autonomously interact with environments through rich observations and actions. They will continue to adapt over the course of lifelong streams of experience. Their goals will be directable towards any combination of grounded signals. Furthermore, agents will utilise powerful non-human reasoning, and construct plans that are grounded in the consequences of the agent's actions upon its environment. Ultimately, experiential data will eclipse the scale and quality of human generated data. This paradigm shift, accompanied by algorithmic advancements in RL, will unlock in many domains new capabilities that surpass those possessed by any human.
يمثّل عصر التجربة لحظة محورية في تطور الذكاء الاصطناعي. فبالبناء على أسس اليوم القوية، لكن مع تجاوز حدود البيانات المستمدة من البشر، ستتعلم العملاء على نحو متزايد من تفاعلاتها الخاصة مع العالم. وستتفاعل العملاء بشكل مستقل مع البيئات عبر مشاهدات وأفعال غنية. وستواصل التكيف على مدى تدفقات تجربة تمتد مدى الحياة. وستكون أهدافها قابلة للتوجيه نحو أي مزيج من الإشارات المتجذرة. علاوة على ذلك، ستستخدم العملاء استدلالًا غير بشري قويًا، وتبني خططًا متجذرة في عواقب أفعال العميل على بيئته. وفي نهاية المطاف، ستتجاوز بيانات التجربة حجم وجودة البيانات التي يولّدها البشر. وهذا التحول في المذهب، مصحوبًا بتطورات خوارزمية في التعلّم المعزَّز، سيطلق في مجالات عديدة قدرات جديدة تفوق ما يملكه أي إنسان.
Acknowledgements
شكر وتقدير
The authors would like to acknowledge helpful comments and discussion from Thomas Degris, Rohin Shah, Tom Schaul and Hado van Hasselt.
يودّ المؤلفان أن يشكرا كلًا من توماس ديغري، وروهين شاه، وتوم شول، وهادو فان هاسيلت على ملاحظاتهم ونقاشاتهم المفيدة.
Preprint chapter for "Designing an Intelligence" (MIT Press) | فصل أولي من كتاب "تصميم الذكاء" (منشورات معهد ماساتشوستس للتقنية)