المستوى: متوسط · الدرس 7 من 14
أهداف الدرس
- أن تفهم «التفكير قبل الإجابة».
- أن تعرف كيف تُدرَّب بالمكافآت التي يمكن التحقّق منها.
نماذج الاستدلال تكتب خطوات تفكير قبل الإجابة، وتصرف وقتاً أطول على المسائل الصعبة. تُدرَّب بالتعلّم المعزّز على مسائل يمكن التحقّق من إجابتها آلياً — رياضيات وبرمجة — فتُكافَأ حين تصيب. ومن الطرق الشائعة أن تُجرَّب عدّة إجابات للسؤال نفسه ويُقارَن كلٌّ منها بمتوسّط المجموعة.
import numpy as np
def reward(answer: str, correct: str) -> float:
return 1.0 if answer.strip() == correct else 0.0 # verifiable: no human needed
answers = ["42", "41", "42", "40"] # 4 tries at the same question
r = np.array([reward(a, "42") for a in answers])
advantage = (r - r.mean()) / (r.std() + 1e-6) # better than the group -> pushed up
print(r, advantage.round(2)) # [1. 0. 1. 0.] [ 1. -1. 1. -1.]| المهمة | نموذج سريع | نموذج استدلال |
|---|---|---|
| ترجمة أو تلخيص | ✓ | |
| مسألة رياضيات متعدّدة الخطوات | ✓ | |
| تصحيح كود وتحليل خطأ | ✓ |
اطلب من النموذج أن يتحقّق من إجابته قبل أن يعطيها؛ هذا يحسّن الدقّة كثيراً في المسائل الصعبة.
تمرين
لماذا تُدرَّب نماذج الاستدلال على الرياضيات والبرمجة خصوصاً؟
الإجابة
لأن صحة الإجابة فيها يمكن التحقّق منها آلياً، فتكون المكافأة دقيقة وبلا حاجة إلى إنسان.
Level: Intermediate · Lesson 7 of 14
Lesson goals
- Understand «thinking before answering».
- Know how they are trained with verifiable rewards.
Reasoning models write out steps before answering, and spend longer on hard problems. They are trained with reinforcement learning on problems whose answers can be checked automatically — maths and code — and are rewarded when right. A common method samples several answers to one question and compares each with the group average.
import numpy as np
def reward(answer: str, correct: str) -> float:
return 1.0 if answer.strip() == correct else 0.0 # verifiable: no human needed
answers = ["42", "41", "42", "40"] # 4 tries at the same question
r = np.array([reward(a, "42") for a in answers])
advantage = (r - r.mean()) / (r.std() + 1e-6) # better than the group -> pushed up
print(r, advantage.round(2)) # [1. 0. 1. 0.] [ 1. -1. 1. -1.]| Task | Fast model | Reasoning model |
|---|---|---|
| Translate or summarise | ✓ | |
| Multi-step maths problem | ✓ | |
| Debugging and error analysis | ✓ |
Ask the model to check its answer before giving it; that improves accuracy a lot on hard problems.
Exercise
Why are reasoning models trained especially on maths and code?
Answer
Their answers can be checked automatically, so the reward is exact and needs no human.
التعليقات / Comments
لا تعليقات بعد. كن أول من يسأل. / No comments yet. Be the first to ask.