المستوى: متوسط · الدرس 6 من 14
أهداف الدرس
- أن تفهم RLHF وDPO.
- أن تعرف لماذا تجعل التفضيلات الإجابات أفضل.
بعد الضبط بالتعليمات نعرض إجابتين للسؤال نفسه ونقول أيّهما أفضل. يتعلّم النموذج من آلاف المقارنات أن يميل إلى الإجابات الأنفع والأصدق والأكثر أماناً.
import numpy as np
def dpo_loss(lp_chosen, lp_rejected, ref_chosen, ref_rejected, beta=0.1):
# How much more the model prefers the chosen answer than the reference model does
margin = (lp_chosen - ref_chosen) - (lp_rejected - ref_rejected)
return float(np.log1p(np.exp(-beta * margin))) # = -log(sigmoid(beta * margin))
print(round(dpo_loss(-12.0, -15.0, -13.0, -14.0), 4)) # 0.5981| RLHF | DPO | |
|---|---|---|
| نموذج مكافأة | نعم | لا |
| التعقيد | أعلى | أبسط |
| الشيوع اليوم | في المختبرات الكبيرة | واسع، ومفتوح المصدر |
يمكن أن يأتي التفضيل من نموذج آخر يحكم وفق مبادئ مكتوبة بدل البشر وحدهم؛ وهذا يوسّع البيانات ويثبّت القيم.
تمرين
أيّ الإجابتين تفضّل لطالب سأل «كيف أحسّن معدّلي؟»: نصيحة عامة أم خطوات محدّدة بمواعيد؟ ولماذا؟
الإجابة
الخطوات المحدّدة: أنفع وقابلة للتنفيذ. هذه هي الإشارة التي يتعلّمها النموذج من التفضيلات.
Level: Intermediate · Lesson 6 of 14
Lesson goals
- Understand RLHF and DPO.
- Know why preferences make answers better.
After instruction tuning we show two answers to the same question and say which is better. From thousands of comparisons the model learns to lean towards answers that are more helpful, honest and safe.
import numpy as np
def dpo_loss(lp_chosen, lp_rejected, ref_chosen, ref_rejected, beta=0.1):
# How much more the model prefers the chosen answer than the reference model does
margin = (lp_chosen - ref_chosen) - (lp_rejected - ref_rejected)
return float(np.log1p(np.exp(-beta * margin))) # = -log(sigmoid(beta * margin))
print(round(dpo_loss(-12.0, -15.0, -13.0, -14.0), 4)) # 0.5981| RLHF | DPO | |
|---|---|---|
| Reward model | Yes | No |
| Complexity | Higher | Simpler |
| Common today | In large labs | Widespread, open source |
Preferences can come from another model judging by written principles, not only from people; that scales data and keeps values steady.
Exercise
Which do you prefer for «how do I raise my GPA?»: general advice or concrete steps with dates? Why?
Answer
Concrete steps: more useful and actionable. That is the signal a model learns from preferences.
التعليقات / Comments
لا تعليقات بعد. كن أول من يسأل. / No comments yet. Be the first to ask.