المستوى: مبتدئ · الدرس 4 من 14
أهداف الدرس
- أن تعرف كيف يتعلّم النموذج من النصوص الخام.
- أن تفهم قوانين التوسّع ومزيج الخبراء.
في التدريب المسبق يقرأ النموذج كمّاً هائلاً من النصوص — ويب وكتب وكود — ويتعلّم مهمة واحدة: توقّع الرمز التالي. كلما أخطأ زادت «الخسارة»، فتُعدَّل أوزانه قليلاً. وبعد تريليونات الرموز يكتسب اللغة والمعرفة العامة.
import numpy as np
vocab = ["الطالب", "يقرأ", "الكتاب", "البحر"]
# The model's probabilities for the word after «الطالب يقرأ»
probs = np.array([0.05, 0.05, 0.80, 0.10])
target = vocab.index("الكتاب")
loss = -np.log(probs[target])
print(round(float(loss), 3)) # 0.223 — low loss: the model expected «الكتاب»| المفهوم | الفكرة |
|---|---|
| قوانين التوسّع | مزيد من البيانات والحجم والحوسبة يعطي تحسّناً يمكن توقّعه |
| مزيج الخبراء | أجزاء متخصّصة يعمل بعضها فقط لكل رمز: قدرة كبيرة بكلفة أقل |
| جودة البيانات | تنظيف، وإزالة التكرار والبيانات الشخصية، وتوازن اللغات |
النموذج بعد التدريب المسبق وحده «يكمل النص» ولا يتبع التعليمات؛ هذا ما يعالجه التدريب اللاحق.
تمرين
لماذا تهمّ نسبة العربية في بيانات التدريب المسبق؟
الإجابة
لأن النموذج يتعلّم اللغة بقدر ما يرى منها؛ قلّة العربية تعني فهماً أضعف ورموزاً أكثر.
Level: Beginner · Lesson 4 of 14
Lesson goals
- Know how a model learns from raw text.
- Understand scaling laws and mixture of experts.
In pre-training the model reads an enormous amount of text — web, books, code — and learns one task: predict the next token. Every miss raises the «loss», and its weights are nudged. After trillions of tokens it has picked up language and general knowledge.
import numpy as np
vocab = ["الطالب", "يقرأ", "الكتاب", "البحر"]
# The model's probabilities for the word after «الطالب يقرأ»
probs = np.array([0.05, 0.05, 0.80, 0.10])
target = vocab.index("الكتاب")
loss = -np.log(probs[target])
print(round(float(loss), 3)) # 0.223 — low loss: the model expected «الكتاب»| Concept | The idea |
|---|---|
| Scaling laws | More data, size and compute give a predictable improvement |
| Mixture of experts | Specialised parts, only some active per token: big capacity, lower cost |
| Data quality | Cleaning, de-duplication, removing personal data, balancing languages |
After pre-training alone a model «continues text» and does not follow instructions; post-training fixes that.
Exercise
Why does the share of Arabic in pre-training data matter?
Answer
A model learns a language as much as it sees it; little Arabic means weaker understanding and more tokens.
التعليقات / Comments
لا تعليقات بعد. كن أول من يسأل. / No comments yet. Be the first to ask.