المستوى: متوسط · الدرس 10 من 14
أهداف الدرس
- أن تبني مساعداً يجيب من مستنداتك.
- أن تفهم التقطيع والتضمين والاسترجاع.
RAG يعطي النموذج «كتاباً مفتوحاً»: نقطّع المستندات، ونحوّل كل قطعة إلى تضمين، ونخزّنها. ومع كل سؤال نسترجع أقرب القطع ونضعها في السياق، فيجيب النموذج منها ويذكرها.
from sentence_transformers import SentenceTransformer
import numpy as np
docs = [ # example documents
"يبدأ التسجيل للفصل الثاني يوم الأحد الأول من فبراير.",
"الحد الأقصى للغياب 15٪ من ساعات المساق.",
"تُسلَّم مشاريع التخرج قبل نهاية الأسبوع الرابع عشر.",
]
model = SentenceTransformer("intfloat/multilingual-e5-small")
doc_vecs = model.encode(["passage: " + d for d in docs], normalize_embeddings=True)
question = "متى يبدأ التسجيل؟"
q_vec = model.encode("query: " + question, normalize_embeddings=True)
best = int(np.argmax(doc_vecs @ q_vec))
prompt = f"أجب من المصدر فقط واذكره.\nالمصدر: {docs[best]}\nالسؤال: {question}"
print(prompt)| قرار | نصيحة |
|---|---|
| حجم القطعة | فقرة أو فقرتان (نحو 200–500 كلمة) |
| التداخل | جملة أو جملتان بين القطع |
| البيانات الوصفية | العنوان والتاريخ والمصدر مع كل قطعة |
| عدد النتائج | 3–5 قطع في السياق |
علّم المساعد أن يقول «لم أجد هذا في المصادر» بدل أن يخترع إجابة.
تمرين
لماذا نضع «passage:» و«query:» قبل النص في هذا النموذج؟
الإجابة
لأنه تدرّب هكذا: يميّز بين النص المخزّن والسؤال، فيتحسّن الاسترجاع.
Level: Intermediate · Lesson 10 of 14
Lesson goals
- Build an assistant that answers from your documents.
- Understand chunking, embedding and retrieval.
RAG gives the model an «open book»: we chunk the documents, turn each chunk into an embedding and store it. For each question we retrieve the closest chunks and put them in the context, so the model answers from them and cites them.
from sentence_transformers import SentenceTransformer
import numpy as np
docs = [ # example documents
"يبدأ التسجيل للفصل الثاني يوم الأحد الأول من فبراير.",
"الحد الأقصى للغياب 15٪ من ساعات المساق.",
"تُسلَّم مشاريع التخرج قبل نهاية الأسبوع الرابع عشر.",
]
model = SentenceTransformer("intfloat/multilingual-e5-small")
doc_vecs = model.encode(["passage: " + d for d in docs], normalize_embeddings=True)
question = "متى يبدأ التسجيل؟"
q_vec = model.encode("query: " + question, normalize_embeddings=True)
best = int(np.argmax(doc_vecs @ q_vec))
prompt = f"أجب من المصدر فقط واذكره.\nالمصدر: {docs[best]}\nالسؤال: {question}"
print(prompt)| Decision | Advice |
|---|---|
| Chunk size | One or two paragraphs (about 200–500 words) |
| Overlap | A sentence or two between chunks |
| Metadata | Title, date and source with each chunk |
| How many results | 3–5 chunks in the context |
Teach the assistant to say «I didn't find this in the sources» instead of inventing an answer.
Exercise
Why prefix «passage:» and «query:» for this model?
Answer
It was trained that way: it tells stored text from questions, which improves retrieval.
التعليقات / Comments
لا تعليقات بعد. كن أول من يسأل. / No comments yet. Be the first to ask.