AMAN
Toxic speech detection in Moroccan Darija
A multi-label NLP system that detects toxic speech in Moroccan Darija, with a large teacher model distilled into a compact, faster student.
- Context
- Supervised team project · 4 members
- Year
- 2026
- Language
- Moroccan Darija
- Task
- Multi-label classification
Macro F1, model by model
Macro F1 averages the F1 score of every toxicity label, so rare labels count as much as frequent ones.
- 010.623XLM-R teacherThe starting point: a multilingual XLM-RoBERTa fine-tuned on the comment corpus.
- 020.815Darija teacherA second teacher, trained for Darija, then distilled into the student.
- 030.895DistilBERT studentThe final, compact model, distilled from the Darija teacher.
- Smaller model
- 4.2×
- Reported size of the DistilBERT student, compared with its teacher.
- Faster model
- 4×
- Reported speed of the DistilBERT student, compared with its teacher.
- Comments
- ~160k
- Comments the XLM-RoBERTa teacher was fine-tuned on.
01Context
Moroccan Darija is a dialect with no native dataset for this task. AMAN is a supervised team project from the AI & Data Science curriculum that treats toxic-speech detection as a multi-label problem: one comment can be toxic in several ways at once.
02Problem
A large multilingual transformer is a natural starting point, but it is slow and heavy to serve. The goal: accurate multi-label detection in Darija with a model small and fast enough to deploy.
03Architecture
01Data
Comment corpus
Teacher training data
02Teacher
XLM-RoBERTa
Multilingual, fine-tuned
03Teacher
Darija teacher
Trained for Darija
04Transfer
Knowledge distillation
Teacher → student
05Student
DistilBERT
Compact and faster
Macro F1 0.895
06Serving
FastAPI + Docker
06App
Streamlit inference
04Team
Team project (4 members) focused on Moroccan Darija toxic-speech detection using transformer fine-tuning, knowledge distillation and an inference application. The work progressed from XLM-RoBERTa and a Darija teacher to a distilled DistilBERT student.
05Technical decisions
- 01
Start from a multilingual model
The first teacher is a multilingual XLM-RoBERTa, fine-tuned on the comment corpus.
- 02
A Darija-focused teacher
Before distillation, a second teacher trained for Darija improved on the initial multilingual one.
- 03
Teacher–student distillation
The Darija teacher’s knowledge is distilled into a compact DistilBERT student.
- 04
Multi-label output
Each comment can carry several labels at once, rather than a single toxic / not-toxic verdict.
06Results
- Macro F1 improved at every step, from the first teacher to the distilled student
- DistilBERT student — reported as 4.2× smaller and 4× faster than its teacher
- Served through FastAPI and Docker, with a Streamlit inference app
07Stack
- Python
- PyTorch
- Hugging Face Transformers
- scikit-learn
- FastAPI
- Docker
- Streamlit
08GitHub
Building something similar?
Start a project