Nawah-VL Detection — ما الأجسام الموجودة في الصورة؟
Give it an image; it names every object it finds, in Arabic, each with a bounding box. Nothing to type — the prompt is fixed.
Running on CPU: about 4–6 s per image, and ~12 s for the very first request while the model loads. Generation is autoregressive, so a crowded picture takes a little longer than an empty one.
What to expect — read the second column first.
| model | image-blind floor | |
|---|---|---|
| det_f1@0.5 | 0.487 | 0.010 |
| label_f1 (naming only) | 0.669 | 0.129 |
| mean IoU on matched pairs | 0.830 | 0.587 |
A hit needs the Arabic name and IoU ≥ 0.5. Emitting the 4 commonest labels at their mean box, never looking at the image, scores 0.010 — so unlike referring grounding, nearly all of this headline is real signal. Naming alone is the exception: it has a 0.129 floor.
It sees big things better than small ones. Recall@0.5 by how much of the frame the object fills: large → 0.770 , medium → 0.711 , small → 0.553 , tiny → 0.294.
It predicts 3.267 objects per image against 2.901 in the gold, and repeats an identical line 0.082 of the time — collapsing those repeats would score 0.509.
The labels come from a YOLO detector over Open Images' 601 classes, not from people, so its vocabulary and its blind spots are inherited.
Examples are held-out rows the model never trained on, spanning the object-count and size range on purpose — the crowded ones are genuinely hard.