Nawah-VL Detection — ما الأجسام الموجودة في الصورة؟

Give it an image; it names every object it finds, in Arabic, each with a bounding box. Nothing to type — the prompt is fixed.

Running on CPU: about 4–6 s per image, and ~12 s for the very first request while the model loads. Generation is autoregressive, so a crowded picture takes a little longer than an empty one.

What to expect — read the second column first.

model image-blind floor
det_f1@0.5 0.487 0.010
label_f1 (naming only) 0.669 0.129
mean IoU on matched pairs 0.830 0.587

A hit needs the Arabic name and IoU ≥ 0.5. Emitting the 4 commonest labels at their mean box, never looking at the image, scores 0.010 — so unlike referring grounding, nearly all of this headline is real signal. Naming alone is the exception: it has a 0.129 floor.

It sees big things better than small ones. Recall@0.5 by how much of the frame the object fills: large → 0.770 , medium → 0.711 , small → 0.553 , tiny → 0.294.

It predicts 3.267 objects per image against 2.901 in the gold, and repeats an identical line 0.082 of the time — collapsing those repeats would score 0.509.

The labels come from a YOLO detector over Open Images' 601 classes, not from people, so its vocabulary and its blind spots are inherited.

Examples are held-out rows the model never trained on, spanning the object-count and size range on purpose — the crowded ones are genuinely hard.

جرب واحدة من دول