Image-to-Text Endpoints

Curated picks for OCR, captioning/VQA, and detection/segmentation. Moondream 3 is the dominant pick across all three; Florence-2 and SAM-3 complete the toolset. Verify with genmedia models --endpoint_id <id> --json before running.

OCR, extract text from image

Caption / VQA

Image description and visual question-answering.

Detection / Segmentation

Nesne tespit ve maskeleme.

Common parameters

genmedia schema fal-ai/moondream3-preview/query --json
genmedia schema fal-ai/got-ocr/v2 --json
genmedia schema fal-ai/sam-3/image/embed --json

Frequently exposed:

Discovery

genmedia models --category vision --limit 10 --json
genmedia models "ocr" --json
genmedia models "image segmentation" --json
genmedia docs "vision" --json

See also