Image-to-Text Endpoints
Curated picks for OCR, captioning/VQA, and detection/segmentation. Moondream 3 is the dominant pick across all three; Florence-2 and SAM-3 complete the toolset. Verify with genmedia models --endpoint_id <id> --json before running.
OCR, extract text from image
fal-ai/got-ocr/v2: GOT OCR 2.0fal-ai/florence-2-large/ocr: Florence-2 Large (OCR head)fal-ai/moondream3-preview/segment: Moondream 3 Preview (segment also reads text regions)fal-ai/moondream3-preview/query: Moondream 3 Preview (query for text content)
Caption / VQA
Image description and visual question-answering.
fal-ai/moondream3-preview/caption: Moondream 3 · Captionfal-ai/moondream3-preview/query: Moondream 3 · Query (VQA)fal-ai/florence-2-large/caption: Florence-2 Largefal-ai/florence-2-large/detailed-caption: Florence-2 Large · Detailedfal-ai/florence-2-large/more-detailed-caption: Florence-2 Large · More Detailedfal-ai/video-understanding: Video Understandingfal-ai/auto-caption: Auto-Captionerperceptron/isaac-01: Perceptron · Isaac 0.1perceptron/isaac-01/openai/v1/chat/completions: Perceptron · Isaac 0.1 (OpenAI-compatible)
Detection / Segmentation
Nesne tespit ve maskeleme.
fal-ai/moondream3-preview/detect: Moondream 3 · Detect (open-vocabulary detection)fal-ai/moondream3-preview/point: Moondream 3 · Pointfal-ai/moondream2/object-detection: Moondream 2 · Object Detectionfal-ai/moondream2/point-object-detection: Moondream 2 · Point Object Detectionfal-ai/sam-3/image/embed: SAM 3 · Image Embed (segmentation backbone)fal-ai/florence-2-large/region-to-category: Florence-2 · Region-to-Categoryfal-ai/florence-2-large/region-to-description: Florence-2 · Region-to-Descriptionperceptron/isaac-01: Perceptron · Isaac 0.1perceptron/isaac-01/openai/v1/chat/completions: Perceptron · Isaac 0.1 (OpenAI-compatible)
Common parameters
genmedia schema fal-ai/moondream3-preview/query --json
genmedia schema fal-ai/got-ocr/v2 --json
genmedia schema fal-ai/sam-3/image/embed --json
Frequently exposed:
image_url: source imageprompt/query/question, for VQA or guided segmentationthreshold: confidence cutoff (detection)output_format: for masks:pngalpha,binary,coco-rle, etc.
Discovery
genmedia models --category vision --limit 10 --json
genmedia models "ocr" --json
genmedia models "image segmentation" --json
genmedia docs "vision" --json
See also
- For mask manipulation utilities, see fal-workflow/references/utility-endpoints.md
- For document scan cleanup before OCR, see fal-recipes/references/image-restoration.md