Image-to-Video Endpoints

Curated picks across 6 use cases. Seedance 2.0 dominates for general I2V; Kling O3 / V3 specializes in reference-to-video and 4K; Sora 2 is included for I2V. Avatar/lipsync has its own large bucket. Verify with genmedia models --endpoint_id <id> --json before running.

Premium realism

Final-quality image-to-video.

Fast / cheap

Economical / fast I2V.

First/last frame interpolation

Controlled transition between start and end frames.

Reference-to-video

Multiple reference images (person / style / element) → video.

4K capable

Endpoints with native 4K output.

Avatar / talking head / lipsync

Talking head, avatar, lip-sync video. The widest bucket in this modality, models differ meaningfully in input requirements, motion style, and language support.

For the multi-step TTS → lipsync recipe, see fal-recipes/references/character-lipsync.md.

Family-specific prompting

Discovery

genmedia models --category image-to-video --limit 10 --json
genmedia docs "image to video" --json