Models

Browse AI models. Click any model to open its detail page with playground, API, and JSON docs.

Model Picker Visualization

This dropdown demonstrates the new grouping strategy (Top-Level Intent + Subcategory). Models with multiple subcategories will appear in each relevant group.

NameCategorySubcategoryInput TypeOutput TypeTagsOpen
Apply Speed Curve
renderer-ease-curve
Apply cinematic speed curves (Ease, Bezier) to video clips with audio sync.
clip-editor
editing
video
video
slow motionslomospeed ramptime warpcinematicvideo edit
View
Claude Sonnet 4.6
claude-sonnet-4-6
SOTA reasoning, coding, and long-context analysis (1M+ tokens).
text
chat
textimageaudio
text
chatcodingreasoningplanningagentcomplex
View
Compress Video
renderer-video-downsize
Compress video to a specific target file size using FFmpeg.
clip-editor
editing
video
video
compressshrinkoptimizefile sizemp4ffmpeg
View
Extract Frame
renderer-extract-frame
Extract the first frame, last frame, or a frame at an exact non-negative timestamp from a video without re-encoding.
clip-editor
editing
video
image
thumbnailpreviewframe extractionpostercapture
View
Loop Video
renderer-video-loop
Repeat a video to a target duration or loop count.
clip-editor
editing
video
video
videolooprepeatffmpeg
View
Merge Audio-Video
renderer-audio-video-merge
Utility for merging video and audio streams into a single high-quality file.
clip-editor
editing
videoaudio
video
mergeaudio-videoffmpegmuxingeditor
View
Overlay / Watermark
renderer-video-overlay
Overlay an image or logo on top of a video.
clip-editor
editing
videoimage
video
videooverlaywatermarklogoimageffmpeg
View
Reverse Video
renderer-video-reverse
Reverse video and audio when audio is present.
clip-editor
editing
video
video
videoreverserewindffmpeg
View
Slideshow Video
renderer-image-slideshow
Create a video from one or more images with optional audio.
clip-editor
editing
imageaudio
video
imagevideoslideshowffmpeg
View
Stitch Videos
renderer-video-merge
Append videos using the first clip canvas, with optional transitions and explicit fit behavior.
clip-editor
editing
video
video
audio-videoeditorffmpegmergemuxingtransition
View
Trim Video
renderer-video-trim
Cuts a clip to a start/end range. Deterministic FFmpeg operation — no generation, no cost variance.
clip-editor
editing
video
video
video-edittrimdeterministicffmpeg
View
Video to WebP
renderer-video-to-webp
Convert a video segment to animated WebP.
clip-editor
editing
video
image
videowebpanimatedffmpeg
View
Gemini 3.5 Flash-Lite
gemini-3-5-flash-lite-vertex
Fast, cost-efficient GA Gemini model for high-throughput text, document, image, audio, and video understanding.
text
chat
textimageaudiovideo
text
geminigooglegaflash-litemultimodalhigh-throughput
View
Runway Gen 4.5
runway-gen-4-5
Professional AI video generation and editing tools for cinematic production.
video
Text to Video
textimage
video
runwayvideo gencinematicvfxcreative
View
Gemini 3.6 Flash
gemini-3-6-flash-vertex
GA Gemini Flash model for coding, multimodal reasoning, knowledge work, and multi-step workflows.
text
chat
textimageaudiovideo
text
geminigooglegaflashmultimodalreasoning
View
Kling v3
kling-v3-video
Professional cinematic video generation with advanced motion and temporal consistency.
video
Text to Video
textimage
video
agent:canvasagent:workflowcapability:image-to-videocapability:text-to-videocinematiccost:highhigh-fidelityklingmotionquality:premiumselection:primaryuse-case:cinematic-videouse-case:motion-controlvideo gen
View
Kling V3 Omni
kling-v3-omni-video
Professional cinematic video generation with advanced motion and temporal consistency.
video
Text to Video
textimagevideoaudio
video
agent:canvasagent:workflowcapability:audio-conditioned-videocapability:image-to-videocapability:text-to-videocapability:video-editcinematiccost:highhigh-fidelityklingmotionquality:premiumselection:primaryuse-case:cinematic-videovideo gen
View
Nano Banana 2 Lite (G)
nano-banana-2-lite-g
Fast, cost-effective Google Gemini 3.1 Flash Image Lite model via Vertex AI. Optimized for ultra-fast turnarounds and low latency.
static
-
textimage
image
agent:canvasagent:workflowcapability:image-editcapability:image-generationcost:lowflashgooglelatestlitequality:balancedquality:draftselection:secondaryspeed:fast
View
Qwen Multiangle
qwen-edit-multiangle
Advanced multimodal model for complex reasoning, vision, and language tasks.
angles
motion-controlangles
textimage
image
qwenmultimodalreasoningvisionlogic
View
Sora2 Pro
sora-2-pro
OpenAI's state-of-the-art cinematic video generation model for high-end production.
video
Text to Video
textimage
video
cinematiccost:highflagship videomasterpieceopenaiselection:deprecatedselection:explicit-onlysora
View
Gemini 3.1 Pro
gemini-3-1-pro-vertex
Enhanced version of Gemini 3 Pro with refined multimodal logic and performance.
text
chat
textimageaudiovideo
text
v3.1expertmultimodalreasoningpreview
View
GPT-5
gpt-5
Strong general reasoning and long-form copy. Use for strategy, ad-structure deconstruction, and briefs that need judgement rather than throughput.
text
chat
textimageaudio
text
text-generationreasoningcopywritingmultimodal
View
ElevenLabs Music
eleven-music-v1
Text-to-Music generation for instrumental and melodic tracks.
music
music
text
audio
agent:canvasagent:workflowbeatsbgmcapability:music-generationcompositiongenerate musicquality:balancedselection:secondarysoundtrackuse-case:bgm
View
ElevenLabs SFX
eleven-sfx-v2
AI Foley and Sound Effects (SFX) from text descriptions.
sfx
sfx
text
audio
agent:canvasagent:workflowaudio assetscapability:foleycapability:sfx-generationcinematic soundfoleyquality:balancedselection:primarysfxsound effects
View
ElevenLabs V3
elevenlabs-v3-alpha
Highly expressive and emotional speech synthesis for character dialogue and audiobooks.
audio
voice
text
audio
agent:canvasagent:workflowaudiobookscapability:expressive-ttscapability:verbatim-ttscost:lowemotional ttsexpressivenarrationquality:balancedquality:draftrequirement:voice-selectionselection:secondaryvoice acting
View
Flash Reframe
luma-photon-flash-reframe
Re-compose images by shifting camera focus or framing via text prompts.
reframe
reframe
image
image
reframecompositionphotographyzoomcrop
View
Flux 2 Klein 9B
workers-ai-flux-2-klein-9b
Flux 2 Klein 4B image generation and editing
static
-
textimage
image
agent:canvasagent:workflowcapability:image-editcapability:image-generationcloudflarecost:lowfluxquality:draftselection:draft-primaryspeed:fastuse-case:rapid-prototypeworkers-ai
View
GPT Image 2
gpt-image-2
Generates and edits high-quality images from text prompts with precise instruction following, sharp text rendering, and configurable aspect ratios.
static
-
textimage
image
agent:canvasagent:workflowcapability:image-editcapability:image-generationcapability:text-renderingcreative-toolsgenerative-aihigh-fidelityimage-editingimage-generationquality:balancedquality:premiumselection:primarytext-to-imageuse-case:cinematic-stilluse-case:infographicuse-case:storyboard
View
Image Upscale (P)
p-image-upscale
AI-powered resolution enhancement for images or video to restore clarity.
upscale
upscale
image
image
upscaleenlargesharpenhd fixresolution
View
Topaz Video Upscale
topaz-video-upscale
AI-powered resolution enhancement for images or video to restore clarity.
video
Text to Video
video
video
upscaleenlargesharpenhd fixresolution
View
Veo 3.1 Fast (G)
veo3-1-fast-google
Google cinematic video generation for storytelling and professional creative work.
video
ReferencesTransitions
textimage
video
agent:canvasagent:workflowcapability:first-last-framecapability:image-to-videocapability:native-audiocapability:text-to-videocinematicgooglequality:balancedselection:secondaryspeed:faststorytellingveovideo gen
View
Gemini Omni Flash (G)
gemini-omni-flash-google
Google Gemini Omni Flash - fast conversational 720p video generation and editing.
video
ReferencesTransitions
textimagevideo
video
agent:canvasagent:workflowcapability:multimodal-videocapability:video-editconversationalgeminigoogleomniquality:balancedselection:primaryspeed:fastvideo gen
View
BirefNet2
birefnet2
Precise background removal with advanced edge matting for complex subjects like hair.
remove-bg
remove-bg
image
image
remove backgroundno-bgtransparent pngmattingsubject extraction
View
Bria Expand
bria-expand
Outpaint image borders to fit new aspect ratios while maintaining style.
expand
outpaintexpand
textimage
image
outpaintexpand canvasuncropaspect ratiofill
View
Gemini 3.1 Flash TTS
gemini-3-1-flash-tts-vertex
Controllable Gemini speech generation with automatic language detection and optional two-speaker dialogue.
audio
voice
text
audio
agent:canvasagent:workflowcapability:directed-ttscapability:multi-speaker-ttsgeminigooglemultispeakerpreviewquality:balancedquality:premiumrequirement:voice-selectionselection:primaryspeechttsvertexvoiceover
View
Luma Photon Reframe
luma-photon-reframe
Reframes to a new aspect ratio by generating the newly exposed edges. Use when widening or heightening beyond the original frame.
reframe
reframe
image
image
reframeoutpaintaspect-ratiogenerative-fill
View
Topaz Upscale
topaz-upscale-fal
AI-powered resolution enhancement for images or video to restore clarity.
upscale
upscale
image
image
upscaleenlargesharpenhd fixresolution
View
Veo 3.1 Lite (G)
veo3-1-lite-google
Value-tier Veo for short cinematic drafts with first/last-frame control and native audio. Use when P Video or Seedance Mini cannot preserve the requested treatment.
video
ReferencesTransitions
textimage
video
agent:canvasagent:workflowcapability:first-last-framecapability:image-to-videocapability:native-audiocapability:text-to-videocinematiccost:lowgooglequality:balancedquality:draftselection:secondaryspeed:faststorytellinguse-case:cinematic-draftveovideo gen
View
Bria Genfill
bria-genfill
Generative fill for missing image areas based on surrounding context.
inpaint
inpaint
textimage
image
inpaintgenfillcontent-awarerestorefix image
View
Image Reframer
reframe
Re-crops an image to a target aspect ratio without inventing new pixels. Use when the subject already fits and only the framing changes.
reframe
reframe
image
image
reframecropaspect-ratiodeterministic
View
Lyria 3 Clip
lyria-3-clip-vertex
Google Lyria 3 model for fixed 30-second music clips, loops, previews, instrumentals, and songs.
music
music
textimage
audio
30-secondagent:canvasagent:workflowaudiocapability:music-generationcost:lowgooglelyriamusicpreviewquality:draftselection:draft-primaryspeed:fastuse-case:music-previewvertex
View
Nano Banana 2 (G)
nano-banana-2-g
Fast, versatile image generation and prompt-based editing. The safe default for most stills; renders short in-image copy legibly.
static
-
textimage
image
agent:canvasagent:workflowcapability:image-editcapability:image-generationcapability:text-renderingfastgeneral-purposeimage-editingimage-generationquality:balancedselection:primaryspeed:fasttext-rendering
View
Veo 3.1 (G)
veo3-1-google
Google cinematic video generation for storytelling and professional creative work.
video
ReferencesTransitions
textimage
video
agent:canvasagent:workflowcapability:image-to-videocapability:native-audiocapability:reference-to-videocapability:text-to-videocinematiccost:highgooglequality:premiumselection:secondarystorytellinguse-case:cinematic-videoveovideo gen
View
Lyria 3 Pro
lyria-3-pro-vertex
Google Lyria 3 model for full-length songs with coherent verses, choruses, bridges, vocals, lyrics, and instrumental arrangements.
music
music
textimage
audio
agent:canvasagent:workflowaudiocapability:music-generationfull-songgooglelyrialyricsmusicpreviewquality:premiumselection:primaryuse-case:full-songvertex
View
Bria Product Cutout
bria-product-cutout
Automated background removal optimized specifically for e-commerce products.
remove-bg
remove-bg
image
image
product photographycutoutwhite backgrounde-commerceamazon
View
Nano Banana Pro (G)
nano-banana-pro-g
Higher-fidelity sibling of Nano Banana 2. Use for finished brand assets and any ad whose headline or CTA must render crisply in the image.
static
-
textimage
image
agent:canvasagent:workflowbrand-assetscapability:image-editcapability:image-generationcapability:text-renderingcost:highhigh-fidelityimage-editingimage-generationquality:premiumselection:primarytext-renderinguse-case:brand-asset
View
P Video
p-video
Lowest-cost video preview model with a native draft switch. Prefer for simple text/image-to-video drafts and high-volume motion tests; it is not the final-quality default.
video
-
textimage
video
agent:canvasagent:workflowcapability:first-last-framecapability:image-to-videocapability:text-to-videocost-effectivecost:lowhigh-volumeimage-to-videoquality:draftselection:draft-primaryspeed:fastuse-case:high-volumeuse-case:rapid-prototype
View
Qwen - Inpaint
qwen-inpaint
Advanced multimodal model for complex reasoning, vision, and language tasks.
static
inpaint
textimage
image
qwenmultimodalreasoningvisionlogic
View
Seedance 2 Fast
seedance-2-0-fast
High-speed video generation model supporting multimodal inputs, character consistency, and synchronized audio-driven animation.
video
-
textimageaudiovideo
video
agent:canvasagent:workflowaudio-drivencapability:audio-conditioned-videocapability:image-to-videocapability:native-audiocapability:reference-to-videocapability:text-to-videocharacter-consistencyfast-inferencemotion-transfermultimodalquality:balancedselection:primaryspeed:fastvideo-generation
View
Seedance 2
seedance-2-0
Multimodal video generation model with native audio synthesis, character consistency via reference inputs, and intelligent duration control.
video
-
textimageaudiovideo
video
agent:canvasagent:workflowaudio-drivencapability:audio-conditioned-videocapability:image-to-videocapability:native-audiocapability:reference-to-videocapability:text-to-videocharacter-consistencycost:highimage-to-videomultimodalquality:premiumselection:primarytext-to-videouse-case:character-consistencyuse-case:cinematic-videovideo-generation
View
Seedream 4.5
seedream-4-5
Immersive image and video generation based on high-level conceptual prompts.
static
-
textimage
image
seedreamconceptualartimmersivevisuals
View
Seedream 4
seedream-4-byteplus
Seedream 4.0 called directly on BytePlus ModelArk. Text-to-image and image editing with multi-reference support.
static
-
textimage
image
seedreambytedancebyteplusconceptualartimmersivevisuals
View
Seedance 1.5 Pro
seedance-1-5-pro
Very low-cost text/image-to-video fallback for silent 480p motion drafts and rapid iteration.
video
Text to Video
textimage
video
agent:canvasagent:workflowbytedancecapability:image-to-videocapability:text-to-videocost:lowcreativefast videoquality:draftresolution:480pseedanceselection:secondarysocialspeed:fastuse-case:rapid-prototype
View
Seedream 5 Lite
seedream-5-lite
Immersive image and video generation based on high-level conceptual prompts.
static
-
textimage
image
seedreamconceptualartimmersivevisuals
View
Seedream 5 Pro
seedream-5-pro
Seedream 5.0 Pro (Dola) called directly on BytePlus ModelArk. Highest-precision Seedream tier; single-image output.
static
-
textimage
image
agent:canvasagent:workflowartbytedancebytepluscapability:image-editcapability:image-generationcapability:multi-referenceconceptualimmersivemodel-pair:seedance-2quality:balancedquality:premiumseedreamselection:primaryuse-case:character-consistencyuse-case:layered-edituse-case:storyboardvisuals
View
Seedance 1.0 Pro
seedance-1-0-pro-byteplus
Seedance 1.0 Pro called directly on BytePlus ModelArk. Text/image-to-video (silent; the 1.0 family has no audio generation).
video
Text to Video
textimage
video
agent:canvasagent:workflowbytedancebytepluscapability:image-to-videocapability:silent-videocapability:text-to-videocost:lowcreativequality:draftresolution:480pseedanceselection:secondarysocialspeed:fastuse-case:rapid-prototypevideo gen
View
LTX 2.3 Fast
ltx-2-3-fast
Fast image-to-video for previewing motion before committing to a premium render.
video
-
textimage
video
image-to-videofastdraftingcost-effective
View
Bytedance Video Upscale
bytedance-video-upscaler
Increases video resolution as a finishing step. Run once, last, after the edit is locked.
video
remove-bg
video
video
video-upscaleresolutionenhancementfinishing
View
Generate Soundtrack
video-to-music
Generates a soundtrack that follows an existing video's pacing and mood. Use when music must fit a cut rather than the cut fitting music.
audio
-
videotext
audio
musicsoundtrackvideo-to-audioscoring
View
Happyhorse 1
happyhorse-1
Generates high-quality videos from text prompts or animates static images. Supports durations up to 15 seconds and resolutions up to 1080p.
video
Image to Video
textimage
video
video-generationimage-to-videotext-to-videoanimationcreative-toolshigh-definition
View
Happyhorse 1.1
happyhorse-1-1
Generates high-quality videos from text prompts or animates static images. Supports durations up to 15 seconds and resolutions up to 1080p.
video
Image to Video
textimage
video
video-generationimage-to-videotext-to-videoanimationcreative-toolshigh-definition
View
Kling Avatar v2
kling-avatar-v2
Professional cinematic video generation with advanced motion and temporal consistency.
video
Text to Video
textimage
video
klingvideo gencinematichigh-fidelitymotion
View
Omnihuman 1.5
omnihuman-1-5
Animates a portrait into natural full-body and facial motion. Strongest option for lifelike avatar presenters.
avatar-ugc
Avatar
textimageaudio
video
avatarhuman-animationportraitlipsync
View
P Video Animate
p-video-animate
video
-
textimageaudiovideo
video
-
View
P Video Avatar
p-video-avatar
Video
-
textimageaudio
video
-
View
P Video Replace
p-video-replace
video
-
textimagevideo
video
-
View
Pruna TryOn
p-image-try-on
static
-
textimage
image
-
View
Recraft Vectorize
recraft-vectorize
Converts an existing raster image into scalable vector artwork. Use to vectorize a supplied logo or a generated mark.
vectorize
vectorize
image
image
vectorsvgimage-to-vectorlogoscalable
View
Seedance 2 Mini
seedance-2-0-mini
Low-cost multimodal Seedance preview model. Prefer 480p for drafts that need references, people, choreography, source audio, or native generated audio.
video
-
textimageaudiovideo
video
agent:canvasagent:workflowbytedancecapability:audio-conditioned-videocapability:image-to-videocapability:native-audiocapability:reference-to-videocapability:text-to-videocost:lowcreativefast videoquality:draftresolution:480presolution:720pseedanceselection:draft-primarysocialspeed:fastuse-case:storyboard-animatic
View
Subtitles
veed-subtitles
Burns subtitles into a video, transcribing automatically or from a supplied SRT. Most social video should ship captioned.
video
Subtitles
textvideo
video
subtitlescaptionsaccessibilitysocial
View
Z image
z-image-base
Lightweight text-to-image for quick concepts and volume drafting.
static
-
text
image
text-to-imagefastdrafting
View
LTX 2.3 Pro
ltx-2-3-pro
Higher-fidelity image-to-video with stronger motion coherence. Use for delivery once the keyframes are approved.
video
-
textimageaudiovideo
video
image-to-videohigh-fidelitymotion
View
Recraft V4
recraft-v4
Text-to-image with strong stylistic control. Use for editorial, design-led, and illustrative looks rather than photoreal product work.
static
-
text
image
agent:canvasagent:workflowcapability:image-generationdesigneditorialquality:premiumselection:primarystyle-controltext-to-imageuse-case:design-led-graphicuse-case:editorial-graphic
View
Recraft V4 SVG
recraft-v4-svg
Generates true vector (SVG) output. The only text-to-vector model available — use whenever the deliverable must scale losslessly, such as a logo or icon.
static
-
text
image
agent:canvasagent:workflowcapability:vector-generationlogoquality:premiumscalableselection:primarysvgtext-to-imageuse-case:iconuse-case:logovector
View
P Image Edit
p-image-edit
High-speed multi-image editing model for production tasks including relighting, style transfer, and subject consistency. Optimized for sub-second performance.
static
-
textimage
image
agent:canvasagent:workflowcapability:image-editcost:lowfast-inferenceimage-editmulti-imageproduction-readyquality:draftselection:draft-primaryspeed:faststyle-transfersubject-consistency
View
Grok Image
grok-imagine-image
Quick image generation for drafting and broad exploration, where iteration speed matters more than final polish.
static
-
textimage
image
agent:canvasagent:workflowcapability:image-editcapability:image-generationcost:lowdraftingfastimage-editingimage-generationquality:draftselection:secondaryspeed:fastuse-case:rapid-prototype
View
Veed Fabric 1 fast
veed-fabric1-fast
Faster, lower-cost variant of Veed Fabric 1 for drafting a talking-head take before the final render.
avatar-ugc
Avatar
audioimage
video
lipsyncavatartalking-headugcfast
View
Veed Fabric 1
veed-fabric1
Turns a still portrait plus an audio track into a talking presenter. Takes no prompt — wire image_url and audio_url. The default for UGC-style pieces.
avatar-ugc
Avatar
audioimage
video
lipsyncavatartalking-headugc
View
Sync Lipsync
sync-linsync-2
Applies lipsync to EXISTING footage rather than a still. Use to re-voice or translate a clip you already have.
lipsync
Avatar
imagevideo
video
lipsyncvideo-dubbingtalking-headre-voicing
View