Apply Speed Curve renderer-ease-curve Apply cinematic speed curves (Ease, Bezier) to video clips with audio sync. | clip-editor | editing | video | video | slow motionslomospeed ramptime warpcinematicvideo edit | View |
Claude Sonnet 4.6 claude-sonnet-4-6 SOTA reasoning, coding, and long-context analysis (1M+ tokens). | text | chat | textimageaudio | text | chatcodingreasoningplanningagentcomplex | View |
Compress Video renderer-video-downsize Compress video to a specific target file size using FFmpeg. | clip-editor | editing | video | video | compressshrinkoptimizefile sizemp4ffmpeg | View |
Extract Frame renderer-extract-frame Extract the first frame, last frame, or a frame at an exact non-negative timestamp from a video without re-encoding. | clip-editor | editing | video | image | thumbnailpreviewframe extractionpostercapture | View |
Loop Video renderer-video-loop Repeat a video to a target duration or loop count. | clip-editor | editing | video | video | videolooprepeatffmpeg | View |
Merge Audio-Video renderer-audio-video-merge Utility for merging video and audio streams into a single high-quality file. | clip-editor | editing | videoaudio | video | mergeaudio-videoffmpegmuxingeditor | View |
Overlay / Watermark renderer-video-overlay Overlay an image or logo on top of a video. | clip-editor | editing | videoimage | video | videooverlaywatermarklogoimageffmpeg | View |
Reverse Video renderer-video-reverse Reverse video and audio when audio is present. | clip-editor | editing | video | video | videoreverserewindffmpeg | View |
Slideshow Video renderer-image-slideshow Create a video from one or more images with optional audio. | clip-editor | editing | imageaudio | video | imagevideoslideshowffmpeg | View |
Stitch Videos renderer-video-merge Append videos using the first clip canvas, with optional transitions and explicit fit behavior. | clip-editor | editing | video | video | audio-videoeditorffmpegmergemuxingtransition | View |
Trim Video renderer-video-trim Cuts a clip to a start/end range. Deterministic FFmpeg operation — no generation, no cost variance. | clip-editor | editing | video | video | video-edittrimdeterministicffmpeg | View |
Video to WebP renderer-video-to-webp Convert a video segment to animated WebP. | clip-editor | editing | video | image | videowebpanimatedffmpeg | View |
Gemini 3.5 Flash-Lite gemini-3-5-flash-lite-vertex Fast, cost-efficient GA Gemini model for high-throughput text, document, image, audio, and video understanding. | text | chat | textimageaudiovideo | text | geminigooglegaflash-litemultimodalhigh-throughput | View |
Runway Gen 4.5 runway-gen-4-5 Professional AI video generation and editing tools for cinematic production. | video | Text to Video | textimage | video | runwayvideo gencinematicvfxcreative | View |
Gemini 3.6 Flash gemini-3-6-flash-vertex GA Gemini Flash model for coding, multimodal reasoning, knowledge work, and multi-step workflows. | text | chat | textimageaudiovideo | text | geminigooglegaflashmultimodalreasoning | View |
Kling v3 kling-v3-video Professional cinematic video generation with advanced motion and temporal consistency. | video | Text to Video | textimage | video | agent:canvasagent:workflowcapability:image-to-videocapability:text-to-videocinematiccost:highhigh-fidelityklingmotionquality:premiumselection:primaryuse-case:cinematic-videouse-case:motion-controlvideo gen | View |
Kling V3 Omni kling-v3-omni-video Professional cinematic video generation with advanced motion and temporal consistency. | video | Text to Video | textimagevideoaudio | video | agent:canvasagent:workflowcapability:audio-conditioned-videocapability:image-to-videocapability:text-to-videocapability:video-editcinematiccost:highhigh-fidelityklingmotionquality:premiumselection:primaryuse-case:cinematic-videovideo gen | View |
Nano Banana 2 Lite (G) nano-banana-2-lite-g Fast, cost-effective Google Gemini 3.1 Flash Image Lite model via Vertex AI. Optimized for ultra-fast turnarounds and low latency. | static | - | textimage | image | agent:canvasagent:workflowcapability:image-editcapability:image-generationcost:lowflashgooglelatestlitequality:balancedquality:draftselection:secondaryspeed:fast | View |
Qwen Multiangle qwen-edit-multiangle Advanced multimodal model for complex reasoning, vision, and language tasks. | angles | motion-controlangles | textimage | image | qwenmultimodalreasoningvisionlogic | View |
Sora2 Pro sora-2-pro OpenAI's state-of-the-art cinematic video generation model for high-end production. | video | Text to Video | textimage | video | cinematiccost:highflagship videomasterpieceopenaiselection:deprecatedselection:explicit-onlysora | View |
Gemini 3.1 Pro gemini-3-1-pro-vertex Enhanced version of Gemini 3 Pro with refined multimodal logic and performance. | text | chat | textimageaudiovideo | text | v3.1expertmultimodalreasoningpreview | View |
GPT-5 gpt-5 Strong general reasoning and long-form copy. Use for strategy, ad-structure deconstruction, and briefs that need judgement rather than throughput. | text | chat | textimageaudio | text | text-generationreasoningcopywritingmultimodal | View |
ElevenLabs Music eleven-music-v1 Text-to-Music generation for instrumental and melodic tracks. | music | music | text | audio | agent:canvasagent:workflowbeatsbgmcapability:music-generationcompositiongenerate musicquality:balancedselection:secondarysoundtrackuse-case:bgm | View |
ElevenLabs SFX eleven-sfx-v2 AI Foley and Sound Effects (SFX) from text descriptions. | sfx | sfx | text | audio | agent:canvasagent:workflowaudio assetscapability:foleycapability:sfx-generationcinematic soundfoleyquality:balancedselection:primarysfxsound effects | View |
ElevenLabs V3 elevenlabs-v3-alpha Highly expressive and emotional speech synthesis for character dialogue and audiobooks. | audio | voice | text | audio | agent:canvasagent:workflowaudiobookscapability:expressive-ttscapability:verbatim-ttscost:lowemotional ttsexpressivenarrationquality:balancedquality:draftrequirement:voice-selectionselection:secondaryvoice acting | View |
Flash Reframe luma-photon-flash-reframe Re-compose images by shifting camera focus or framing via text prompts. | reframe | reframe | image | image | reframecompositionphotographyzoomcrop | View |
Flux 2 Klein 9B workers-ai-flux-2-klein-9b Flux 2 Klein 4B image generation and editing | static | - | textimage | image | agent:canvasagent:workflowcapability:image-editcapability:image-generationcloudflarecost:lowfluxquality:draftselection:draft-primaryspeed:fastuse-case:rapid-prototypeworkers-ai | View |
GPT Image 2 gpt-image-2 Generates and edits high-quality images from text prompts with precise instruction following, sharp text rendering, and configurable aspect ratios. | static | - | textimage | image | agent:canvasagent:workflowcapability:image-editcapability:image-generationcapability:text-renderingcreative-toolsgenerative-aihigh-fidelityimage-editingimage-generationquality:balancedquality:premiumselection:primarytext-to-imageuse-case:cinematic-stilluse-case:infographicuse-case:storyboard | View |
Image Upscale (P) p-image-upscale AI-powered resolution enhancement for images or video to restore clarity. | upscale | upscale | image | image | upscaleenlargesharpenhd fixresolution | View |
Topaz Video Upscale topaz-video-upscale AI-powered resolution enhancement for images or video to restore clarity. | video | Text to Video | video | video | upscaleenlargesharpenhd fixresolution | View |
Veo 3.1 Fast (G) veo3-1-fast-google Google cinematic video generation for storytelling and professional creative work. | video | ReferencesTransitions | textimage | video | agent:canvasagent:workflowcapability:first-last-framecapability:image-to-videocapability:native-audiocapability:text-to-videocinematicgooglequality:balancedselection:secondaryspeed:faststorytellingveovideo gen | View |
Gemini Omni Flash (G) gemini-omni-flash-google Google Gemini Omni Flash - fast conversational 720p video generation and editing. | video | ReferencesTransitions | textimagevideo | video | agent:canvasagent:workflowcapability:multimodal-videocapability:video-editconversationalgeminigoogleomniquality:balancedselection:primaryspeed:fastvideo gen | View |
BirefNet2 birefnet2 Precise background removal with advanced edge matting for complex subjects like hair. | remove-bg | remove-bg | image | image | remove backgroundno-bgtransparent pngmattingsubject extraction | View |
Bria Expand bria-expand Outpaint image borders to fit new aspect ratios while maintaining style. | expand | outpaintexpand | textimage | image | outpaintexpand canvasuncropaspect ratiofill | View |
Gemini 3.1 Flash TTS gemini-3-1-flash-tts-vertex Controllable Gemini speech generation with automatic language detection and optional two-speaker dialogue. | audio | voice | text | audio | agent:canvasagent:workflowcapability:directed-ttscapability:multi-speaker-ttsgeminigooglemultispeakerpreviewquality:balancedquality:premiumrequirement:voice-selectionselection:primaryspeechttsvertexvoiceover | View |
Luma Photon Reframe luma-photon-reframe Reframes to a new aspect ratio by generating the newly exposed edges. Use when widening or heightening beyond the original frame. | reframe | reframe | image | image | reframeoutpaintaspect-ratiogenerative-fill | View |
Topaz Upscale topaz-upscale-fal AI-powered resolution enhancement for images or video to restore clarity. | upscale | upscale | image | image | upscaleenlargesharpenhd fixresolution | View |
Veo 3.1 Lite (G) veo3-1-lite-google Value-tier Veo for short cinematic drafts with first/last-frame control and native audio. Use when P Video or Seedance Mini cannot preserve the requested treatment. | video | ReferencesTransitions | textimage | video | agent:canvasagent:workflowcapability:first-last-framecapability:image-to-videocapability:native-audiocapability:text-to-videocinematiccost:lowgooglequality:balancedquality:draftselection:secondaryspeed:faststorytellinguse-case:cinematic-draftveovideo gen | View |
Bria Genfill bria-genfill Generative fill for missing image areas based on surrounding context. | inpaint | inpaint | textimage | image | inpaintgenfillcontent-awarerestorefix image | View |
Image Reframer reframe Re-crops an image to a target aspect ratio without inventing new pixels. Use when the subject already fits and only the framing changes. | reframe | reframe | image | image | reframecropaspect-ratiodeterministic | View |
Lyria 3 Clip lyria-3-clip-vertex Google Lyria 3 model for fixed 30-second music clips, loops, previews, instrumentals, and songs. | music | music | textimage | audio | 30-secondagent:canvasagent:workflowaudiocapability:music-generationcost:lowgooglelyriamusicpreviewquality:draftselection:draft-primaryspeed:fastuse-case:music-previewvertex | View |
Nano Banana 2 (G) nano-banana-2-g Fast, versatile image generation and prompt-based editing. The safe default for most stills; renders short in-image copy legibly. | static | - | textimage | image | agent:canvasagent:workflowcapability:image-editcapability:image-generationcapability:text-renderingfastgeneral-purposeimage-editingimage-generationquality:balancedselection:primaryspeed:fasttext-rendering | View |
Veo 3.1 (G) veo3-1-google Google cinematic video generation for storytelling and professional creative work. | video | ReferencesTransitions | textimage | video | agent:canvasagent:workflowcapability:image-to-videocapability:native-audiocapability:reference-to-videocapability:text-to-videocinematiccost:highgooglequality:premiumselection:secondarystorytellinguse-case:cinematic-videoveovideo gen | View |
Lyria 3 Pro lyria-3-pro-vertex Google Lyria 3 model for full-length songs with coherent verses, choruses, bridges, vocals, lyrics, and instrumental arrangements. | music | music | textimage | audio | agent:canvasagent:workflowaudiocapability:music-generationfull-songgooglelyrialyricsmusicpreviewquality:premiumselection:primaryuse-case:full-songvertex | View |
Bria Product Cutout bria-product-cutout Automated background removal optimized specifically for e-commerce products. | remove-bg | remove-bg | image | image | product photographycutoutwhite backgrounde-commerceamazon | View |
Nano Banana Pro (G) nano-banana-pro-g Higher-fidelity sibling of Nano Banana 2. Use for finished brand assets and any ad whose headline or CTA must render crisply in the image. | static | - | textimage | image | agent:canvasagent:workflowbrand-assetscapability:image-editcapability:image-generationcapability:text-renderingcost:highhigh-fidelityimage-editingimage-generationquality:premiumselection:primarytext-renderinguse-case:brand-asset | View |
P Video p-video Lowest-cost video preview model with a native draft switch. Prefer for simple text/image-to-video drafts and high-volume motion tests; it is not the final-quality default. | video | - | textimage | video | agent:canvasagent:workflowcapability:first-last-framecapability:image-to-videocapability:text-to-videocost-effectivecost:lowhigh-volumeimage-to-videoquality:draftselection:draft-primaryspeed:fastuse-case:high-volumeuse-case:rapid-prototype | View |
Qwen - Inpaint qwen-inpaint Advanced multimodal model for complex reasoning, vision, and language tasks. | static | inpaint | textimage | image | qwenmultimodalreasoningvisionlogic | View |
Seedance 2 Fast seedance-2-0-fast High-speed video generation model supporting multimodal inputs, character consistency, and synchronized audio-driven animation. | video | - | textimageaudiovideo | video | agent:canvasagent:workflowaudio-drivencapability:audio-conditioned-videocapability:image-to-videocapability:native-audiocapability:reference-to-videocapability:text-to-videocharacter-consistencyfast-inferencemotion-transfermultimodalquality:balancedselection:primaryspeed:fastvideo-generation | View |
Seedance 2 seedance-2-0 Multimodal video generation model with native audio synthesis, character consistency via reference inputs, and intelligent duration control. | video | - | textimageaudiovideo | video | agent:canvasagent:workflowaudio-drivencapability:audio-conditioned-videocapability:image-to-videocapability:native-audiocapability:reference-to-videocapability:text-to-videocharacter-consistencycost:highimage-to-videomultimodalquality:premiumselection:primarytext-to-videouse-case:character-consistencyuse-case:cinematic-videovideo-generation | View |
Seedream 4.5 seedream-4-5 Immersive image and video generation based on high-level conceptual prompts. | static | - | textimage | image | seedreamconceptualartimmersivevisuals | View |
Seedream 4 seedream-4-byteplus Seedream 4.0 called directly on BytePlus ModelArk. Text-to-image and image editing with multi-reference support. | static | - | textimage | image | seedreambytedancebyteplusconceptualartimmersivevisuals | View |
Seedance 1.5 Pro seedance-1-5-pro Very low-cost text/image-to-video fallback for silent 480p motion drafts and rapid iteration. | video | Text to Video | textimage | video | agent:canvasagent:workflowbytedancecapability:image-to-videocapability:text-to-videocost:lowcreativefast videoquality:draftresolution:480pseedanceselection:secondarysocialspeed:fastuse-case:rapid-prototype | View |
Seedream 5 Lite seedream-5-lite Immersive image and video generation based on high-level conceptual prompts. | static | - | textimage | image | seedreamconceptualartimmersivevisuals | View |
Seedream 5 Pro seedream-5-pro Seedream 5.0 Pro (Dola) called directly on BytePlus ModelArk. Highest-precision Seedream tier; single-image output. | static | - | textimage | image | agent:canvasagent:workflowartbytedancebytepluscapability:image-editcapability:image-generationcapability:multi-referenceconceptualimmersivemodel-pair:seedance-2quality:balancedquality:premiumseedreamselection:primaryuse-case:character-consistencyuse-case:layered-edituse-case:storyboardvisuals | View |
Seedance 1.0 Pro seedance-1-0-pro-byteplus Seedance 1.0 Pro called directly on BytePlus ModelArk. Text/image-to-video (silent; the 1.0 family has no audio generation). | video | Text to Video | textimage | video | agent:canvasagent:workflowbytedancebytepluscapability:image-to-videocapability:silent-videocapability:text-to-videocost:lowcreativequality:draftresolution:480pseedanceselection:secondarysocialspeed:fastuse-case:rapid-prototypevideo gen | View |
LTX 2.3 Fast ltx-2-3-fast Fast image-to-video for previewing motion before committing to a premium render. | video | - | textimage | video | image-to-videofastdraftingcost-effective | View |
Bytedance Video Upscale bytedance-video-upscaler Increases video resolution as a finishing step. Run once, last, after the edit is locked. | video | remove-bg | video | video | video-upscaleresolutionenhancementfinishing | View |
Generate Soundtrack video-to-music Generates a soundtrack that follows an existing video's pacing and mood. Use when music must fit a cut rather than the cut fitting music. | audio | - | videotext | audio | musicsoundtrackvideo-to-audioscoring | View |
Happyhorse 1 happyhorse-1 Generates high-quality videos from text prompts or animates static images. Supports durations up to 15 seconds and resolutions up to 1080p. | video | Image to Video | textimage | video | video-generationimage-to-videotext-to-videoanimationcreative-toolshigh-definition | View |
Happyhorse 1.1 happyhorse-1-1 Generates high-quality videos from text prompts or animates static images. Supports durations up to 15 seconds and resolutions up to 1080p. | video | Image to Video | textimage | video | video-generationimage-to-videotext-to-videoanimationcreative-toolshigh-definition | View |
Kling Avatar v2 kling-avatar-v2 Professional cinematic video generation with advanced motion and temporal consistency. | video | Text to Video | textimage | video | klingvideo gencinematichigh-fidelitymotion | View |
Omnihuman 1.5 omnihuman-1-5 Animates a portrait into natural full-body and facial motion. Strongest option for lifelike avatar presenters. | avatar-ugc | Avatar | textimageaudio | video | avatarhuman-animationportraitlipsync | View |
P Video Animate p-video-animate | video | - | textimageaudiovideo | video | - | View |
P Video Avatar p-video-avatar | Video | - | textimageaudio | video | - | View |
P Video Replace p-video-replace | video | - | textimagevideo | video | - | View |
Pruna TryOn p-image-try-on | static | - | textimage | image | - | View |
Recraft Vectorize recraft-vectorize Converts an existing raster image into scalable vector artwork. Use to vectorize a supplied logo or a generated mark. | vectorize | vectorize | image | image | vectorsvgimage-to-vectorlogoscalable | View |
Seedance 2 Mini seedance-2-0-mini Low-cost multimodal Seedance preview model. Prefer 480p for drafts that need references, people, choreography, source audio, or native generated audio. | video | - | textimageaudiovideo | video | agent:canvasagent:workflowbytedancecapability:audio-conditioned-videocapability:image-to-videocapability:native-audiocapability:reference-to-videocapability:text-to-videocost:lowcreativefast videoquality:draftresolution:480presolution:720pseedanceselection:draft-primarysocialspeed:fastuse-case:storyboard-animatic | View |
Subtitles veed-subtitles Burns subtitles into a video, transcribing automatically or from a supplied SRT. Most social video should ship captioned. | video | Subtitles | textvideo | video | subtitlescaptionsaccessibilitysocial | View |
Z image z-image-base Lightweight text-to-image for quick concepts and volume drafting. | static | - | text | image | text-to-imagefastdrafting | View |
LTX 2.3 Pro ltx-2-3-pro Higher-fidelity image-to-video with stronger motion coherence. Use for delivery once the keyframes are approved. | video | - | textimageaudiovideo | video | image-to-videohigh-fidelitymotion | View |
Recraft V4 recraft-v4 Text-to-image with strong stylistic control. Use for editorial, design-led, and illustrative looks rather than photoreal product work. | static | - | text | image | agent:canvasagent:workflowcapability:image-generationdesigneditorialquality:premiumselection:primarystyle-controltext-to-imageuse-case:design-led-graphicuse-case:editorial-graphic | View |
Recraft V4 SVG recraft-v4-svg Generates true vector (SVG) output. The only text-to-vector model available — use whenever the deliverable must scale losslessly, such as a logo or icon. | static | - | text | image | agent:canvasagent:workflowcapability:vector-generationlogoquality:premiumscalableselection:primarysvgtext-to-imageuse-case:iconuse-case:logovector | View |
P Image Edit p-image-edit High-speed multi-image editing model for production tasks including relighting, style transfer, and subject consistency. Optimized for sub-second performance. | static | - | textimage | image | agent:canvasagent:workflowcapability:image-editcost:lowfast-inferenceimage-editmulti-imageproduction-readyquality:draftselection:draft-primaryspeed:faststyle-transfersubject-consistency | View |
Grok Image grok-imagine-image Quick image generation for drafting and broad exploration, where iteration speed matters more than final polish. | static | - | textimage | image | agent:canvasagent:workflowcapability:image-editcapability:image-generationcost:lowdraftingfastimage-editingimage-generationquality:draftselection:secondaryspeed:fastuse-case:rapid-prototype | View |
Veed Fabric 1 fast veed-fabric1-fast Faster, lower-cost variant of Veed Fabric 1 for drafting a talking-head take before the final render. | avatar-ugc | Avatar | audioimage | video | lipsyncavatartalking-headugcfast | View |
Veed Fabric 1 veed-fabric1 Turns a still portrait plus an audio track into a talking presenter. Takes no prompt — wire image_url and audio_url. The default for UGC-style pieces. | avatar-ugc | Avatar | audioimage | video | lipsyncavatartalking-headugc | View |
Sync Lipsync sync-linsync-2 Applies lipsync to EXISTING footage rather than a still. Use to re-voice or translate a clip you already have. | lipsync | Avatar | imagevideo | video | lipsyncvideo-dubbingtalking-headre-voicing | View |