Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Produce up to 2K open-weight video at 24fps with built-in stereo sound — text, stills, or references all work as input via the comfyui minimax h3 workflow.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Kling Motion Control
Turn reference images into amazing motion videos in minutes
Why Build Video Pipelines with the comfyui minimax h3 Workflow
With the comfyui minimax h3 workflow, MiniMax's all-in-one omni-modal model runs directly inside ComfyUI as open weights. It processes text, visuals, footage, and sound together, producing clips with built-in stereo audio — voice, effects, and score generated in a single pass. You can push output to roughly 2K at 24fps for around 15 seconds while keeping direct command of every parameter.
- Every Clip Ships with Synced Stereo AudioSpeech, effects, and score are rendered together with the footage and saved as one MP4, so everything stays in sync without extra steps in the comfyui minimax h3 workflow.
- Local Control Without API LimitsOperate the comfyui minimax h3 model on your own hardware and tune resolution, length, and every diffusion variable freely — no API caps or external bottlenecks.
- Mix Text, Images, Video, and AudioFeed a single generation with text, pictures, footage, and audio cues to hold a subject, aesthetic, movement, framing, or voice — all through the comfyui minimax h3 nodes.
Three Simple Steps to Start with the comfyui minimax h3 Workflow
Follow this quick guide to get open-weight video with audio running through the comfyui minimax h3 workflow in minutes.
Key Features of the comfyui minimax h3 Workflow
A complete local video production stack: three native ComfyUI templates, omni-modal open-weight generation, native stereo audio, reference-driven control, and optional Sage Attention acceleration all live inside the comfyui minimax h3 workflow.
Three Native Workflow Templates
The comfyui minimax h3 template library includes text-to-video, image-to-video, and reference-to-video examples, each covering one generation mode out of the box.
Omni-Modal Context
The comfyui minimax h3 model reads text, images, video, and audio together in a single context, letting you combine all reference types in one generation.
Reference-Driven Generation
Lock a character's identity, a style, a motion, a camera move, or a voice from reference materials — up to 9 images, 3 videos, and 3 audio clips via the comfyui minimax h3 R2V node.
Accurate Text and Brand Rendering
Spelled-out text and brand elements render cleanly with the comfyui minimax h3 model, and instruction following describes reference relationships in natural language.
Sage Attention Speedup
Roughly double generation speed with minimal quality loss by adding the Patch Sage Attention KJ node to the comfyui minimax h3 workflow.
Resolution and Duration Grid
The comfyui minimax h3 Resolution Selector computes width and height from aspect ratio and megapixels, snapped to the model's 32-multiple grid and 17-frame-per-block duration at 24fps.
Common Questions About the comfyui minimax h3 Workflow
Quick answers on running MiniMax H3 inside ComfyUI — setup, quality, modes, audio, and speed.
What exactly is the comfyui minimax h3 workflow?
This workflow is ComfyUI's native integration of MiniMax H3, MiniMax's general-purpose omni-modal generation model released as open weights. It produces video with native stereo audio from text, images, video, and audio references in a single forward pass.
What output quality can I expect?
The comfyui minimax h3 workflow outputs up to 2K resolution at 24fps for about 15 seconds. Its native canvas uses a 768px short edge, capped at 768x1344 pixels and rounded to a multiple of 32.
Which generation modes are included?
The comfyui minimax h3 template library ships with three examples: text-to-video (T2V), image-to-video (I2V) with optional first/last-frame control, and reference-to-video (R2V) that locks in character, style, motion, camera, or voice.
Does it generate audio?
Yes — the comfyui minimax h3 model produces native stereo audio including voice, sound effects, and music, modeled together with the video in one pass and synced in a single MP4 file.
How do I get started?
Update ComfyUI to version 0.30.0 or later, open Template Library > Video, choose a comfyui minimax h3 workflow, and follow the pop-up to download models from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Can I speed up generation?
Yes — install SageAttention and the KJNodes custom nodes, then add a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow to roughly double generation speed.
Kick Off Video Creation with the comfyui minimax h3 Workflow
Run MiniMax H3 locally in ComfyUI with native stereo audio, open weights, and full parameter control — T2V, I2V, and R2V templates are ready to go. Press Generate to start your first clip.
