comfyui minimax h3
Render audio-synced clips straight from the comfyui minimax h3 node graph
AI Video Prompt Generator
10s

Feedback

AI Ad Video Example

Loading...

comfyui minimax h3

Generate 2K clips with synced stereo sound from prompts, photos, or footage — the comfyui minimax h3 workflow gives you open weights and node-level control.

All Tools

Discover our comprehensive AI-powered animation toolkit

Open-Weight, Audio-Native Video: the comfyui minimax h3 Workflow

Shipped as open weights, the comfyui minimax h3 workflow embeds MiniMax's omni-modal engine directly into ComfyUI. It reads text, images, video, and audio inside one shared context, then renders footage together with stereo sound — speech, effects, and music produced in a single forward pass. Clips reach 2K at 24fps for roughly 15 seconds, and every parameter stays under node-level control.

  • Stereo Sound, Generated Natively
    Speech, sound effects, and music arrive in the same MP4 as the picture — the comfyui minimax h3 workflow syncs all of it in one pass instead of stitching audio on afterward.
  • Open Weights, Local Runs
    Load the comfyui minimax h3 model on your own GPU and tune resolution, clip length, and every diffusion setting yourself — nothing is capped by an external API.
  • Mix Text, Image, Video, Audio
    Feed several reference types into one generation to pin down a face, a look, a motion, a camera path, or a voice through the comfyui minimax h3 nodes.

Running the comfyui minimax h3 Workflow in Three Steps

From a fresh install to an audio-synced MP4 — here is the shortest path through the comfyui minimax h3 workflow.

What the comfyui minimax h3 Workflow Can Do

Three ready-made templates, omni-modal context, reference locking, clean text rendering, an optional speed patch, and a smart resolution grid — that is the full stack the comfyui minimax h3 workflow puts on your machine.

Three Ready-Made Templates

Text-to-video, image-to-video, and reference-to-video examples come bundled in the comfyui minimax h3 template library, one for each generation mode.

One Shared Context for Every Modality

Text, stills, footage, and audio are interpreted together rather than separately, so the comfyui minimax h3 model can blend reference types inside a single generation.

References That Lock Details

Pin a character's face, an art style, a motion, a camera move, or a voice using up to 9 images, 3 videos, and 3 audio clips through the comfyui minimax h3 R2V node.

Text and Logos That Read Cleanly

On-screen words and brand marks come out legible with the comfyui minimax h3 model, and plain-language prompts can describe how each reference relates to the rest.

Optional Sage Attention Patch

Drop the Patch Sage Attention KJ node into the comfyui minimax h3 graph to roughly double throughput while keeping quality loss minimal.

Resolution and Duration Grids

The comfyui minimax h3 Resolution Selector derives width and height from aspect ratio and megapixels, snapped to the model's 32-pixel grid and 17-frame blocks at 24fps.

FAQ

Answers About the comfyui minimax h3 Workflow

Straight answers to the questions people ask most about running MiniMax H3 inside ComfyUI.

1

What exactly is the comfyui minimax h3 workflow?

It is ComfyUI's built-in integration of MiniMax H3, an omni-modal generation model MiniMax released as open weights. A single forward pass turns text, images, video, and audio references into video that already carries stereo sound.

2

How high does the output resolution go?

The comfyui minimax h3 workflow renders up to 2K at 24fps for roughly 15 seconds. Its native canvas uses a 768px short edge, tops out at 768x1344 pixels, and rounds dimensions to a multiple of 32.

3

Which generation modes are included?

Three examples ship in the comfyui minimax h3 template library: text-to-video (T2V), image-to-video (I2V) with optional first- and last-frame control, and reference-to-video (R2V) for locking character, style, motion, camera, or voice.

4

Does it produce audio as well?

It does. Speech, sound effects, and music are modeled alongside the picture by the comfyui minimax h3 model and delivered in sync inside a single MP4 file.

5

What do I need before running it?

ComfyUI 0.30.0 or later is the baseline. Open Template Library > Video, pick a comfyui minimax h3 workflow, and accept the pop-up that downloads models from the Hugging Face Comfy-Org/MiniMax-H3 repository.

6

Is there a way to make it render faster?

Yes. Install SageAttention plus the KJNodes custom nodes, then insert a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow to roughly double generation speed.

Put the comfyui minimax h3 Workflow to Work Today

Open weights, synced stereo audio, and every parameter in your hands — run MiniMax H3 locally in ComfyUI with text-to-video, image-to-video, and reference-to-video graphs ready to launch.