One model across modalities
Start from a prompt, continue from an image, or restage a source clip while carrying over the central subject, motion cues, and tone. MiniMax H3 Max is trained on images, videos, and audio together, so generation and understanding share one world model instead of fragmented pipelines.















