Jul. 31, 2026
MiniMax Unveils H3 Multimodal Model With Open-Weight Plans
TMTPOST — Shanghai-based artificial intelligence developer MiniMax formally launched its H3 general-purpose multimodal generation model on Friday, introducing unified cross-modal processing for text, images, video, and audio. The system natively processes multi-format contexts to output video and audio streams featuring native stereo sound at a maximum resolution of 2K for durations up to 15 seconds. Built on architectures including the H3-Omni Transformer and H3-VAE, the model delivers default 2K rendering designed for commercial content production. Company executives confirmed that the organization intends to publish the model weights within days, pending compliance checks across relevant regulatory frameworks. Commercial video generators have historically operated under closed-source development paradigms that constrain ecosystem integration and hardware adaptation. Releasing open weights shifts deployment access toward developers and enterprises seeking customized architectures on local or cloud hardware infrastructure.
More News

  • Subscribe To Our News