Models / MiniMax H3

Generate avatar video with MiniMax H3

Likna's video engine — MiniMax H3, rendered on fal.ai's cloud, with native audio included at no extra charge. Pick one of two resolution rungs per generation: 768p or 2K, each priced per second. Open-weight and genuinely self-hostable: Likna also runs the same weights on its own hardware, though that leg isn't part of the paid video flow today. See sample output, pricing, and how it compares below.

Get MiniMax H3 — on Creator and Studio

Speed

Cloud, i2v

Cost

7–10 credits / second

Best for

Highest-fidelity video with audio

Sample output

How it compares

FAQ

Can I generate a video with MiniMax H3 on Likna today?

Yes — it's Likna's only video engine, rendering on fal.ai's cloud at whichever of its two resolution rungs (768p or 2K) you pick per generation. Available on Creator and Studio plans.

What makes H3 different from Wan 2.2?

It's open-weight — Likna has the weights and runs it on its own hardware too, unlike the now-retired Wan 2.2, which only ran through a third-party RunPod cloud endpoint. It also renders native audio, which Wan 2.2 didn't do at all.

Does it generate audio?

Yes, natively — dialogue, sound effects and music render in the same pass as the picture, included in the price with no separate opt-in or surcharge.

How much does a clip cost?

7 credits per second at 768p, 10 at 2K — a 5-second clip is 35 or 50 credits and a 10-second clip is 70 or 100, depending on the rung. Audio is included either way, not an add-on.

Which resolution rung should I pick?

768p is the cheaper default for everyday clips; 2K costs more per second but holds more detail. Both include native audio and every duration — it's a quality/price choice, not a feature gate.

How is it prompted?

With an explicit shot timeline — [0s-2s], [2s-3.5s] and so on — rather than one freeform sentence, and it supports scripted, non-English-capable dialogue through a documented <d>[Language] ...</d> markup. It's also guidance-free: no negative prompt, no CFG scale, so steering happens entirely through the prompt text.

Does it preserve a supplied source photo?

In image-to-video mode, yes — it pins the first frame to the supplied still almost exactly (a measured mean pixel difference of 4.28/255, consistent with ordinary VAE/h264 roundtrip loss rather than reinterpretation). A separate reference-to-video variant trades that pinning for flexibility: up to 9 reference images plus reference video and reference audio, with framing left free and reference audio able to drive voice timbre and delivery.

What durations are available?

5, 10 or 15 seconds, all at 24fps, at either resolution rung.

Is it included in every plan?

No — video starts at Creator. Starter and the free trial don't include video generation at all.

Why did Likna switch to it?

Audio, mainly. The prior premium engine, Kling v3 Pro (now retired), charged a separate opt-in for sound; H3 renders native stereo audio in the same pass at no extra provider cost, so the same price now includes it.