SAN FRANCISCO — xAI is pushing its creative tools further, giving Grok's video model a set of upgrades that make it faster to use and sharper to watch. The company announced that Grok Imagine Video 1.5 now supports text-to-video, image-to-video and reference-to-video generation, along with native 1080p output, in an update detailed on xAI's news page.
The headline change is that users no longer need a starting image. Describe the shot in words and Grok pairs its image generation with image-to-video to produce the clip, now rendered at full 1080p on grok.com/imagine, iOS and Android. It is a meaningful step for a model xAI already billed as its best, and it lands as Elon Musk's AI company keeps up an aggressive release cadence.
References That Hold a Character Together
The most creator-friendly addition is reference control. Pass Grok a character image alongside a voice reference, and both stay consistent — the same face and the same voice across every scene. Creators can supply up to seven references per generation, locking a face, a product or a location in place while swapping everything else around it. Keep a character and change the scene, keep the scene and change the character, or hold both steady and change only the action.
That consistency has been one of the hardest problems in AI video, where faces and voices tend to drift from shot to shot. By letting users anchor the elements that matter, xAI is targeting the exact pain point that has kept generative video from feeling production-ready. The move builds on the momentum xAI showed when it brought Grok 4.5 to GitHub Copilot for millions of developers, extending Grok's reach across creative and technical workflows alike.



