Chinese Interactive World Model Update: Text and Images Generate 3D Worlds with One Click
Bot Telegraph August 18th – HiDream.ai recently released its native multimodal interactive world model, HiDream-O1-World. This model supports multimodal inputs such as text, images, and interaction, and integrates three core functions: roaming, editing, and interaction, enabling users to generate a spatiotemporally consistent, physically plausible interactive world with a single click.
In the third-party evaluation benchmark WBench, HiDream-O1-World topped the Navi sub-ranking with an average score of 80.9. It scored 73.3 in the physical dimension, ranking first, and achieved a consistency dimension score of 88.0. The model is equipped with HiDream’s self-developed native multimodal (UiT) architecture, achieving breakthroughs in key spatiotemporal and physical consistency: when the camera moves, the scene geometry remains stable, and objects do not deform or disappear; collisions, occlusions, and gravity responses between objects also highly conform to real-world causal logic.
Users can start from text or images to explore generated scenes from a first- or third-person perspective and control character actions, weather, and object states. HiDream.ai plans to apply this model in fields such as AI interactive film and gaming, embodied intelligence simulation, and 3D scene production. To address common issues like scene drift in interactive world models, the model maintains stability during long interactions by writing 3D spatial information into the Memory context and combining it with Test-Time Training (TTT) for continuous adjustment during inference.
Additionally, a paper co-authored by HiDream.ai and Fudan University, titled “DreamWorld: Geometry-Driven Video Diffusion for 3D-Consistent World Modeling,” has been accepted by the top international conference ECCV 2026. The related research focuses on spatial consistency.
Source: Bot Telegraph China
