Xpeng Unveils TuringViT Efficient Visual Encoder to Fully Support IRON Humanoid Robot
Bot Telegraph July 21st report, XPeng recently officially released the TuringViT high-efficiency visual encoder, systematically reconstructing the architectural design, data paradigm, and training process of visual encoders for the VLM/VLA era. TuringViT will comprehensively support three major business scenarios: intelligent driving, intelligent cockpit, and the IRON humanoid robot. This encoder achieves a stronger accuracy-efficiency trade-off than leading open-source baseline models using only 10% of the training data.
XPeng stated that TuringViT achieves an inference throughput of 3.04 times that of Seed1.5-ViT at a resolution of 1536×1536, clearing core obstacles for the edge deployment of high-resolution perception, multi-camera input, and long-term video processing. In the technical system of the IRON humanoid robot, TuringViT serves as the core role of the visual retina, providing foundational perception capabilities for embodied intelligence, such as fine-grained object recognition, spatial relationship understanding, operable area detection, and dynamic environment tracking.
Source: Bot Telegraph China
