Key Takeaways
- LatentVerse raised a new round (Seed) from GL Ventures, Crystal Stream, Agibot, Robot Era, Innoangel Fund.
- Sector: Artificial Intelligence (AI), Technology, Software & Gaming.
- Geography: China.
Analysis
A new venture, LatentVerse, is charting a distinct course in the rapidly evolving embodied intelligence arena, seeking to unify disparate AI modeling approaches. The startup, founded by researcher Hu Yucheng, aims to move beyond the prevailing Vision-Language-Action (VLA) paradigms and traditional world models. This strategic divergence comes as the company secures a significant nine-figure RMB seed funding round from a consortium of prominent investors including GL Ventures, Crystal Stream, Agibot, Robot Era, and Innoangel Fund.
Unlike many contemporaries focusing on pure world models, which predict future states through video generation, or VLA systems that directly link perception to action, LatentVerse is developing a unified architecture. This approach integrates the semantic understanding capabilities of large vision-language models with the predictive power of world models. The goal is to create foundation models that inherently possess understanding, prediction, and action generation within a single framework, addressing what Hu Yucheng identifies as a critical bottleneck in achieving physical artificial general intelligence.
The limitations of existing methods are substantial. VLA systems often require costly, action-labeled robotic manipulation data, with collection costs potentially exceeding RMB 1,000 per hour. Conversely, pure world models, while leveraging cheaper internet video data, may struggle with nuanced semantic interpretation, hindering their ability to grasp complex instructions like "hand the cup to the guest." LatentVerse's proposed third path seeks to harness the strengths of both, mitigating their respective drawbacks.
LatentVerse's initial offering, the Unified Tactile Action Model (UTAM), is designed to output simultaneous visual, language, and tactile signals. This model employs a four-module execution chain: a vision-language expert for intent interpretation, a world model expert for task planning, an action expert for generating coarse action representations, and a tactile expert for high-frequency, real-time end-effector control. This layered approach, combining open-loop execution with closed-loop tactile correction, mirrors human motor control, allowing for precise adjustments based on physical feedback.
The company's focus on generalization, long-horizon execution, and dexterity is crucial for commercial viability in robotics. Hu Yucheng emphasizes that these are not isolated capabilities but essential, simultaneous requirements for robots operating in real-world scenarios such as domestic assistance or complex industrial tasks. By collecting data across multiple dimensions and training for extended task durations, LatentVerse aims to equip robots with both cognitive and contact intelligence, enabling them to navigate and interact effectively in unstructured environments.
The founding team comprises talent from prestigious institutions and industry leaders, including Tsinghua University, Nanyang Technological University, Peking University, and former members from ByteDance and Xiaomi. This blend of academic rigor and industry experience positions LatentVerse to tackle the intricate challenges of embodied AI. The startup's innovative architectural design and strategic funding position it as a noteworthy player to watch in the competitive AI hardware and software sector.