Gasgoo Munich-On September 15, Gasgoo hosted The Symposium on Embodied Perception Fusion & Multimodal LargeModel nnovation 2026in Shanghai. During a panel discussion, Mao Jiming, partner and vice president at GIGAAI, joined Liao Yongxing, AI algorithm director at Zhejiang Fulai New Materials Co., Ltd., Lu Yao, chief scientist at Sunrising AI Lab, Zhuo Weifeng, multi-sensor fusion lead at AiMOGA, and Ji Haifeng, solutions director at RealMan Group. Centered on the theme "From Seeing to Doing: Where Embodied AI Breaks Through to General Intelligence," the group tackled the toughest hurdles standing between the lab and the real world. How to choose a model route? How does multi-modal perception translate into physical interaction? How to combine real-world, simulation, and human data? And where exactly do the bottlenecks lie—in models, data, perception, or the hardware itself?While opinions varied, the experts reached a shared verdict on the industry's trajectory: the breakthrough for embodied AI won't come from a single technology winning, but from the maturation of system capabilities. Rapid advancements in VLA, world models, and physics-native models are pushing robots from merely seeing the world toward understanding it, making decisions, and taking action. Yet, once deployed in real-world scenarios, constraints emerge simultaneously across perception fusion, data acquisition, model generalization, and hardware application.Image source: GasgooModel Routes Aren't a Multiple-Choice QuestionThe discussion opened with a direct hit on the core controversy: among VLA, world models, and physics-native models, which serves as the brain of embodied AI?The industry often simplifies this into a battle of technical routes, but Liao Yongxing offered a different take: it’s less about which model dominates and more about a layered architecture. You can have a primary "brain" handling high-level semantic understanding, spatial awareness, and object localization—interpreting task prompts as a grasp of the physical environment before outputting motion trajectories. The critical next step is executing those trajectories. This involves a "cerebellum" role—breaking down motion paths or tasks into actual interactions within the physical world.This distinction highlights the fundamental gap between embodied AI and pure language models. Large language models can reason within a relatively closed symbolic space, but robots must contend with contact, friction, inertia, deformation, and real-time control. Even if a single end-to-end model dazzles in a demo, it struggles to balance high-level semantic understanding with low-level, high-frequency control.Layered fusion isn't a compromise; it is an engineering necessity dictated by the constraints of the physical world.Lu Yao approached the debate from a different angle, dismissing the fixation on paradigms. Much of the current noise centers on whether to use VLA or world models, but the paradigm itself shouldn't be the point of contention. Instead, the focus should be on the actual development needs of general intelligence. As Lu Yao put it, the ultimate goal of model development is to better perceive and understand the world—and then get the task done.Whether the resulting model resembles VLA or a world model is just that—a result. Developers shouldn't limit capabilities by locking into a paradigm too early. A better path involves leveraging the power of large models while tapping into the accumulated wisdom of traditional industrial sectors, integrating them to build a truly mature generalization solution.This perspective is particularly relevant right now. The industry easily gets trapped in label-driven debates, yet application scenarios for embodied AI are highly fragmented. Industrial assembly, home services, retail replenishment, and automotive final assembly all impose different requirements on models. Picking a side in the paradigm war too early only stifles the space for technical exploration.Zhuo Weifeng echoed this sentiment. The "brain" of embodied AI, he suggested, might not be a single unified large model, but rather a system composed of multiple modules: task planning, trajectory generation, and motion control.Ultimately, a model's value is measured in specific scenarios by cycle time, precision, stability, and cost.Ji Haifeng also explicitly backed a fusion approach. "For robots to function in homes and diverse environments, a single model likely won't solve the problem," he said. "VLA offers action sequences but lacks physical constraints and the ability to predict the next moment, whereas world models and physics models hold inherent advantages." In his view, several "brains" will ultimately fuse together to form a robot's super-brain, steering the technology toward something that increasingly resembles the human brain."Slow thinking" handles task understanding and long-term planning, while "fast reaction" manages contact control and immediate adjustments—both are indispensable. It is clear that the embodied sector has moved past the stage of simply comparing model concepts and has entered the deep waters of architectural design and engineering implementation.The debate over model routes will persist, but what truly determines success is the ability to organize the strengths of different models into a system that is deployable, iterable, and scalable.Multi-modal Perception Must Enter the Physical Interaction LoopAs the focus shifts from model debates to layered fusion, a threshold closer to the physical world emerges. The industry has relied heavily on vision because it dominates human perception and the technology is mature. However, embodied AI isn't just about seeing the world; it is about using hands—or end-effectors—to change it.Once tasks involve precision assembly, flexible grasping, or surface finishing, vision can only offer preliminary judgment. It cannot replace the feedback from force and touch during contact. The critical question in moving from "seeing" to "doing" is whether multi-modal perception can integrate into the control loop, rather than remaining a feature for demonstration purposes.Liao Yongxing highlighted a common dilemma based on his work with tactile data: how to integrate touch or force into existing robotic models. Often, the industry’s understanding of tactile sensing remains stuck at a binary level—simply detecting presence or absence—rather than capturing the nuances of real contact: changes in elasticity, direction, and reaction to an object's features.As sensor technology evolves, the primary focus of data fusion has shifted to reliability, consistency, and synchronization. It is essential to ensure that sensors provide data that is sufficiently consistent and accurate, and that they synchronize effectively with other modules. "If these two goals can be met," Liao said, "integrating data with embodied AI will proceed much more smoothly."Visual data can be collected en masse as images or video, but tactile data is distributed, continuous, high-frequency, and tightly linked to the state of contact. If sensor consistency is lacking or time synchronization is off, tactile information cannot enter the control loop; it remains relegated to an auxiliary signal for post-judgment.Lu Yao added to the necessity of tactile and force sensing from the perspective of learning mechanisms. Embodied AI operates in the physical world and requires an action system to interact with it. Data is essential here as it reflects the true state of that world—whether through video or the data morphology required for actual contact operations. Lu Yao drew a parallel to human infants: before they develop long-term reasoning, they explore by touching, grasping, and feeling, gradually collecting signals through interaction and summarizing patterns.Embodied AI should follow the same path: placing models in interactive environments to collect sufficient physical signals for learning.Currently, most embodied data is standardized around images, actions, and body states, with force data remaining relatively scarce. In this data-scarce environment, Lu Yao proposed the concept of explicit states versus implicit results. Explicit states relate to the actual mechanical arm joints, velocity, and contact force. The goal is to inject knowledge already compressed from human experience—such as dynamic equations and system models—directly into the learning process. This allows the model to learn from less data by standing on human shoulders. Meanwhile, data like images is inferred via neural networks and aligned with large-scale empirical knowledge.This approach seeks a middle ground between data-driven and physics-based models. Pure data-driven methods require massive samples, while pure physics models struggle to cover complex, open scenarios. Embedding dynamic priors into the learning process could be a crucial path to reducing data dependency and enhancing generalization.However, integration into the model requires validation through the control loop. The reason tactile and force sensing are hard to replace is that they participate directly in real-time feedback during contact. At the control level, touch isn't just an additional modality; it is the critical link determining whether an action can be completed stably. As Zhuo Weifeng noted, robots use sensor data for perception, then make decisions based on predictive information, followed by behavioral planning and control. Ultimately, the end-effector must provide feedback to the real world and the interaction.Without tactile sensing, a robot relies on vision to predict contact states; once a deviation occurs, it cannot adjust in time. With tactile and force feedback, the robot can instantly sense whether it has a firm grip, if it is slipping, or if it is applying too much force.The value of tactile sensing in the control loop is most evident in precision operation scenarios. Ji Haifeng illustrated this with the example of precision assembly. "Building robots is essentially the process of humans recreating themselves," he said. Humans primarily observe the world through their eyes, with visual information accounting for perhaps 80% of input—which is why 'V' in VLA stands for Vision, as it is the most used and mature modality. However, in precision assembly scenarios, tasks simply cannot be completed without tactile and force sensing. If a threaded hole is just 5 millimeters wide, being off by even 1 millimeter or 0.1 millimeter means the screw won't go in. That is when impedance control is needed, relying on force and tactile perception to assist in completing the operation.For multi-modal perception to truly enter the control loop, data must be stably collected, synchronized, annotated, and used for training. The difficulty of integrating tactile and force data isn't just a sensor or algorithm issue; it is a data supply issue. How do we divide labor among real-world, simulation, and human data? How do we balance them at different training stages? How do we form a loop from scenario to model and back to scenario? These questions determine whether perception fusion can evolve from a localized function into a systemic capability.Without a data loop, tactile data remains stuck in single-point verification. With a loop, model layering and perception fusion gain the foundation needed for continuous iteration.Data Strategy Isn't a Single Choice EitherAs multi-modal perception attempts to enter the physical interaction loop, the data problem can no longer be sidestepped. Fusing vision, touch, and force requires a vast amount of high-quality samples, yet what embodied AI currently lacks most is precisely real-world interaction data.Simulation data can cover a broader range at a lower cost. Real-world robot data is expensive, but it captures the most authentic interaction with the physical world and remains the most reliable. Human-centric data, meanwhile, compensates for gaps by mimicking human methods, covering a wider scope with greater ease and lower cost.Humans have accumulated operational experience from childhood—picking up a microphone, folding clothes, handling tools. These are real skills. Liao Yongxing believes that as simulation data matures, the next major role will be played by human-centric approaches: collecting data on real human operational skills, then using a small amount of real-world robot data for verification and adjustment to complete the transfer of skills from human to robot.Regarding the value of simulation, Lu Yao noted that different data sources and even different data qualities offer distinct guidance. He cautioned against a common misconception: viewing simulation merely as a low-cost substitute for real-world data.If simulation is used only as a replacement, you hit the next bottleneck: does the simulation accurately reflect the real-world situation? Lu Yao argues that simulation is a significant source of data generation, and its value lies in generating a wide variety of actions within the same state at scale.Simulation acts as a scaffold for dynamics. Through this method, the model can learn the underlying transfer laws of the world. Once these systemic laws are distilled, they can be aligned to the specific dynamic parameters and characteristics of the real robot using real-world data, giving the model strong generalization capabilities.Therefore, different data sources and qualities—including both successful and failed samples—hold unique and critical value. They must be supplemented into the model learning process with varying ratios at different training stages, ultimately achieving true general intelligence through a comprehensive approach.Simulation is not a cheap substitute; it is a generator of laws. Its true value is not in mimicking real robots, but in generating massive, controlled variations that allow the model to learn the underlying laws of the physical world, which are then aligned with specific parameters using real-world data.Yet, a clear classification of data does not equate to a clear path to deployment. Once the industry has defined the respective roles of real-world, simulation, and human data, a thornier problem emerges: why, with all this data, does embodied AI still struggle to move from demo to scale? The bottleneck is clearly not just the data itself.Liao Yongxing believes it isn't a failure of a single link, but a combination of factors—all presenting resistance.If one had to rank them, the first issue is data volume. The amount of real-world interaction data is genuinely insufficient. Compared to the data volumes used for language models or visual data, the data for embodied AI's interaction with the physical world is on a completely different scale. Data scarcity is the first bottleneck.Next is the algorithm layer. How can we better integrate multi-modal data—such as tactile or force knowledge—into the decision-making or motion planning layers of embodied AI? Enabling the system to better understand the world and execute both fast and slow thinking and reactions represents another bottleneck at the algorithmic level.At the same time, a systemic loop is required: how to apply data to the hardware, give the robot a degree of intelligence through the model, put it to work in a scenario, and then collect more data from that work to continuously improve its intelligence.Without a loop, data is an island, the model is an exhibit, and the robot is a prototype. With a loop, the robot can continuously accumulate capabilities in real-world scenarios.Summary: The key breakthrough for embodied AI, from seeing to doing, lies not in a single-point technology but in system capability.At the model level, the industry is shifting from a battle of routes to layered fusion. Combining slow thinking with fast reaction, semantic understanding, physical prediction, motion control, and real-time feedback are forming stable interfaces.At the perception level, vision remains the most mature modality, but precision operations and physical interaction require tactile and force sensing to enter the control loop. Multi-modal fusion must resolve issues of reliability, consistency, synchronization, and the embedding of physical priors.At the data level, real-world, simulation, and human data each hold irreplaceable value. Simulation provides law generation and large-scale variation; human data offers skill priors and contact interaction; real-world data provides final verification and a closed loop with real scenarios. The key to data lies in training ratios and phased application.At the scaling level, bottlenecks exist in data volume, algorithms, hardware, computing power, and talent, but the most significant shortcoming is the lack of a systemic loop. Only by forming an autonomous iteration loop among the hardware, data, models, and scenarios can embodied AI truly move from demo to scale.The next phase of competition in embodied AI will not be merely about model parameters or hardware form factors, but about the efficiency of system iteration. Whoever can most rapidly convert real-world scenario data into model capabilities, and then feed those capabilities back into the scenario for verification and optimization, will be the first to cross the threshold from seeing to doing.