Gasgoo Munich-If robots truly need 100 million hours of data, a practical question looms: Who will produce it?In our first installment, we noted that Maniformer CEO Yao Maoqing pegged data requirements for the next stage of embodied intelligence at tens of millions—or even hundreds of millions—of hours.But scaling from 1 million to 100 million isn’t just a matter of adding two zeros.It means the current data production infrastructure must expand a hundredfold.During a group interview on August 31, Yao acknowledged that linearly replicating the existing model is "extremely difficult"—a complete overhaul of operations is necessary.And one answer rising to the surface is clear: Get more ordinary people involved.Gasgoo Embodied Intelligence learned on the scene that Mifeng’s "Mifeng Pai" app has already launched. Once fully open to the public, ordinary users can apply to rent collection gear, claim tasks through the app, perform the required actions, and upload valid video and data to earn cash rewards upon approval.It sounds a lot like another kind of gig platform.Except that ride-hailing drivers sell a trip, and food delivery couriers complete a drop-off.This time, what’s being traded is human movement, skill, and experience.Moreover, this is no longer an isolated experiment by a single Chinese firm.Just six days ago, the U.S. humanoid robotics company Figure unveiled a project with a very direct name: Index.Figure had been running the system in secret for four months. Ordinary people could download the app, claim recording tasks, and perform real-world chores like cooking, cleaning, laundry, or shelf organization at home or work, getting paid for the data they uploaded.By August 25, Index had reached 108 countries with 264,000 app downloads. More than 44,000 weekly active users had uploaded over 16 million videos, and Figure had paid $15 million to data contributors. The company also announced plans to invest more than $1 billion in data and computing power over the next 12 months.As China’s Mifeng and JD.com, along with the U.S.-based Figure, turn their attention to ordinary people almost simultaneously, a trend is becoming clear: Robot data is shifting from "factory production" to "social production."From "Building Data Factories" to "Mobilizing People"How was robot data produced in the past?The most typical method was building a data collection factory. Companies would prepare robots, teleoperation gear, operators, and simulated environments, having operators repeat tasks like grasping, moving, and organizing.In terms of production organization, this was essentially traditional manufacturing. Equipment, workers, and facilities were all centralized; the only difference was that the final product was data rather than parts.This model won’t disappear. Yao even emphasized that many well-run physical data factories are still operating at full capacity, with some projects requiring two shifts a day.The real problem is that scaling it infinitely isn’t easy.You can build 10 more factories. But if data demand suddenly increases a hundredfold, you can’t simply build a hundred times the facilities, buy a hundred times the robots, or hire a hundred times the teleoperators.So, data companies are shifting their thinking: Why must we move people and scenes into data factories? Can’t we send the collection gear to where people are already working?Restaurant staff are already wiping tables; warehouse workers are already sorting goods; households do laundry, wash dishes, and cook every day; maintenance technicians are already dismantling equipment. If these naturally occurring actions could be recorded synchronously, the entire real world could itself become a massive data factory.This is also why Mifeng’s equipment is starting to leave professional BPO and data collection bases.Simply wear the device, connect to the app, and start accepting tasks for data collection.Yao revealed that the devices are already widely deployed. "Many retirees and stay-at-home moms in communities" have started helping generate data, while companies are largely deploying equipment through local partners who have access to specific scenes and human resources.This isn’t just about adding a few more collectors; what’s truly changing is the method of production organization.A Nationwide Data Collection Experiment Has BegunIn fact, this kind of exploration in China isn’t limited to Mifeng.In May this year, an embodied intelligence data collection community jointly built by JD.com and Suqian officially launched.Local residents can wear JoyEgoCam devices to record data during normal household chores. According to JD.com’s plan, Suqian alone aims to mobilize over 100,000 citizens, covering more than 100 specific scenarios including homes, offices, factories, logistics, stores, and sanitation.JD.com’s broader goal is to mobilize over 100,000 internal employees and 500,000 external workers across various industries. It aims to accumulate more than 10 million hours of real-world scenario data within two years.Another experiment is underway in Huishui, Guizhou.There, ordinary residents fold quilts, wipe tables, and sort vegetables at home, while hotel staff collect data during actual room service. According to local disclosures, collection scenarios now cover 8 major industry categories and over 50 real-world settings, delivering tens of thousands of hours of high-quality data annually.The difference is that this model isn’t solely organized by private enterprise.The region is trying a model of "government leadership, state-owned platform operation, and enterprise as scenario operators." It uses administrative and industrial resources to coordinate access to communities, supermarkets, factories, and medical institutions.This is intriguing. These three cases actually represent three emerging organizational models for embodied data:Mifeng is closer to platform-based crowdsourcing; JD.com leans toward corporate scenarios plus community mobilization; Huishui adds government and local resource coordination. The paths differ, but they solve the same puzzle: how to decentralize data production, once concentrated in factories, out into the real world.Figure’s Index proves this isn't a uniquely Chinese challenge; it may well be an infrastructure restructuring facing the global embodied intelligence industry.What’s Truly Scarce Isn’t People, But "Access Passes"But the real difficulty in nationwide data collection may not be finding enough people—China certainly doesn’t have a labor shortage.Figure has already proven that an app can quickly organize tens of thousands of contributors.The real challenge is: Where can these people go to collect?Homes have privacy; businesses have trade secrets. Hospitals, factories, warehouses, and even nursing homes all involve distinct data security and third-party rights issues.Just because a mechanic knows how to repair a car doesn’t mean they can wear a camera into the workshop and upload the entire process to a third-party platform.Thus, there is a fundamental difference between embodied data and traditional internet crowdsourcing.An image labeler only needs a computer; an embodied data collector needs access to the physical world.This explains why Huishui needs local government help to open up scenarios, and why JD.com’s advantage largely stems from owning real-world businesses like logistics and retail.Yao also admitted in the group interview that operating "no-body" data collection might be even more complex than running a fixed factory. Equipment must circulate across different industries, cities, and people; workers need training; and the problem of accessing real-world scenarios must be solved.So, while it looks like "no-body" collection dismantles the factory, it actually transforms the management problem of a single facility into a coordination problem across thousands of real-world scenarios.Centralized production decreases, but social organization costs rise.This may be the true moat for future crowdsourcing platforms. It’s not about who can build an app, but who can consistently secure "passes" to the real world—factories, restaurants, homes, malls, warehouses.When Human Actions Become CommoditiesAnother deeper question is: Who actually owns this data?Today, this question isn’t entirely without a regulatory framework.The Personal Information Protection Law stipulates that when personal consent is the basis for processing personal information, voluntary and explicit consent must be obtained from individuals who are fully informed. The Data Security Law also requires that data collection be done through legal and proper means.The 2022 "Opinions of the CPC Central Committee and The State Council on Building a Data Basic System to Better Play the Role of Data Elements" is known as the "Data Twenty Articles." This document further proposes exploring a "structural separation" of data rights. This means data resource holding rights, data processing and usage rights, and data product operation rights can be held by different entities, while emphasizing the protection of the legitimate rights and interests of data sources.But these remain relatively high-level regulatory frameworks.When embodied data truly enters real-world labor scenarios, new questions quickly become concrete.Take a maintenance worker performing a task at a company. The skill belongs to the worker; the scenario belongs to the enterprise; the equipment is provided by the data company; the data is processed by the platform; and finally, it’s purchased by a robotics company to train a model.So, how should the value generated by this data be distributed among these parties?Data companies are already exploring their own commercial rules.Yao introduced that there are currently two main methods of data trading. One is selling usage rights, where the data provider retains ownership and the customer cannot resell it commercially. The other is an exclusive buyout, where ownership transfers and the provider must delete the corresponding data from their servers.But this primarily addresses the relationship between data companies and model clients.If vast amounts of data come from ordinary people in the future, another question will eventually need an answer. Is the person contributing the action just a gig worker picking up tasks? Or are they part of the data value chain?There is no mature answer to this question yet.A "Robot Uber" Model Can't Run Just YetTherefore, it may be too early to directly compare nationwide data collection to the next generation of food delivery or ride-hailing.Figure has 264,000 app downloads; JD.com is preparing to mobilize hundreds of thousands; Mifeng has already deployed devices into communities.It seems people are no longer the biggest bottleneck. What’s truly missing is a complete set of platform rules.Where will tasks continuously come from?Who decides how much different skills are worth?Who handles equipment rental, maintenance, and retrieval?How do families, businesses, and third parties grant authorization?How are profits distributed between the platform and the collector after a data transaction?Until these questions are resolved, "working for robots" looks more like an emerging mode of production than a mature new profession.There isn’t even a unified price for data itself yet.Yao stated in the interview that there is no standard pricing for data across different scenarios currently; the scarcity of the scenario, the difficulty of access, and delivery timeliness all affect the final price.This is also why defining the business model of embodied data crowdsourcing by a project paying "15 yuan per hour" or "120 yuan a day" is inaccurate. Even if similar quotes appear in the industry, they primarily represent how much a specific project is willing to pay for a certain type of collection labor today. This does not mean a stable labor price has formed for embodied data.What’s truly worth watching now is how a new market mechanism is trying to take shape.In the past, once a chef finished a dish, the action vanished instantly.A mechanic tightening a screw left behind no reusable "digital asset."But in the era of embodied intelligence, these fleeting human skills can, for the first time, be continuously recorded and become raw material for machine learning.So, the most noteworthy aspect of embodied data crowdsourcing isn’t "how much a stay-at-home mom can earn a day." Instead, it is a larger shift: human physical experience is being incorporated into the production system of the AI industry.From data factories to communities, and then to the homes and workplaces of ordinary individuals, this migration has only just begun.But as more and more human actions are converted into data, the next question becomes unavoidable: how can data generated by so many different people, in different environments, using different devices, actually be used by robots?Being able to produce it doesn’t mean being able to train with it.