Human demonstration data acquisition system and method based on robot-like first visual angle
By using a human demonstration data acquisition system based on a robot-like first-person perspective, and combining head and limb data acquisition modules with inertial measurement and multi-camera technology, the problems of high cost and inconsistent perspective of remote-operated equipment are solved, achieving efficient and natural data acquisition and supporting the training of embodied intelligent models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU WUZHIBO TECHNOLOGY CO LTD
- Filing Date
- 2026-04-10
- Publication Date
- 2026-05-15
AI Technical Summary
In existing technologies, remote-operated devices are costly and slow in data acquisition. Furthermore, the perspective of wearable or handheld devices is inconsistent with that of the robot when performing tasks, resulting in low data acquisition efficiency and making it difficult to conduct efficient natural data acquisition in open environments.
A human demonstration data acquisition system based on a robot-like first-person perspective is adopted. Through head and limb data acquisition modules, combined with inertial measurement and multi-camera technology, the system can realize the synchronous acquisition and attitude calibration of global and local image data, construct a unified world coordinate system, and ensure the naturalness of hand operations and the continuity of data.
It enables efficient and interference-free data acquisition in open environments, eliminates visual domain bias, improves the accuracy and speed of data acquisition, supports efficient data support for embodied intelligence, and is suitable for training embodied models, visual-language-action models, and world models.
Smart Images

Figure CN122033986A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and robot control technology, specifically relating to a human demonstration data acquisition system and method based on a robot-like first-person perspective. Background Technology
[0002] In the fields of embodied intelligence and robot learning, the improvement of model capabilities highly depends on large-scale, high-quality operational data. Currently, the industry widely adopts teleoperation to collect robot execution data, but this method has the following problems: teleoperation equipment is usually expensive, the system deployment is cumbersome, and the data collection process relies heavily on slow fine-tuning by professionals, resulting in high cost and slow speed of data collection, which cannot meet the demand for massive amounts of data in embodied intelligence scaling. Therefore, directly collecting natural demonstration data from humans in real-world environments has become a highly promising and efficient data acquisition method.
[0003] However, existing human demonstration data acquisition methods and equipment still have significant limitations. On the one hand, the visual images acquired using wearable smart glasses are limited to a first-person perspective from the top of the human head. This not only lacks the microscopic details of finger-to-object interaction but also completely differs from the observation perspective of actual dual-arm robots performing tasks (which typically rely on the collaboration of multiple cameras mounted on the head and the grippers at the ends of the arms). This severe visual domain bias leads to a sharp drop in the generalization performance of the model when deployed to real robots. On the other hand, handheld data acquisition devices, such as UMI (Universal Manipulation Interface), require the operator to hold a specific device with cameras and sensors for recording. This method severely disrupts the natural operating habits of the human hand, reducing the efficiency of acquiring data for complex and detailed tasks. Furthermore, the equipment is cumbersome and makes it difficult to conduct routine, interference-free generalization data collection in complex in-the-wild environments.
[0004] Therefore, there is an urgent need for a bare-hand demonstration data acquisition system that can simulate the real multi-camera observation perspective of a robot, is lightweight and flexible, and does not require handheld devices. Summary of the Invention
[0005] To address the shortcomings of existing technologies and achieve the goals of improving data acquisition accuracy and efficiency, and enhancing generalization performance, this invention adopts the following technical solution:
[0006] A human demonstration data acquisition system based on a robot-like first-person perspective includes a head data acquisition module and a limb data acquisition module. The system also includes a demonstration data generation module.
[0007] The demonstration data generation module uses the head inertial measurement data acquired by the head data acquisition module and global image data from at least two limbs to execute visual and inertial algorithms to calculate the pose transformation matrix of the head data acquisition module in the world coordinate system; it then performs limb tracking using the global image data to obtain the limb pose in the head data acquisition module coordinate system; based on the pose transformation matrix and the position information, it performs a homogeneous transformation to convert the limb pose in the head data acquisition module coordinate system to the world coordinate system; finally, it acquires individual local image data of the limbs acquired by the limb data acquisition module, and constructs demonstration data based on the limb pose in the world coordinate system, the global image data, and the local image data.
[0008] The head-mounted camera and IMU work together to calibrate the hand's end-effector posture. The limb-mounted camera primarily captures images of the hand, while the head-mounted camera locks onto the overall hand workspace and the global position of the object. The limb-mounted camera focuses on the fine interaction area between the fingers and the object. The head-mounted camera uses a wide-angle lens and is positioned centrally in front of the forehead to ensure that the hands remain within the complete field of view in the normal workspace in front of the operator when the head is naturally lowered. Simultaneously, the IMU monitors head posture in real time. If the head tilt angle exceeds a preset threshold, potentially causing the hands to leave the field of view, the system automatically issues a voice or vibration prompt to guide the operator to quickly adjust their posture, ensuring that the head-mounted camera stably captures the entire hand operation. The head-mounted camera locks onto the overall hand workspace and the global position of the object, while the limb-mounted camera focuses on the fine interaction area between the fingers and the object. The two cameras achieve dynamic complementarity across multiple scenarios, avoiding blind spots or loss of interaction details from a single perspective. Ultimately, multiple sensors jointly acquire visual and posture information about the hand.
[0009] Furthermore, the pose transformation involves multiplying the three-dimensional coordinates of the limb in the coordinate system of the head data acquisition module with the calculated time-by-time pose transformation matrix of the head data acquisition module in the world coordinate system to obtain the three-dimensional coordinates of the limb in the world coordinate system. This time-by-time, frame-by-frame pose transformation mechanism fundamentally solves the core problem of the continuous dynamic change of the coordinate system origin and posture caused by the free movement of the head data acquisition module with the operator's head: without pre-fixing the head position or restricting the operator's head movement, the relative hand poses acquired at different times and under different head postures can be uniformly mapped to the globally static world coordinate system, ensuring the spatial consistency and continuity of hand pose data throughout the entire demonstration process;
[0010] Furthermore, the collected global image data, local image data, and head inertial measurement data are time-aligned before the algorithm is executed to ensure time consistency across all acquisition channels.
[0011] Furthermore, the global image data, the local image data, and the limb poses in the world coordinate system are encapsulated by timestamps to form a demonstration dataset.
[0012] Furthermore, the head data acquisition module includes a head camera and an inertial measurement unit. The head camera is used to acquire global top-down image data of the hand, and the inertial measurement unit is used to acquire the velocity and acceleration of the head.
[0013] The extremity data acquisition module is installed on a single hand and is used to acquire local image data of the hand and / or the opposite hand.
[0014] The human demonstration data acquisition method based on a robot-like first-person perspective employs the aforementioned human demonstration data acquisition system based on a robot-like first-person perspective. It utilizes head inertial measurement data and global image data from at least two limbs acquired by the head data acquisition module, and individual local image data from the limbs acquired by the limb data acquisition module. The demonstration data generation module calculates the pose transformation matrix of the head data acquisition module in the world coordinate system and performs a homogeneous transformation based on the tracked limb pose. The resulting limb pose in the world coordinate system, combined with the global image data and the local image data, constructs demonstration data.
[0015] A human demonstration data acquisition system based on a robot-like first-person perspective includes a head data acquisition module and a limb data acquisition module. The system also includes a calibration tag and a demonstration data generation module.
[0016] The calibration tag has multiple tag corner points and a specific physical size, and at least one calibration tag is deployed in the calibration scene;
[0017] The demonstration data generation module acquires global image data of at least two limbs through the head data acquisition module, identifies the calibration labels from them, sets the world coordinate system of the calibration scene based on the calibration labels, and acquires the position information of multiple label corner points in the calibration labels in the coordinate system of the head data acquisition module. Combining the known physical dimensions of the calibration labels, a multi-point perspective pose algorithm is used to calculate the pose transformation matrix of the head data acquisition module in the world coordinate system. Using the global image data, limb tracking is performed to obtain the pose of the limbs in the head coordinate system. Based on the pose transformation matrix and the position information, a homogeneous transformation is performed to convert the limb pose in the head data acquisition module coordinate system to the world coordinate system. Individual local image data of the limbs acquired by the limb data acquisition module is acquired. Based on the limb pose in the world coordinate system, the global image data, and the local image data, demonstration data is constructed.
[0018] Furthermore, the limb pose in the world coordinate system is strictly aligned with the global image data and the local image data according to the timestamp to obtain a temporally consistent demonstration dataset.
[0019] Furthermore, the demonstration data generation module is equipped with an exception handling mechanism. When the calibration label is not recognized, the pose conversion is paused, the valid data of the previous frame is retained and a prompt is issued. After the calibration label is recognized, the normal pose conversion is resumed to avoid data loss or errors.
[0020] The human demonstration data acquisition method based on a robot-like first-person perspective employs the aforementioned human demonstration data acquisition system based on a robot-like first-person perspective. It acquires global image data from at least two limbs through the head data acquisition module, identifies calibration labels containing multiple corner points, establishes a world coordinate system for the calibration scene based on these labels, and acquires the position information of the multiple corner points in the head data acquisition module's coordinate system. Combining this with the known physical dimensions of the calibration labels, it calculates the pose transformation matrix of the head data acquisition module in the world coordinate system. This matrix is then combined with the tracked limb poses and subjected to homogeneous transformation to obtain the limb poses in the world coordinate system. Finally, it combines the global image data and the local image data to construct demonstration data.
[0021] The advantages and beneficial effects of this invention are as follows:
[0022] The architecture and data acquisition mechanism proposed in this invention eliminate the visual domain deviation between human demonstration data and robot execution data at the physical source, constructing a highly consistent robot-like first-person perspective. Due to the fully wearable hardware design, this invention does not rely on handheld intervention devices such as handles or UMIs, allowing operators to perform natural operations directly with their bare hands without the burden of additional handheld devices. Limb movements do not need to adapt to intermediate control logic such as handles or buttons, and the movement forms are more in line with human instinctive operating habits, intuitive and natural, and in line with operational intuition. At the same time, it eliminates redundant steps such as holding, controlling, and posture transition of handheld devices, and the acquisition movements are smooth and uninterrupted, effectively improving the data acquisition speed. Furthermore, the field of view setting of the limb-end camera ensures that it can capture images of the end of the arm, effectively ensuring the complete capture of grasping details and spatial pose. In addition, the system acquisition module and aggregation module support wireless networking, are concealed and lightweight, and are not limited by fixed experimental platforms, allowing them to be integrated into daily life and realize large-scale, interference-free human demonstration data acquisition in open and unconstrained environments. The data output by this invention is widely compatible with various SLAM and hand tracking algorithms, fully supporting the diverse needs of supervised robot control model training with precise hand trajectory calibration, as well as unsupervised robot control model training without trajectory calibration, providing efficient data support for embodied intelligence scaling. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the structure of a human demonstration data acquisition system based on a robot-like first-person perspective in an embodiment of the present invention.
[0024] Figure 2 This is a diagram illustrating the execution process in an open, uncalibrated scenario according to an embodiment of the present invention.
[0025] Figure 3 This is a diagram illustrating the execution process in an indoor calibration scenario according to an embodiment of the present invention. Detailed Implementation
[0026] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0027] To address the problems of high cost and slow speed in existing robotic remote-controlled data acquisition technologies, as well as the limited field of view of existing wearable or handheld acquisition devices, which affect natural operation and are difficult to generalize in open environments, this invention proposes a human demonstration data acquisition system based on a robot-like first-person perspective. This system acquires human demonstration data from a robot-like first-person perspective (RoboEgo-Centric), and can be applied to embodied models, visual-language-action models, and world models, etc. Figure 1 As shown, the system mainly includes a head-mounted data acquisition module, at least two limb-mounted data acquisition modules, a demonstration data generation module, and a data receiving and transmission module.
[0028] The head-mounted data acquisition module, which includes a head-mounted camera with a fixed IMU (Inertial Measurement Unit) worn on the operator's head, is used to simulate the robot's head vision and acquire global top-down image data.
[0029] Two wearable data acquisition modules are fixed to the left and right wrists or forearms of the user, respectively. Each module integrates a wrist-mounted top-down camera, whose field of view flexibly covers the operator's hand end, the entire hand, or an area including the wrist. This simulates the vision above the grippers of the robot's arms, acquiring local top-down image data of finger-object interaction. A head camera and an IMU work together to calibrate the hand end's posture, while the end-user camera primarily captures hand observation images. Simultaneously, the head camera locks the overall hand workspace and the global position of objects, while the end-user camera focuses on the fine interaction area between fingers and objects. The fields of view of these two modules dynamically complement each other, avoiding blind spots or loss of interaction details from a single perspective. Ultimately, multiple sensors jointly acquire visual and posture information about the hand.
[0030] In this embodiment of the invention, the system mainly comprises three acquisition nodes. The head node is worn on the operator's forehead via a headband and integrates a wearable head camera and a high-frequency IMU. The left and right wrist nodes are fixed to the operator's forearms near the wrists via flexible straps and integrate wrist-mounted top-view cameras. The camera lenses are pointed towards the fingertips, and adjusted according to the specific lens focal length to cover the entire area from the wrist to the fingertips. When the operator performs actions such as grasping, the head camera records the overall top-view scene and the relative positions of the hands, while the wrist cameras record the microscopic details of the interaction between the fingers and objects. This configuration strictly corresponds to the hardware layout of a typical dual-arm robot in spatial geometry, achieving a RoboEgo-Centric perspective.
[0031] The demonstration data generation module collects data from multi-view observation videos simultaneously, connects relevant computer vision algorithms to perform pose calculations in a unified coordinate system, and generates a multi-view demonstration dataset for robot training.
[0032] Specifically, in open, uncalibrated scenarios, the acquired head camera vision and IMU information can be used to obtain the world coordinate system pose through a visual-inertial SLAM algorithm. Combined with the 3D hand pose extracted by the hand tracking algorithm, a homogeneous coordinate system transformation is performed. The specific execution process is as follows:
[0033] First, define the core variable: the world coordinate system is... The head camera coordinate system is , Head camera coordinate system In the world coordinate system The pose transformation matrix in the middle, Hand in head camera coordinate system The three-dimensional coordinates below For the hand in the world coordinate system The three-dimensional coordinates below For the first Frame data acquisition timestamp , , They are respectively Real-time global top-down image of the head, local top-down image of the extremities, and IMU acceleration and angular velocity data.
[0034] Then, the head images, limb images, and IMU data are time-aligned to ensure time consistency across all acquisition channels.
[0035] Then, and Inputting the visual-inertial SLAM algorithm, the time-by-time pose transformation matrix of the head camera in the world coordinate system is obtained. Used for subsequent hand coordinate transformation, and simultaneously for the global top-down view of the head image. Perform a hand tracking algorithm to obtain the three-dimensional coordinates of the hand in the head camera coordinate system. Then, through the core homogeneous transformation formula The 3D hand pose in the head camera coordinate system is transformed to the world coordinate system to obtain globally consistent 3D hand coordinates. Finally, the synchronously acquired multi-view images and the 3D hand pose sequence in the world coordinate system are encapsulated by timestamp to generate a demonstration dataset for robot training. ,in This represents the total number of data frames collected.
[0036] In embodiments of the present invention, such as Figure 2 As shown, the operator can perform unconstrained hand pose acquisition based on IMU and SLAM in outdoor environments or unfamiliar rooms without calibration objects. The specific execution process is as follows:
[0037] 1. After the system starts, it synchronously acquires a global top-down view of the head and a partial top-down view of the extremities;
[0038] 2. Align and register multi-channel visual and IMU sensor data based on hardware timestamps;
[0039] 3. Perform pose calculation steps:
[0040] First, it receives images from the head camera and IMU data (e.g., IMU output angular velocity, linear acceleration, etc.), and runs a vision-inertial SLAM algorithm, using the world coordinate system as the reference. The head camera coordinate system is The SLAM algorithm, based on head-mounted camera images and IMU data, can output in real time. exist pose transformation matrix in ;
[0041] Simultaneously, by utilizing existing hand tracking algorithms, the joints of both hands are identified in the head camera images, and the three-dimensional coordinates of the hands in the head camera coordinate system are obtained. ;
[0042] Then, through the formula Calculate globally consistent world coordinates for the hand. .
[0043] Finally, a demonstration dataset containing multi-view vision and human hand pose is generated, and the model is trained using hand data in a unified coordinate system.
[0044] In open, uncalibrated scenarios, this invention also provides a human demonstration data acquisition method based on a robot-like first-person perspective. Using the aforementioned human demonstration data acquisition system based on a robot-like first-person perspective, the method acquires head inertial measurement data and global image data from at least two limbs through a head data acquisition module, and individual local image data from the limbs through a limb data acquisition module. A demonstration data generation module calculates the pose transformation matrix of the head data acquisition module in the world coordinate system, combines it with the tracked limb poses, performs a homogeneous transformation to obtain the limb poses in the world coordinate system, and combines the global and local image data to construct demonstration data.
[0045] In indoor calibration scenarios, without the need for an IMU, calibration tags (such as AprilTags) placed in the scene can be used. Algorithms can then be employed to enable the head camera to decode the tags and obtain its own pose in the world coordinate system, thereby completing the coordinate transformation of the hand pose. The specific execution is as follows:
[0046] First, evenly distribute several AprilTag visual calibration tags in a fixed, accessible, and unobstructed area of the indoor calibration scene (such as table corners, walls, etc.), and ensure that the head camera can capture at least two or more complete tags in any posture within its operating range.
[0047] Then, based on the reference coordinate system of any calibration label, a unified world coordinate system is defined for the entire scene. <w>The label corner points are clearly defined in the world coordinate system, and their three-dimensional coordinates are pre-calibrated based on their known physical dimensions, serving as a reference for subsequent pose calculation.
[0048] Subsequently, the head-mounted camera module continuously acquires a global top-down view of the scene, ensuring that at least one complete AprilTag is clearly captured in each frame, avoiding tag occlusion, blurring, or exceeding the field of view. At the same time, each frame image undergoes preprocessing such as grayscale conversion, Gaussian denoising, and edge enhancement to improve the accuracy of tag corner point recognition.
[0049] Next, the PnP (Perspective-n-Point) algorithm combined with AprilTag recognition is used to complete the head camera pose calculation. Specifically, the two-dimensional pixel coordinates of the tag corner points in the head camera image coordinate system are first extracted, and then the corresponding relationship is constructed by combining the known physical size of the tag and the three-dimensional coordinates in the world coordinate system to solve the problem including the head camera coordinate system. Relative to the world coordinate system pose transformation matrix If multiple labels are captured, the nonlinear least squares method is used to fuse and optimize the solution results to improve stability.
[0050] Subsequently, based on the preprocessed global image from the head camera, a hand detection and tracking algorithm is run to accurately identify the hand region and extract key hand nodes in the head camera coordinate system. 3D pose coordinates Then, using the homogeneous coordinate transformation formula... Transform the hand pose from the head camera coordinate system to the world coordinate system. Obtain the global pose of the hand This achieves global consistency of pose data.
[0051] Subsequently, the hand pose data is strictly aligned with the head camera image and the hand partial image according to the timestamp, providing high-quality time-consistent data for subsequent applications.
[0052] Finally, an exception handling mechanism is set up so that when the camera fails to capture enough tags or tag recognition fails, the pose conversion is paused, the valid data of the previous frame is retained, and a prompt is issued. The normal conversion process is resumed after the tag recognition stabilizes, thus avoiding data loss or errors.
[0053] In embodiments of the present invention, such as Figure 3 As shown, in certain workbench scenarios with defined boundaries, the operator may not need to configure or use an IMU. Instead, several AprilTags can be affixed to the corner of the table or the wall as visual calibration labels. Lightweight hand pose acquisition can be performed based on these scene-marked labels. The specific execution process is as follows:
[0054] 1. Pre-place visual labeling tags (AprilTag) at fixed locations in the operational scenario;
[0055] 2. When the head camera captures images, it acquires a global image containing calibration tags;
[0056] 3. The reference coordinate system defined by the AprilTag label is the world coordinate system. By identifying tags using algorithms such as PnP and combining the tag's physical dimensions, the pose of the head camera relative to the calibrated tag AprilTag in the world coordinate system can be directly calculated. ;
[0057] 4. The global top-down view of the head can obtain the position of the hand in the head coordinate system through the hand tracking algorithm, and convert it into the 3D pose of the hand in the world coordinate system using the same matrix transformation logic.
[0058] This method has lower equipment requirements and consumes less computing resources, making it suitable for data collection in large-scale indoor fixed scenarios.
[0059] The aforementioned synchronous multi-view video streams and high-precision human hand pose sequence data with the same world coordinate system can be directly used for training robot control models, including but not limited to embodied models, visual-language-action models, and world models.
[0060] In an indoor calibration scenario, this invention also provides a human demonstration data acquisition method based on a robot-like first-person perspective. Using the aforementioned human demonstration data acquisition system based on a robot-like first-person perspective, the head data acquisition module acquires global image data from at least two limbs, identifies calibration labels containing multiple label corner points, sets the world coordinate system of the calibration scenario based on the calibration labels, and acquires the position information of multiple label corner points in the coordinate system of the head data acquisition module. Combining the known physical dimensions of the calibration labels, the pose transformation matrix of the head data acquisition module in the world coordinate system is calculated. Combined with the tracked limb poses, a homogeneous transformation is performed to obtain the limb poses in the world coordinate system. Finally, combining global and local image data, demonstration data is constructed.
[0061] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.< / w>
Claims
1. A human demonstration data acquisition system based on a robot-like first-person perspective, comprising a head data acquisition module and a limb data acquisition module, characterized in that: The system also includes a demonstration data generation module; The demonstration data generation module uses the head inertial measurement data acquired by the head data acquisition module and global image data from at least two limbs to execute visual and inertial algorithms to calculate the pose transformation matrix of the head data acquisition module in the world coordinate system; it then performs limb tracking using the global image data to obtain the limb pose in the head data acquisition module coordinate system; based on the pose transformation matrix and the position information, it performs a homogeneous transformation to convert the limb pose in the head data acquisition module coordinate system to the world coordinate system; finally, it acquires individual local image data of the limbs acquired by the limb data acquisition module, and constructs demonstration data based on the limb pose in the world coordinate system, the global image data, and the local image data.
2. The human demonstration data acquisition system based on a robot-like first-person perspective as described in claim 1, characterized in that: The homogeneous transformation is to multiply the three-dimensional coordinates of the limb in the coordinate system of the head data acquisition module with the calculated time-by-time pose transformation matrix of the head data acquisition module in the world coordinate system to obtain the three-dimensional coordinates of the limb in the world coordinate system.
3. The human demonstration data acquisition system based on a robot-like first-person perspective as described in claim 1, characterized in that: The algorithm is then executed after the collected global image data, local image data, and head inertial measurement data are time-aligned.
4. The human demonstration data acquisition system based on a robot-like first-person perspective as described in claim 1, characterized in that: The global image data, the local image data, and the limb poses in the world coordinate system are encapsulated by timestamps to form a demonstration dataset.
5. The human demonstration data acquisition system based on a robot-like first-person perspective as described in claim 1, characterized in that: The head data acquisition module includes a head camera and an inertial measurement unit. The head camera is used to acquire global top-down image data of the hand, and the inertial measurement unit is used to acquire the velocity and acceleration of the head. The extremity data acquisition module is installed on a single hand and is used to acquire local image data of the hand and / or the opposite hand.
6. A method for collecting human demonstration data based on a robot-like first-person perspective, characterized in that: The human demonstration data acquisition system based on a robot-like first-person perspective, as described in claims 1 to 5, utilizes head inertial measurement data and global image data of at least two limbs acquired by the head data acquisition module, and individual local image data of the limbs acquired by the limb data acquisition module. The demonstration data generation module calculates the pose transformation matrix of the head data acquisition module in the world coordinate system, combines it with the tracked limb pose, performs a homogeneous transformation to obtain the limb pose in the world coordinate system, and combines the global image data and the local image data to construct demonstration data.
7. A human demonstration data acquisition system based on a robot-like first-person perspective, comprising a head data acquisition module and a limb data acquisition module, characterized in that: The system also includes a calibration label and a demonstration data generation module; The calibration tag has multiple tag corner points and a specific physical size, and at least one calibration tag is deployed in the calibration scene; The demonstration data generation module acquires global image data of at least two limbs through the head data acquisition module, identifies the calibration labels from them, sets the world coordinate system of the calibration scene based on the calibration labels, and acquires the position information of multiple label corner points in the calibration labels in the coordinate system of the head data acquisition module. Combining the known physical dimensions of the calibration labels, a multi-point perspective pose algorithm is used to calculate the pose transformation matrix of the head data acquisition module in the world coordinate system. Using the global image data, limb tracking is performed to obtain the pose of the limbs in the head coordinate system. Based on the pose transformation matrix and the position information, a homogeneous transformation is performed to convert the limb pose in the head data acquisition module coordinate system to the world coordinate system. Individual local image data of the limbs acquired by the limb data acquisition module is acquired. Based on the limb pose in the world coordinate system, the global image data, and the local image data, demonstration data is constructed.
8. The human demonstration data acquisition system based on a robot-like first-person perspective according to claim 6, characterized in that: The limb pose in the world coordinate system is strictly aligned with the global image data and the local image data according to the timestamp to obtain a temporally consistent demonstration dataset.
9. The human demonstration data acquisition system based on a robot-like first-person perspective as described in claim 6, characterized in that: The demonstration data generation module is equipped with an exception handling mechanism. When the calibration label is not recognized, the pose conversion is paused, the valid data of the previous frame is retained and a prompt is issued. Normal pose conversion is resumed after the calibration label is recognized.
10. A method for acquiring human demonstration data based on a robot-like first-person perspective, characterized in that: The human demonstration data acquisition system based on a robot-like first-person perspective, as described in claims 7 to 9, acquires global image data of at least two limbs through the head data acquisition module, identifies calibration labels containing multiple label corner points, sets the world coordinate system of the calibration scene based on the calibration labels, and acquires the position information of the multiple label corner points in the coordinate system of the head data acquisition module. Combining the known physical dimensions of the calibration labels, the pose transformation matrix of the head data acquisition module in the world coordinate system is calculated. Combined with the tracked limb pose, a homogeneous transformation is performed to obtain the limb pose in the world coordinate system. The demonstration data is constructed by combining the global image data and the local image data.