Body-aware data acquisition method, device, medium and product

By combining an exoskeleton system with a large language-visual model, the collection and annotation of embodied intelligence data are automated, solving the problems of high cost and low efficiency. This generates a high-quality dataset that meets the training requirements of the VLA model, improving the model's generalization ability and adaptability.

CN120781874BActive Publication Date: 2025-11-11SHANGHAI COOPERS TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511241789.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-11-11
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

Existing embodied intelligence data collection is costly and inefficient, data processing lacks automation and standardization, and data scenarios and diversity are relatively limited, making it difficult to support the training needs of generalized VLA models.

Method used

By collecting human demonstration data through an exoskeleton system and combining it with a large language-visual model for data processing, a high-quality embodied intelligence dataset is generated, including automated annotation and segmentation of video streams, motion data, and force data, generating task annotations that conform to natural language descriptions.

Benefits of technology

It significantly reduces data collection costs, improves collection efficiency, generates large-scale, high-quality, and diverse datasets, meets the training needs of VLA models, and enhances the model's generalization ability and cross-embodied adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120781874B_ABST
    Figure CN120781874B_ABST
Patent Text Reader

Abstract

This application relates to the field of information technology and discloses a method, device, medium, and product for acquiring embodied intelligence data. The method includes: collecting artificial demonstration data through an exoskeleton system, the artificial demonstration data including video stream data, motion data, and force data; importing the video stream data into a first large language-visual model, segmenting it into multiple independent task segments, and generating descriptions for the independent task segments; importing the descriptions of the independent task segments into the large language model to generate task annotations based on natural language descriptions; and exporting the independent task segments, the task annotations, and the corresponding motion data and force data to generate embodied intelligence data. This establishes a low-cost, high-efficiency, and scalable method for acquiring and processing embodied intelligence data, achieving automated conversion from human operation data to high-quality VLA training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information technology, and in particular to a method, device, medium and product for acquiring embodied intelligent data. Background Technology

[0002] In recent years, embodied AI, as a key vehicle for bringing artificial intelligence into the real world, is becoming an important path to achieving artificial general intelligence (AGI). Embodied AI essentially integrates cognitive intelligence with physical execution systems, enabling machines to complete complex tasks through perception, understanding, and coordinated action. With the rapid development of multimodal large models (MLMs) and vision-language-action (VLA) models, the perception, interaction, and reasoning capabilities of embodied agents have been enhanced like never before.

[0003] However, the research and application of embodied intelligence still face significant challenges in data acquisition. High-quality training data is fundamental to building powerful embodied intelligence systems, but current data acquisition methods have significant limitations:

[0004] 1. Traditional teleoperation data acquisition is costly and inefficient.

[0005] While existing robotic teleoperation systems can provide high-fidelity datasets, they rely on specialized robotic equipment, require highly trained operators, and their high cost severely limits the scalability of data acquisition.

[0006] 2. Data processing workflows lack automation and standardization.

[0007] Current data processing relies heavily on manual operations, including data cleaning, spatiotemporal alignment, segmentation, and labeling. This approach is not only inefficient but also prone to inconsistencies. Data scarcity remains a persistent challenge in embodied intelligence research, and collecting real-world robot data faces numerous technical and cost hurdles.

[0008] 3. Existing data scenarios and diversity are relatively limited.

[0009] VLA models require large-scale, diverse pairings of visual-language-action data for training, including precise spatiotemporally synchronized data, structured annotations, and metadata. Existing datasets, limited by acquisition equipment and facilities, often lack sufficient diversity and scale to support the training needs of generalized VLA models.

[0010] Therefore, how to establish a low-cost, high-efficiency, and scalable method for acquiring and processing embodied intelligence data, and to achieve the automated conversion from human-operated data to high-quality VLA training data, has become a key technical problem that urgently needs to be solved in the field of embodied intelligence. Summary of the Invention

[0011] One objective of this application is to provide a method, device, medium, and product for acquiring embodied intelligence data, at least to address the problems of high cost and low efficiency in acquiring embodied datasets.

[0012] To achieve the above objectives, some embodiments of this application provide the following aspects:

[0013] This application provides a method for acquiring embodied intelligence data, the method comprising:

[0014] Artificial demonstration data is collected through an exoskeleton system, including video stream data, motion data, and force data.

[0015] The video stream data is imported into the first language-visual model, segmented into multiple independent task segments, and descriptions of the independent task segments are generated.

[0016] The descriptions of the independent task fragments are imported into a large language model to generate task annotations based on natural language descriptions.

[0017] The independent task segments, task labels, and corresponding motion and force data are exported to generate embodied intelligence data.

[0018] Secondly, some embodiments of this application also provide an electronic device, the electronic device comprising: one or more processors; and a memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method described above.

[0019] Thirdly, some embodiments of this application also provide a computer-readable medium having computer program instructions stored thereon, which can be executed by a processor to implement the method described above.

[0020] Fourthly, some embodiments of this application also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the method described above.

[0021] Compared with related technologies, the solution provided in this application significantly reduces data acquisition costs and improves efficiency by collecting human operational data through an exoskeleton device. Then, by combining the collected data with a large language-visual model, the data processing flow can be automated to generate a high-quality dataset that can be directly used for VLA model training. This establishes a low-cost, high-efficiency, and scalable method for acquiring and processing embodied intelligence data, achieving automated conversion from human operational data to high-quality VLA training data. Attached Figure Description

[0022] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0023] Figure 1 A flowchart illustrating an embodied intelligence data acquisition method provided as an exemplary embodiment of this disclosure;

[0024] Figure 2 A flowchart of another method for acquiring embodied intelligence data provided as an exemplary embodiment of this disclosure;

[0025] Figure 3 An exemplary structural diagram of the electronic device provided for some embodiments of this application. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] Figure 1 An embodied intelligence data acquisition method is provided as an exemplary embodiment of this disclosure, the embodied intelligence data acquisition method comprising:

[0028] S101. Collect artificial demonstration data through the exoskeleton system. The artificial demonstration data includes video stream data, motion data, and force data.

[0029] Specifically, a wearable exoskeleton system is used to collect various human demonstration data, such as picking up a cup and tidying a table. The exoskeleton system can simultaneously collect the following data streams: visual data (including RGB-D image sequences, including third-person and hand perspectives), motion data (joint angles, end effector pose, velocity, and acceleration information), and force data (contact force, torque, and grasping force information). The exoskeleton can utilize the commercially available, low-cost AirExo exoskeleton system. This exoskeleton is equipped with high-precision force sensors and cameras, capable of recording the operator's hand movements, posture, and visual motion information during the operation in real time. By utilizing a commercially available, low-cost exoskeleton for human teaching data collection, the high hardware and operational costs required by traditional robot teaching methods are significantly reduced, decreasing costs by more than 80% compared to traditional robot teleoperation systems, while simultaneously increasing data collection efficiency by 3-5 times.

[0030] S102. Import the video stream data into the first language-visual model, divide it into multiple independent task segments, and generate descriptions of the independent task segments.

[0031] Specifically, this step leverages the powerful world knowledge and contextual understanding capabilities of a large language-visual model (VLM) to identify "meaningful" task transition points from a human perspective as boundaries in the data stream. The input is a time-aligned multimodal data stream, primarily video streams. The video stream is segmented based on content semantic segmentation or action unit detection methods. Then, a first large language-visual model (such as Video-XL-2) performs video understanding on each segment and generates scene description text. The time scale of each video data segment within the original video is then calculated, ultimately resulting in independent task segments and their descriptions, such as "0.0s, 2.5s, the robot hand is reaching for the water cup," "2.5s, 4.5s, the robot hand grasps the water cup," and "4.5s, 10.0s, the robot is pouring water."

[0032] S103. Import the independent task fragment description into the large language model to generate task annotations based on natural language description.

[0033] Specifically, this step involves using LLM to refine primitive labels to obtain more natural subtask instructions. Finally, this series of subtask instructions is input into the LLM as contextual information for intent understanding and annotation. Leveraging the powerful common-sense reasoning capabilities of LLM, independent task fragments are transformed into natural language task instructions that better align with human expression habits. For example, the primitive label "the robot's hand grasped the water glass" is refined to "picking up the glass on the table." This results in a series of coherent natural language subtask instructions, which are then analyzed by LLM and subjected to high-level induction and reasoning to identify the core user intent or the "skills" the robot needs to learn from the entire teaching data.

[0034] S104. Export the independent task segments, the task labels, and the corresponding motion data and force data to generate embodied intelligence data.

[0035] Specifically, the individual task segments, task annotations, and corresponding motion and force data from each segment are packaged together and output as embodied intelligence data, which will then form an embodied intelligence dataset for subsequent development and simulation.

[0036] In this embodiment, human operational data is collected through an exoskeleton device, significantly reducing data acquisition costs and improving efficiency. The collected data is then combined with a large language-visual model, automating the data processing flow and generating a high-quality dataset directly usable for VLA model training. This establishes a low-cost, high-efficiency, and scalable method for acquiring and processing embodied intelligence data, achieving automated conversion from human operational data to high-quality VLA training data. The large-scale, high-quality, diverse, and semantically rich accompanying data generated by this method directly meets the core requirements of VLA models for training data, effectively overcoming the challenge of existing VLA models' "lack of large-scale datasets." By utilizing human operational data, VLA models can learn a wider range of skills and stronger generalization capabilities, achieving "few-shot generalization" and "cross-embodied adaptation," thus performing better in unseen tasks and novel environments. This high-quality data is the cornerstone for VLA models to realize their general intelligence potential.

[0037] In one embodiment, such as Figure 2 As shown, the step of importing the video stream data into the first language-visual model, dividing it into multiple independent task segments, and generating descriptions of the independent task segments specifically includes:

[0038] S201. The video stream data is compressed into multiple video segments using a frame extraction algorithm.

[0039] Specifically, in this embodiment, a frame extraction algorithm is used to compress continuous video frames into different small segments. The core principle is to selectively retain or discard video frames, thereby reducing the amount of data while preserving as much key information as possible (such as actions, scene changes, etc.), providing concise and meaningful video segments for subsequent processing (such as VLM analysis).

[0040] Frames can be extracted evenly from consecutive frames based on preset time intervals (e.g., one frame every 0.25 seconds or 0.5 seconds) or frame rates (e.g., extracting 10 frames from a 30fps video). For example, a 10-second, 30fps video (300 frames in total) will ultimately retain 10 frames if frames are extracted at 1-second intervals (one frame is taken from every 30 frames).

[0041] Dynamic threshold sampling can also be used to calculate pixel differences (such as changes in grayscale or RGB values) or feature differences (such as changes in edges, textures, or object contours) between two adjacent frames, and quantify them numerically (e.g., the larger the difference, the more drastic the scene change). A threshold is set; when the difference between two frames exceeds the threshold, the current frame is retained (considered a "critical change point"); if the difference is less than the threshold, the current frame is discarded (considered content duplication).

[0042] Alternatively, a key target tracking approach can be used. First, a target detection algorithm is used to identify key targets in the video (such as a "robot hand"), then its motion state (position and posture changes) is tracked. When the target's motion state changes significantly (e.g., from "stationary" to "moving," or from "reaching towards the water cup" to "grasping the water cup"), a frame is extracted and the frame at that moment is preserved. For example, in a video of a robot pouring water, when "the hand's position moves from above the water cup to the water flow," the frame is extracted and preserved as the turning point of the action.

[0043] S202, Generate video clip description text using the first large language-visual model.

[0044] Specifically, the primary language—visual models (such as Video-XL-2)—is used to perform video understanding on each video segment and generate scene description text.

[0045] S203. Calculate the time scale of the video segment in the video stream data based on the frame extraction frequency in the frame extraction algorithm.

[0046] Specifically, the core of this step is to establish a mapping between video segments and the timeline of video stream data by establishing the correspondence between the frame extraction frequency and the original video frame rate. Essentially, it is to reverse-engineer the time information lost during the frame extraction process. The video segment obtained after frame extraction consists of several keyframes, and each keyframe has a unique "frame number" in the original video. The timestamp in the original video can be deduced from the frame number.

[0047] S204. The video stream data is segmented into independent task segments based on the video segment description text and the time scale, and the description of the independent task segment is generated.

[0048] Specifically, based on the video clip description text and time scale, the final result is an independent task clip and an independent task clip description, such as "0.0s, 2.5s, the robot hand is reaching for the water cup", "2.5s, 4.5s, the robot hand grasps the water cup", and "4.5s, 10.0s, the robot is pouring water".

[0049] In this embodiment, frame extraction reduces processing costs, while time conversion preserves key temporal information, multimodal associations, and semantic rationality. Ultimately, the automatically segmented subtask fragments truly align with "human cognitive logic of task stages," providing a reliable foundation for subsequent processing.

[0050] In one embodiment, the step of importing the video stream data into a first language-visual model, segmenting it into multiple independent task segments, and generating descriptions of the independent task segments further includes:

[0051] S205. Extract motion features from motion data and record time points.

[0052] S206. Search for the motion features in the independent task segment, and adjust the time scale of the independent task segment based on the time points of the searched motion features.

[0053] Specifically, the time scale of video stream data obtained through video semantic understanding may not be precise enough. This step achieves more accurate segmentation by fusing high-level semantic boundaries and low-level temporal features. First, temporal data such as the robot arm's end effector speed is processed (e.g., sensor temporal data of end effector speed, acceleration, joint angle, and gripping force (aligned with the video stream time)). Relevant motion features are extracted and time points are recorded (e.g., determining the start and stop of motion using speed thresholds). For example, kinematic thresholds are set (e.g., speed > 0.05 m / s is considered "moving", speed = 0 is considered "stationary"), and key time points such as "motion start", "motion stop", and "action switch" are identified through threshold intersections. Then, a small time window (e.g., ±0.2s, i.e., 2.3s-2.7s) is set around the time scale boundary of an independent task segment (e.g., 2.5s) to focus on fine features near the boundary. The searched motion feature time point (e.g., 2.48s) replaces the original coarse-grained boundary (e.g., 2.5s) of the VLM output, forming a refined segment and achieving more accurate segmentation. This ultimately yields more refined sub-task segments, such as "0.0s, 2.48s, the robot hand is reaching for the water glass," "2.48s, 4.62s, the robot hand grasps the water glass," and "4.62s, 10.0s, the robot is pouring water." The original video stream is segmented based on these time scales to obtain the final time scales for independent task segments, thus achieving independent task segments with time accuracy down to the millisecond level.

[0054] In this embodiment, through the collaborative logic of "high-level semantics determining direction and low-level features refining details", the semantic understanding ability of VLM for the task stage is preserved (ensuring that the segmentation is "meaningful"), while the physical accuracy of kinematic data is utilized (ensuring that the segmentation is "precise enough"). The resulting sub-task segments can be intuitively understood by humans and meet the technical requirements of robot systems for temporal accuracy, laying a reliable foundation for automated and intelligent task analysis.

[0055] In one embodiment, the step of segmenting the video stream data into the independent task segments based on the video segment description text and the time scale, and generating the independent task segment description, specifically includes:

[0056] S207. Input the independent task fragment into the second language-visual model to generate tags for actions and items in the independent task fragment.

[0057] Specifically, primitive labels are generated for actions and objects in independent task segments using a second large language-visual model. For example, basic action types such as grasping, placing, pushing, and pulling are automatically identified and the objects involved in the operation are labeled. The second large language-visual model can be the same as the first large language-visual model, or a different large language-visual model, or a large language-visual model specifically optimized for different scenarios.

[0058] S208, and construct a structured description of the actions and items through the second major language-visual model.

[0059] Specifically, by inputting a segmented, independent task into VLM (Qwen2.5-VL), a structured prompt is given to VLM, such as: "Please describe the core action of the robotic arm in the video in a short phrase, in the format of 'skill + object'." VLM generates the corresponding primitive tag, such as "pick up + cup".

[0060] S209. The structured description is determined as the description of the independent task fragment.

[0061] Specifically, the structured descriptions will be replaced with natural language task instructions that better align with human expression habits. For example, the primitive label "pick up + cup" will be refined to "pick up the cup on the table." The operational intent will be inferred based on contextual information to generate complete operational skill annotations. This will result in a series of coherent natural language sub-task instructions, which will then be analyzed using LLM (Limited Language Management) for high-level induction and reasoning. This will reveal the core user intent or the "skills" the robot needs to learn from the entire teaching data.

[0062] In one embodiment, human-machine collaborative annotation further controls the quality of the output data. Based on a visual annotation platform, annotators review the original video clips, VLM primitives, LLM-generated subtask instructions, and intent tags, making modifications, confirmations, or rejections. Manual verification and correction of the annotation results ensure data quality and eliminate potential systematic errors. This human-machine collaborative annotation mechanism achieves automation efficiency while ensuring accuracy and consistency through manual review, further improving data quality and avoiding the illusions or biases that purely automated annotation might introduce.

[0063] In this embodiment, the second VLM generates action tags (such as "approach" and "grab") and item tags (such as "cup" and "table"), and constructs a structured description of "skill + object" (such as "take + cup"). This abstracts the visual content in the video into discrete semantic units with clear logical relationships, transforming the raw visual information of independent task segments into standardized, reusable semantic symbols, providing a unified "language" for task understanding, analysis, and reuse. Subsequently, the large language model can generate task annotations based on natural language descriptions using this unified "language."

[0064] To ensure the "high quality" of VLA model training data, the collected data needs to be temporally and spatially aligned to ensure consistency between data from different sensors in time and space. Preprocessing is required before the step of collecting artificial demonstration data via the exoskeleton system, which includes video stream data, motion data, and force data.

[0065] In one embodiment, the embodied intelligent data acquisition method includes timestamping the sensors in the exoskeleton system.

[0066] Specifically, a global time base is established to timestamp and calibrate all sensor data.

[0067] In one embodiment, the embodied intelligent data acquisition method further includes calibrating the global coordinate system of the exoskeleton system using spatial markers.

[0068] Specifically, ArUco markers are placed in space, and the global coordinate system is calibrated using a relative state-action representation method.

[0069] In one embodiment, the embodied intelligent data acquisition method further includes eliminating sensor latency in the exoskeleton system through a time consistency algorithm.

[0070] Specifically, using a time consistency algorithm, the left shoulder camera is selected as the primary sensor, and its timestamp sequence is used as a reference. Each time, data from other sensors that is closest to the reference time is selected. The found multi-sensor data is combined into a single frame and marked as the aligned data output. This eliminates the impact of delays from multiple sensors in the exoskeleton system.

[0071] In the above embodiments, the precise temporal-space alignment technology ensures a high degree of consistency between human operation data and robot operation space, effectively solving the problems of modal mismatch and timing differences.

[0072] To ensure the purity and reliability of the artificial demonstration data, and to guarantee that the VLA model can learn from clean and accurate data, the embodied intelligence data acquisition method further includes data cleaning and noise reduction processing of the collected data after the step of collecting artificial demonstration data through the exoskeleton system, which includes video stream data, motion data, and force data.

[0073] In one embodiment, the embodied intelligence data acquisition method further includes identifying and removing anomalous data from the artificial demonstration data.

[0074] Specifically, abnormal data points caused by sensor malfunctions or operational errors are identified and removed. This is done by checking whether sensor readings fall within a reasonable range. For example, when reading camera data, abnormalities are indicated by images that are entirely black or white, or display a green screen. Similarly, when reading joint data, abnormal values ​​are identified by joint rotation angles or speeds that exceed or fall below the reasonable range achievable by the motor.

[0075] In one embodiment, the embodied intelligence data acquisition method further includes denoising the artificial demonstration data.

[0076] Specifically, Kalman filtering algorithms are applied to remove sensor noise from noisy multi-sensor measurements (such as IMU acceleration).

[0077] In one embodiment, the embodied intelligence data acquisition method further includes performing an integrity check on the artificial demonstration data.

[0078] Specifically, a rule-based script is used to detect whether the data stream is arranged continuously in ascending order of timestamps, to reorder out-of-order frames, to perform linear interpolation compensation for detected missing frames or to directly discard the synchronization frame group at that moment, and to handle problems such as missing data and disordered timing.

[0079] In one embodiment, the embodied intelligence data acquisition method further includes manual sampling quality inspection and secondary cleaning.

[0080] Specifically, samples are taken from the processed data for manual quality inspection, and data that does not meet the quality requirements are manually cleaned a second time.

[0081] In the above embodiments, the comprehensive data cleaning and preprocessing process ensures the purity and reliability of the training data, ensuring that the VLA model can learn from clean and accurate data.

[0082] In one embodiment, the embodied intelligence data acquisition method further includes:

[0083] The exported embodied intelligence data conforms to the LeRobot standard format specification.

[0084] Specifically, through the aforementioned data collection, preprocessing, automatic segmentation, and labeling processes, this invention ultimately generates structured, cleaned, and semantically rich multimodal data. To ensure that this data meets the training requirements of current mainstream VLA models and possesses good scalability and compatibility, the data undergoes further data transformation and adaptation processes to construct a dataset structure conforming to the LeRobot standard format (the standard LeRobot dataset format, containing visual data, state information, and action labels), as follows:

[0085] <dataset_root> /

[0086] ├── data /

[0087] │ ├── chunk-000 /

[0088] │ │ ├── episode_000000.parquet

[0089] │ │ ├── episode_000001.parquet

[0090] │ │ └── ...

[0091] │ └── chunk-001 /

[0092] ├── meta /

[0093] │ ├── episodes.jsonl

[0094] │ ├── info.json

[0095] │ ├── stats.safetensors

[0096] │ └── tasks.jsonl

[0097] └── videos /

[0098] ├── chunk-000 /

[0099] │ ├── observation.images.palm_camera /

[0100] │ │ ├── episode_000000.mp4

[0101] │ │ └── ...

[0102] │ └── observation.images.global_camera /

[0103] └── chunk-001 /

[0104] In this embodiment, in order to ensure that these data meet the training requirements of current mainstream VLA models and have good scalability and compatibility, the above data undergoes a data conversion and adaptation process to be converted into embodied intelligence data that conforms to the LeRobot standard format, thus clearing obstacles for subsequent VLA model training and embodied intelligence system applications.

[0105] Furthermore, some embodiments of this application also provide an electronic device. The electronic device can be various forms of digital computer, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device can also be various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices.

[0106] The electronic device includes: one or more processors; and a memory storing computer program instructions that, when executed, cause the processor to perform the steps of the methods provided in any one or more of the above embodiments. Figure 3 An exemplary structural diagram of the electronic device is disclosed. The electronic device includes one or more processors 1101, a memory 1102, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0107] The electronic device may further include an input device 1103 and an output device 1104. The processor 1101, memory 1102, input device 1103 and output device 1104 may be connected by a bus or other means, as shown in the figure, which is connected by a bus.

[0108] Input device 1103 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 1104 may include a display device, auxiliary lighting device (e.g., LED), and haptic feedback device (e.g., vibration motor). The display device may include, but is not limited to, a liquid crystal display, a light-emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0109] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device (e.g., a cathode ray tube or LCD monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback); and input from the user can be received in any form (e.g., voice input or tactile input).

[0110] In this embodiment, a computer-readable medium stores a computer program / instructions that, when executed by a processor, implement the steps of the methods provided in any one or more of the above embodiments. This computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into that device. The aforementioned computer-readable medium carries one or more computer-readable instructions.

[0111] The memory 1102 can serve as a non-transitory computer-readable storage medium, used to store non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, thereby implementing the program instructions / modules corresponding to the methods provided in any one or more of the embodiments described above in this application.

[0112] The memory 1102 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 1102 may optionally include memory remotely located relative to the processor 1101, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0113] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0114] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technologies, read-only optical discs, digital versatile optical discs or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0115] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0116] In the above embodiments, all or part of the implementation can be achieved through software, hardware, firmware, or any combination thereof. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drives, floppy disks, and similar devices. In addition, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.

[0117] The computer program product provided in this application includes one or more computer programs / instructions. When executed by a processor, these computer programs / instructions generate, in whole or in part, the processes or functions described in this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0118] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0119] The scope of this application is defined by the appended claims rather than the foregoing description, and is therefore intended to encompass all variations falling within the meaning and scope of equivalents of the claims. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device in software or hardware. Terms such as "first," "second," etc., are used only for distinguishing descriptions and do not indicate any particular order, nor should they be construed as indicating or implying relative importance.

[0120] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily made by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.

Claims

1. A method for acquiring embodied intelligence data, characterized in that, The method for acquiring embodied intelligence data includes: Artificial demonstration data is collected through an exoskeleton system, including video stream data, motion data, and force data. The video stream data is divided into multiple independent task segments, and the first major language-visual model is imported to generate descriptions of the independent task segments; The descriptions of the independent task fragments are imported into a large language model to generate task annotations based on natural language descriptions. The independent task segments, task labels, and corresponding motion and force data are exported to generate embodied intelligence data. The step of segmenting the video stream data into multiple independent task segments and importing them into a first language-visual model to generate descriptions of the independent task segments specifically includes: The video stream data is compressed into multiple video segments using a frame-skipping algorithm; The video clip is used to generate a descriptive text for the video clip through the first large language-visual model; The time scale of the video segment in the video stream data is calculated based on the frame extraction frequency in the frame extraction algorithm. The video stream data is segmented into independent task segments based on the video segment description text and the time scale, and descriptions of the independent task segments are generated.

2. The embodied intelligence data acquisition method according to claim 1, characterized in that, The step of dividing the video stream data into multiple independent task segments and importing them into a first language-visual model to generate descriptions of the independent task segments specifically includes: Motion features are extracted from the motion data and time points are recorded. The motion features are searched within the independent task segments, and the time scale of the independent task segments is adjusted based on the time points of the searched motion features.

3. The embodied intelligence data acquisition method according to claim 1, characterized in that, The step of segmenting the video stream data into independent task segments based on the video segment description text and the time scale, and generating descriptions for the independent task segments, specifically includes: The independent task segments are input into the second large language-visual model to generate tags for actions and items in the independent task segments; And a structured description of the actions and items is constructed using the second major language-visual model; The structured description is determined as the description of the independent task fragment.

4. The embodied intelligence data acquisition method according to claim 1, characterized in that, Prior to the step of collecting artificial demonstration data via the exoskeleton system, wherein the artificial demonstration data includes video stream data, motion data, and force data, the embodied intelligence data acquisition method further includes: The sensors in the exoskeleton system are timestamped and calibrated. and / or; The global coordinate system of the exoskeleton system is calibrated using spatial markers; and / or; Sensor latency in the exoskeleton system is eliminated using a time consistency algorithm.

5. The embodied intelligence data acquisition method according to claim 1, characterized in that, Following the step of collecting artificial demonstration data via an exoskeleton system, wherein the artificial demonstration data includes video stream data, motion data, and force data, the embodied intelligence data acquisition method further includes: Identify and remove anomalous data from the artificially generated sample data; and / or; The artificial demonstration data is then denoised. and / or; The integrity of the artificial demonstration data is checked.

6. The embodied intelligence data acquisition method according to claim 1, characterized in that, The method for acquiring embodied intelligence data also includes: The exported embodied intelligence data conforms to the LeRobot standard format specification.

7. An electronic device, characterized in that, The electronic device includes: One or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method as described in any one of claims 1 to 6.

8. A computer-readable medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 6.

9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Methods and systems for determining object activity within a region of interest

    US20190318171A1

  • Large language model-based event processing method and apparatus, device and medium

    WO2025086682A1