Sample data acquisition method, data processing system and storage medium

By collecting and labeling production-accompanying video data in a real production environment, sample data is generated to train the embodied robot, which solves the problem of decreased adaptability of the embodied robot during actual deployment and improves its execution capability in the production environment.

CN121305262APending Publication Date: 2026-01-09ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511518736.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

In existing technologies, the sample data used to train embodied intelligent models is collected in experimental sites or simulation environments, which leads to a decrease in the adaptability and performance degradation of embodied robots when they are actually deployed.

Method used

The operator wears a wearable video capture device to collect accompanying video data from a first-person perspective. The data is then labeled using a data processing system to generate sample data for training the embodied robot. The data collection takes place in a real production environment and includes visual information about the production environment and the task execution process.

Benefits of technology

This improves the matching degree and adaptability of the embodied robot in the actual production environment, and enhances its execution capability and overall performance in actual production tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121305262A_ABST
    Figure CN121305262A_ABST
Patent Text Reader

Abstract

The invention provides a sample data acquisition method, a data processing system and a storage medium. The data processing system can firstly obtain target video data to be labeled, and the target video data to be labeled is production adjoint type video data collected by an operator through carrying wearable video collection equipment at a first person perspective. Wherein the production syndrome video data refers to data generated in the process of executing the actual production task by the operator. Then, the data processing system can label the target video data to obtain sample data; wherein the sample data is used for training the body-equipped robot to execute the actual production task executed by the operator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and in particular to a method for acquiring sample data, a data processing system, and a storage medium. Background Technology

[0002] Embodied intelligence technology combines artificial intelligence with physical entities (such as robots and smart devices), enabling these entities to perceive their environment, understand tasks, and perform tasks autonomously using embodied intelligence models. Training embodied intelligence models is a crucial step in embodied intelligence technology.

[0003] Currently, the sample data used to train embodied intelligence models are usually operational data collected in pre-built dedicated experimental sites or operational data simulated in constructed simulation environments.

[0004] However, when embodied intelligence models trained using these sample data are deployed in practice, they may experience a decline in adaptability.

[0005] The background information is merely information known only to the inventor and does not imply that such information had entered the public domain before the date of this application, nor does it imply that it could be considered prior art in this disclosure. Summary of the Invention

[0006] This specification provides a sample data acquisition method, a data processing system, and a storage medium, which can annotate production-related video data to obtain sample data for training androids so that the androids can perform corresponding actual production tasks.

[0007] To achieve the above objectives, the embodiments in this specification adopt the following technical solutions: Firstly, this specification provides a method for acquiring sample data, comprising: acquiring target video data to be labeled, wherein the target video data to be labeled is production-accompanying video data acquired by an operator from a first-person perspective by wearing a wearable video acquisition device, wherein the production-accompanying video data refers to data generated by the operator during the execution of actual production tasks; and labeling the target video data to obtain sample data, wherein the sample data is used to train a holographic robot to perform the actual production tasks performed by the operator.

[0008] In some embodiments, the step of annotating the target video data to obtain sample data includes: processing the target video data to obtain multiple video segments and sub-tasks corresponding to the multiple video segments respectively; performing annotation tasks for at least one dimension based on the multiple video segments and the sub-tasks corresponding to the multiple video segments respectively to obtain annotated data corresponding to the multiple video segments respectively, wherein the annotated data includes: reference data corresponding to the at least one dimension, and the sample data obtained based on the target video data, the sub-tasks corresponding to the multiple video segments respectively, and the annotated data corresponding to the multiple video segments respectively.

[0009] In some embodiments, the annotation task of at least one dimension includes an annotation task of the task planning dimension, and the reference data corresponding to the task planning dimension includes: task status information; the step of performing the annotation task of at least one dimension based on the plurality of video segments and the sub-tasks corresponding to the plurality of video segments includes: generating a first prompt instruction based on the plurality of video segments and the sub-tasks corresponding to the plurality of video segments, the first prompt instruction being used to guide the visual big language model to perform content recognition on the plurality of video segments, and to perform inference of the task planning dimension based on the content recognition results and sub-tasks corresponding to the plurality of video segments; and inputting the first prompt instruction into the visual big language model to obtain the task status information corresponding to the plurality of video segments output by the visual big language model.

[0010] In some embodiments, the task status information includes at least one task status question-answer pair, which includes question text and answer text. Accordingly, the first prompt instruction is also used to guide the visual big language model to obtain multiple question texts from a preset question library, and to perform task planning dimension reasoning based on the content recognition results and sub-tasks corresponding to each video segment, thereby generating the answer text corresponding to the question text.

[0011] In some embodiments, the annotation task of at least one dimension includes an annotation task of the operable region dimension, and the reference data corresponding to the operable region dimension includes: operable region; the step of performing the annotation task of at least one dimension based on the plurality of video segments and the sub-tasks corresponding to the plurality of video segments includes: using a target recognition model to perform target recognition on the plurality of video segments to obtain candidate objects existing in the plurality of video segments and the annotation boxes of the candidate objects; and based on the sub-tasks corresponding to the plurality of video segments, determining the target object corresponding to the sub-task from the candidate objects, and determining the operable region according to the annotation boxes of the target objects.

[0012] In some embodiments, the operable area includes an area corresponding to an operable item and an area corresponding to the operation position of the operable item. The target object includes a target item and a target operating part, and the target operating part includes a component or hand for operation. Determining the operable area based on the annotation frame of the target object includes at least one of the following steps: if the target object corresponding to the subtask includes the target item, determine that the area marked by the annotation frame of the target item is the area corresponding to the operable item; if the target object corresponding to the subtask includes the target item and the target operating part, determine that the overlapping area of ​​the annotation frame of the target operating part and the annotation frame of the target item is the area corresponding to the operation position of the operable item.

[0013] In some embodiments, the annotation task of at least one dimension includes a navigation trajectory dimension annotation task, and the reference data corresponding to the navigation trajectory dimension includes: navigation guidance text; the step of performing the annotation task of at least one dimension based on the plurality of video segments and the sub-tasks corresponding to the plurality of video segments includes: determining at least one navigation video segment involving navigation among the plurality of video segments based on the sub-tasks corresponding to the plurality of video segments; and generating navigation guidance text corresponding to the at least one navigation video segment using a visual large language model.

[0014] In some embodiments, generating navigation guidance text corresponding to the at least one navigation video segment using a visual big language model includes: determining multiple target images in the navigation video segment, wherein the target images are images in the navigation video segment that meet preset conditions; obtaining the spatial coordinates corresponding to target objects in the multiple target images using a spatial reconstruction model; generating a second prompt instruction based on the multiple target images and the spatial coordinates corresponding to target objects in the multiple target images, wherein the second prompt instruction is used to guide the visual big language model to perform content recognition on the multiple target images and to infer the navigation trajectory dimension based on the spatial coordinates corresponding to target objects in the multiple target images; and inputting the second prompt instruction into the visual big language model to obtain the navigation guidance text output by the visual big language model.

[0015] In some embodiments, the preset condition includes: the entropy of the probability distribution of the predicted navigation action corresponding to the image is greater than a preset threshold; determining multiple target images in the navigation video segment includes: obtaining the entropy corresponding to each frame of the image in the navigation video segment using a pre-trained entropy prediction model, wherein the entropy prediction model has the ability to determine the entropy of the probability distribution of the predicted navigation action based on the image; and determining the image among the multiple images whose entropy is greater than the preset threshold as the target image.

[0016] In some embodiments, the navigation guidance text includes navigation direction, navigation distance, and navigation angle, wherein the navigation direction is obtained by the visual big language model based on the navigation trajectory dimension inference of the multiple target images, and the navigation distance and navigation angle are obtained by the visual big language model based on the spatial coordinates corresponding to the target objects in the multiple target images respectively, based on the navigation trajectory dimension inference of the navigation trajectory.

[0017] In some embodiments, the annotation task of at least one dimension includes an annotation task of the operation trajectory dimension, and the reference data corresponding to the operation trajectory dimension includes the movement trajectory of the target operation unit; the step of performing the annotation task of at least one dimension based on the plurality of video segments and the sub-tasks corresponding to the plurality of video segments includes: determining at least one operation video segment involving operation among the plurality of video segments based on the sub-tasks corresponding to the plurality of video segments; and obtaining a target operation unit coordinate sequence in which the spatial coordinates of the target operation unit change over time in the at least one operation video segment, and generating the movement trajectory of the target operation unit based on the target operation unit coordinate sequence.

[0018] In some embodiments, the target operating part includes a hand or a component for operation. The step of obtaining a target operating part coordinate sequence in the at least one operation video segment, where the spatial coordinates of the target operating part change over time, includes: identifying the target operating part in the operation video segment using a target recognition model, and at least one key point corresponding to the target operating part; obtaining the spatial coordinates of the key point in each frame of the operation video segment using a spatial reconstruction model; and generating a target operating part coordinate sequence corresponding to the operation video segment based on the spatial coordinates of the key point in each frame.

[0019] In some embodiments, the target operating unit includes a hand or a component for operation, and the target operating unit includes at least one positioning module; obtaining the target operating unit coordinate sequence in which the spatial coordinates of the target operating unit change over time in the at least one operating video segment includes: obtaining positioning data corresponding to the operating video segment, the positioning data being collected using the positioning module in the target operating unit when collecting target video data; and generating the target operating unit coordinate sequence corresponding to the target operating unit based on the operating video segment and the positioning data.

[0020] In some embodiments, before annotating the target video data to obtain sample data, the method further includes: performing at least one of the following preprocessing steps on the target video data to obtain preprocessed target video data: replacing at least one candidate object or target operation unit in the target video data; performing derivation generation processing on the target video data to generate a derived video of the target video data, wherein the derived video has a different production task or a different execution result from the production task of the target video data; or editing a portion of the target video data.

[0021] In some embodiments, the target operation part in the target video data is a hand, and replacing the target operation part in the target video data includes: acquiring an image of a component for operation as a replacement object; and using a video editing big model to replace the replacement object at the position of the hand in the target video data.

[0022] In some embodiments, the step of performing derivative generation processing on the target video data to generate a derivative video of the target video data includes: using a large language model to generate content-generating descriptive text related to the preset derivative content as the target, the preset derivative content including: content that is opposite to the execution result of the production task in the target video data, and content that is different from the production task in the target video data; and using a text-generated video model to generate the derivative video based on the target video data and the content-generated descriptive text.

[0023] In some embodiments, after obtaining the annotation data corresponding to the plurality of video segments respectively, the method further includes: performing an accuracy check on the obtained annotation data corresponding to the plurality of video segments respectively, obtaining an accuracy check result, the accuracy check result being used to characterize whether the annotation data is accurate; and correcting the annotation data when the accuracy check result indicates that the annotation data is inaccurate.

[0024] In some embodiments, the android includes a motion model and a navigation model, wherein the sample data is used to train the motion model and navigation model of the android to enable the android to perform the actual production tasks performed by the operator.

[0025] Secondly, this specification also provides a data processing system, comprising: at least one storage medium storing at least one instruction set for acquiring sample data; and at least one processor communicatively connected to the at least one storage medium, wherein, when the data processing system is running, the at least one processor reads the at least one instruction set and implements the method provided in the first aspect according to the instructions of the at least one instruction set.

[0026] Thirdly, this specification also provides a computer-readable non-volatile storage medium, wherein the computer-readable non-volatile storage medium stores at least one instruction set, which, when executed by at least one processor, implements the method provided in the first aspect.

[0027] Other functions of the sample data acquisition method, data processing system, and storage medium provided in this specification will be partially listed in the following description. The inventive aspects of the sample data acquisition method, data processing system, and storage medium provided in this specification can be fully explained through practice or use of the methods, apparatus, and combinations described in the detailed examples below. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 A schematic diagram illustrating an application scenario of the sample data acquisition method provided according to embodiments of this specification is shown. Figure 2 A hardware structure diagram of a computing system provided according to an embodiment of this specification is shown; Figure 3 A flowchart of a sample data acquisition method provided according to an embodiment of this specification is shown; Figure 4 A schematic diagram illustrating a method for preprocessing target video data according to an embodiment of this specification is shown. Figure 5 A flowchart illustrating an annotation of target video data according to an embodiment of this specification is shown; Figure 6 A schematic diagram illustrating a method for splitting target video data according to an embodiment of this specification is shown; Figure 7 A schematic diagram of a task annotation task in the task planning dimension provided according to an embodiment of this specification is shown; Figure 8 A schematic diagram of an operational region dimension annotation task is shown according to an embodiment of this specification; Figure 9 A schematic diagram of a navigation trajectory dimension annotation task provided according to an embodiment of this specification is shown; and Figure 10 A schematic diagram of an operation trajectory dimension annotation task provided according to an embodiment of this specification is shown. Detailed Implementation

[0030] The following description provides specific application scenarios and requirements for this specification, intended to enable those skilled in the art to make and use the contents of this specification. Various partial modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the embodiments shown, but rather to the widest scope consistent with the claims.

[0031] The terminology used herein is for the purpose of describing particular exemplary embodiments only and is not restrictive. For example, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” used herein may also include the plural forms. When used in this specification, the terms “comprising,” “including,” and / or “containing” mean that the associated integers, steps, operations, elements, and / or components are present, but do not exclude the presence of one or more other features, integers, steps, operations, elements, components, and / or groups, or that other features, integers, steps, operations, elements, components, and / or groups may be added to the system / method.

[0032] Considering the following description, these and other features of this specification, as well as the operation and function of the related components of the structure, and the economy of assembly and manufacture of the parts, can be significantly improved. All of these form part of this specification with reference to the accompanying drawings. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.

[0033] The flowcharts used in this specification illustrate operations implemented according to some embodiments of this specification. It should be clearly understood that the operations in the flowcharts may not be implemented in a sequential order. Instead, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more additional operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.

[0034] For ease of description, the terms that will appear later in this manual will be explained first.

[0035] Term 1: Embodied Robot. This refers to a robot that utilizes embodied intelligence technology, in which an embodied intelligence model is deployed. This model can autonomously control the robot to perform actions such as target recognition, manipulation, and movement to execute corresponding production tasks. Training an embodied robot is equivalent to training the embodied intelligence model deployed within it.

[0036] Term 2: Production-Accompanying Data. Production-accompanying data acquisition is a data collection method that occurs simultaneously during actual production or task execution. In production-accompanying data acquisition, data collection and task execution occur concurrently, eliminating the need for a dedicated data collection site. In this specification, production-accompanying data refers to video data. Operators can use wearable video capture devices to collect video data from a first-person perspective while performing actual production tasks.

[0037] Term 3: Production Environment. This refers to the site where actual production tasks are carried out. A real production environment includes materials, equipment, etc., used for production. Production tasks performed in a real production environment will genuinely alter the physical state of that environment.

[0038] Term 4: Production Task. This refers to an operation or process performed in a production environment that achieves a specific goal. For example, assuming the production environment is a warehouse, a production task is to move a piece of material from shelf A to shelf B. Once the production task is completed, the material on shelf A will actually be moved to shelf B.

[0039] In this specification, the Large Language Model (LLM) may also be referred to simply as the Large Model. A Large Language Model is a natural language processing model based on deep learning techniques, typically with billions to hundreds of billions or even more parameters, possessing powerful language understanding and generation capabilities. Large Language Models can employ the Transformer architecture or its variants (such as GPT, BERT, etc.), which utilizes an attention mechanism to globally model sequential data, efficiently handling long-distance dependencies and thus performing excellently in natural language tasks. Large Language Models learn the statistical features and semantic relationships of language through pre-training on large-scale corpora, giving them excellent generalization capabilities. The core capabilities of Large Language Models include, but are not limited to: understanding contextual semantics, generating coherent and grammatically correct text, performing logical reasoning, and handling multi-task scenarios. Its usage typically includes two modes: direct inference and fine-tuning. In direct inference mode, the user guides the Large Language Model to generate specific outputs by designing prompts. Cue words can be task descriptions or instructions in text form, used to stimulate the semantic understanding and generation capabilities of large language models. In fine-tuning mode, large language models are further trained on small-scale datasets in specific domains to optimize their performance on specific tasks. The powerful generalization ability and flexibility of large language models make them an important tool in the field of artificial intelligence, providing efficient and accurate solutions for automated text generation and understanding.

[0040] In some embodiments, large language models can also understand and generate data from other modalities (such as visual and audio data). In this case, large language models can also be called multimodal large language models (MLLMs). MLLMs provide a richer and more natural interactive experience by integrating multiple types of input and output, such as text, images, and sound. The core advantage of MLLMs lies in their ability to process and understand information from different modalities and fuse this information to accomplish complex tasks. For example, the Vision Language Model (VLM) discussed in this specification is a branch of MLLMs; a VLM can analyze an image and generate descriptive text. In other examples, MLLMs can also generate corresponding images or videos based on text descriptions. This cross-modal understanding and generation capability makes MLLMs widely applicable in multiple fields.

[0041] It should be noted that the key technologies of large language models can be found in the detailed description in the paper "A Survey of Large Language Models" (paper number: arXiv:2303.18223v16, published on March 11, 2025, public link: https: / / doi.org / 10.48550 / arXiv.2303.18223), and will not be repeated here.

[0042] In related technologies, acquired video data is labeled to generate sample data, which is then used to train embodied robots to perform actual production tasks. However, the video data in these technologies is often collected during simulated production tasks in experimental settings or simulated environments. Since experimental settings or simulated environments differ significantly from real production scenarios, using labeled sample data from such settings to train embodied robots can lead to decreased adaptability and performance degradation in actual production environments.

[0043] Therefore, this specification provides a method for acquiring sample data, which can be executed by a data processing system. The method includes: the data processing system first acquires target video data to be labeled. This target video data is production-related video data collected by an operator from a first-person perspective using a wearable video acquisition device. The production-related video data refers to data generated by the operator during the execution of actual production tasks. Then, the data processing system labels the target video data to obtain sample data. This sample data is used to train the android to perform the actual production tasks performed by the operator.

[0044] In the solution provided in this manual, the target video data to be labeled is production-related video data acquired from a first-person perspective using a wearable video acquisition device worn by the operator. The data processing system can label the target video data to obtain sample data. The production-related video data acquired during actual production contains visual information of the real production environment and the complete production task execution process. By labeling the production-related video data and obtaining sample data, the sample data can characterize the process of executing production tasks in the actual production environment. Using such sample data to train the embodied robot can improve the matching degree and adaptability of the embodied robot to the actual production environment, making the behavior of the embodied robot more closely match the actual production environment. Furthermore, this can improve the execution capability and overall performance of the embodied robot in actual production tasks.

[0045] In this specification, video data annotation refers to the process of identifying, classifying, locating, or describing the content in video data to generate structured tags. Figure 1 A schematic diagram illustrating an application scenario of the sample data acquisition method provided according to embodiments of this specification is shown. For example... Figure 1 As shown, the application scenario 100 may include an operator 11, a video acquisition device 12, and a data processing system 13.

[0046] refer to Figure 1 Application scenario 100 includes at least two stages: a production-accompanying video data acquisition stage and a sample data acquisition stage. In the production-accompanying video data acquisition stage, the operator 11 can use the video acquisition device 12 to acquire video of the actual production task being performed, thus obtaining production-accompanying video data.

[0047] As an example, the video capture device 12 may be a terminal device with video recording capabilities. For example, the video capture device 12 may include mobile devices, tablets, laptops, built-in devices in motor vehicles, or similar content, or any combination thereof. In some embodiments, the mobile device may include wearable devices, shooting devices, smart mobile devices, virtual reality devices, augmented reality devices, or similar devices, or any combination thereof. In some embodiments, wearable devices include smartwatches, smart bracelets, smart glasses, etc. In some embodiments, the smart mobile device may include smartphones, personal digital assistants, gaming devices, navigation devices, etc., or any combination thereof. In some embodiments, the virtual reality device or augmented reality device may include head-mounted displays, virtual reality helmets, virtual reality glasses, virtual reality patches, augmented reality helmets, augmented reality glasses, augmented reality patches, or similar content, or any combination thereof. In some embodiments, the built-in device in the motor vehicle may include an in-vehicle camera.

[0048] For example, the video capture device 12 can be smart glasses with video recording capabilities. The operator 11 can wear the smart glasses and activate the recording function, then perform the actual production task. The smart glasses can record the operator performing the actual production task from a first-person perspective, and the video data obtained in this way is the collected production-related video data.

[0049] In some embodiments, during the sample data acquisition phase, the data processing system 13 can acquire the accompanying video data (i.e., target video data) collected during the accompanying video data acquisition phase, and annotate the accompanying video data to obtain sample data. The data processing system 13 can be deployed on a device or cluster of devices with data processing capabilities. For example, the data processing system 13 can be deployed on physical devices such as servers, server clusters, or cloud servers. In this case, the physical device corresponding to the data processing system 13 can store data or instructions for executing the sample data acquisition method described in this specification, and can execute or be used to execute the data or instructions. In some embodiments, the physical device corresponding to the data processing system 13 may include hardware devices with data information processing functions and the necessary programs required to drive the hardware devices.

[0050] It should be understood that Figure 1 The number of operators 11, video acquisition devices 12, and data processing systems 13 shown is merely illustrative. Depending on implementation needs, any number of operators 11, video acquisition devices 12, and data processing systems 13 can be included.

[0051] Figure 2 A hardware structure diagram of a computing system provided according to an embodiment of this specification is shown. The computing system 200 can serve as... Figure 1 The data processing system 13 in the specification executes the sample data acquisition method described herein.

[0052] like Figure 2 As shown, the computing system 200 may include at least one storage medium 230 and at least one processor 220. In some embodiments, the computing system 200 may also include a communication port 250 and an internal communication bus 210. The computing system 200 may also include I / O components 260.

[0053] The internal communication bus 210 can connect to different system components. For example, the internal communication bus 210 can connect to storage medium 230, processor 220, communication port 250, and I / O component 260, etc.

[0054] I / O component 260 supports input / output between computing system 200 and other components.

[0055] Communication port 250 is used for data communication between computing system 200 and the outside world. For example, communication port 250 can be used for data communication between computing system 200 and a network. Communication port 250 can be a wired communication port or a wireless communication port.

[0056] Storage medium 230 may include a data storage device. The data storage device may be a non-transitory storage medium or a temporary storage medium. For example, the data storage device may include one or more of a disk 232, a read-only storage medium (ROM) 234, or a random access storage medium (RAM) 235. Storage medium 230 also includes at least one instruction set stored in the data storage device. The instruction set may include computer program code, which may include programs, routines, objects, components, data structures, procedures, modules, etc.

[0057] At least one processor 220 may be communicatively connected to at least one storage medium 230. When the computing system 200 is running, at least one processor 220 reads the at least one instruction set and executes the sample data acquisition method provided in this specification according to the instructions of the at least one instruction set. The processor 220 may perform the steps included in the sample data acquisition method. The processor 220 may be in the form of one or more processors. In some embodiments, the processor 220 may include one or more hardware processors, such as a microcontroller, microprocessor, reduced instruction set computer (RISC), application-specific integrated circuit (ASIC), application-specific instruction set processor (ASIP), central processing unit (CPU), graphics processing unit (GPU), physical processing unit (PPU), microcontroller unit, digital signal processor (DSP), field-programmable gate array (FPGA), advanced RISC machine (ARM), programmable logic device (PLD), any circuit or processor capable of performing one or more functions, or any combination thereof.

[0058] For illustrative purposes only, the accompanying drawings show only one processor 220 for the computing system 200. However, it should be noted that the computing system 200 may also include multiple processors; therefore, the operations and / or method steps disclosed herein may be executed by one processor or by multiple processors in combination. For example, if the processor 220 of the computing system 200 described in this specification executes steps A and B, it should be understood that steps A and B may also be executed jointly or separately by two different processors 220 (e.g., a first processor executes step A, a second processor executes step B, or the first and second processors jointly execute steps A and B).

[0059] Figure 3 A flowchart of a sample data acquisition method according to an embodiment of this specification is shown. As previously described, the computing system 200 can execute the sample data acquisition method of this specification.

[0060] like Figure 3 As shown, sample data acquisition methods may include: S310: Acquire target video data to be labeled. The target video data to be labeled is production-accompanying video data collected by the operator from a first-person perspective using a wearable video acquisition device. Production-accompanying video data refers to the data generated by the operator during the execution of actual production tasks.

[0061] In some embodiments, referring to the description of the application scenario, the target video data is production-related video data collected from a first-person perspective by an operator carrying a wearable video acquisition device during the execution of an actual production task.

[0062] As an example, wearable video capture devices can be smart glasses with shooting capabilities, or action cameras or webcams that are fixed to the human body via straps, braces, or other components. Action cameras offer better shooting results and longer battery life; however, they are heavier, making them less suitable for extended wear. Smart glasses are lighter, but their shooting quality and battery life are inferior. In practical applications, the solutions provided in this manual allow for the selection of different wearable video capture devices based on video capture needs; this manual does not restrict the type of wearable video capture device.

[0063] In some embodiments, the target video data includes at least the execution process of the production task, the execution result of the production task, candidate objects in the production environment, and a target operating unit. In the target video data, the operator uses the target operating unit to manipulate the target object among the candidate objects to perform the production task. The target operating unit can be the operator's hand or a component used for manipulation. Components used for manipulation can include grippers, suction cups, tools, robotic arms, etc.

[0064] For example, a production task can be a material handling task. During the execution of such a task, the operator can pick up the target object by hand or operate a gripper to pick up the target object and move it to the target location.

[0065] Alternatively, production tasks may also include installation tasks. In the execution of such tasks, the operator can pick up the target object with one hand and the installation tool with the other, and use the tool to install the target object in the target location.

[0066] In this embodiment, when collecting target video data, the operator can directly use the components corresponding to the embodied robot to perform production tasks and collect target video data. Since the labeled target video data is used to train the embodied robot, this method of collecting target video data reduces the differences between human hand operations and the embodied robot's operating parts, facilitating generalization and transfer during training and improving the usability and adaptability of the sample data. The embodied robot trained using such sample data can perform production tasks more accurately and efficiently in real-world production environments.

[0067] Those skilled in the art will understand that different types of production tasks will utilize different types of target operating units. The components listed in this specification for operation are merely examples; in practical applications, selections can be made based on actual production needs. This specification does not impose any restrictions on the type of production task or the type of components used for operation.

[0068] Figure 4A schematic diagram of a preprocessing method for target video data provided according to an embodiment of this specification is shown.

[0069] In some embodiments, the data processing system may preprocess the target video data before labeling it. For example, see [reference needed]. Figure 4 The data processing system can perform at least one of the following preprocessing steps on the target video data to obtain preprocessed target video data: replacing at least one candidate object or target operation unit in the target video data; performing derivation generation processing on the target video data to generate a derivative video of the target video data, wherein the derivative video corresponds to a different production task or the execution result of the production task is different from that of the target video data; or editing a portion of the target video data.

[0070] In some embodiments, replacement processing refers to replacing candidate objects or target operating units in target video data using specified replacement objects. Replacement processing can be implemented using MLLMs with video processing capabilities.

[0071] For example, the data processing system can first acquire the replacement object, then generate prompts from the target video data and the replacement object to guide MLLMs in performing the replacement process. The data processing system can then input these prompts into the MLLMs, guiding them to replace the specified candidate object or target operation unit in the target video data with the replacement object, resulting in target video data for either the candidate object or the target operation unit. The target video data for either the candidate object or the target operation unit is the preprocessed target video data.

[0072] As an example, suppose the target manipulation part in the target video data is a hand, and it needs to be replaced with a gripper. The data processing system can then acquire an image of the component used for manipulation (i.e., the image of the gripper) as the replacement object. The data processing system can then use large video editing models (MLLMs) to replace the hand in the target video data.

[0073] For example, the data processing system can first generate prompts based on the target video data and the image of the gripper. These prompts include descriptive text instructing the replacement of the gripper image with the hand image in the target video data. Then, the data processing system can input the prompts into a large video editing model, guiding the model to perform the replacement process and generate target video data for the replaced target operation unit, resulting in preprocessed target video data.

[0074] As an example, suppose the candidate objects in the target video data include a box, a mobile phone, and a cup, and we need to replace the mobile phone with a tablet computer. The data processing system can then acquire an image of the tablet computer as the replacement object. Next, the data processing system can use a large video editing model to replace the mobile phone image in the target video data. The implementation method of using the large video editing model for replacement has been explained earlier and will not be repeated here.

[0075] In this embodiment, the data processing system expands the appearance and types of candidate objects or target operating parts in the target video data by replacing at least one candidate object or target operating part, thereby increasing the diversity of the data. Using the sample data obtained from these labeled data to train the embodied robot can effectively improve its generalization ability.

[0076] In some embodiments, derivative generation processing refers to generating a derivative video containing specified derivative content based on the target video data. Derivative generation processing can be implemented using large language models and text-based video models.

[0077] For example, a data processing system can first acquire preset derived content, and then, using this preset derived content as the target, generate descriptive text related to it using a large language model. The preset derived content includes content that contradicts the execution result of the production task in the target video data or content that differs from the production task in the target video data. Then, the data processing system can generate descriptive text based on the target video data and the content, and use a text-based video model to generate derived videos. Depending on the preset derived content, the derived videos can include derived videos with different production tasks (also called interfering videos), or derived video data with different execution results of the production task (also called counterfactual videos). The text-based video model can be MLLMs; for example, a large video editing model or an image-to-video (Img2Video) model can both be used as text-based video models.

[0078] As an example, suppose the production task in the target video data is to move a box to shelf number 1, and the target video data shows that the box was successfully moved to shelf number 1. Then, the pre-defined derivative content could include: the box falls during the movement (i.e., the production task fails). Or, the box is successfully moved to shelf number 2 (i.e., the production task is different).

[0079] In this scenario, the data processing system can first generate corresponding content description text based on preset derived content using LLM. For example, assuming the preset derived content is "the box fell while being moved," the data processing system can obtain prompts for text generation based on the preset derived content. These prompts guide the LLM to analyze and reason about the derived content, generating a detailed description related to the derived content, i.e., obtaining the content description text. As an example, the content description text can describe scene details, item details, movement trajectories, and operational actions related to the derived content.

[0080] Then, the data processing system can obtain prompts for text-based video generation based on the target video data and the content description text. These prompts guide the text-based video generation model to generate derived videos using the scenes in the target video data as templates. The content of the derived videos matches the scene details, object details, movement trajectories, and operational actions described in the content description text.

[0081] In this embodiment, the data processing system can generate derivative videos that differ from the target video data in terms of production tasks or execution results. These derivative videos can serve as interference or failure sample data for the production tasks, supplementing the erroneous samples in the sample data. Training the embodied robot using these erroneous samples can enhance its adaptability and robustness in complex scenarios.

[0082] In some embodiments, editing refers to modifying certain segments of the target video data, such as removing, splicing, speeding up, slowing down, or reversing certain segments. Editing the target video data can remove segments irrelevant to the production task, making the pre-processed target video data more focused on expressing the execution process of the production task, thereby improving the annotation efficiency and accuracy of the target video data.

[0083] S320: Label the target video data to obtain sample data, which is used to train the embodied robot to perform the actual production tasks performed by the operator.

[0084] In some embodiments, the target video data in S320 may be the preprocessed target video data in S310. The data processing system can annotate multiple preprocessed target video data separately to obtain multiple sample data. Then, the data processing system can generate a training sample set for the embodied robot based on the multiple sample data. By training the embodied robot using the training sample set, the trained embodied robot can perform the actual production tasks performed by the operator in the target video data.

[0085] Figure 5 A flowchart illustrating an annotation of target video data according to an embodiment of this specification is shown.

[0086] In some embodiments, refer to Figure 5 The data processing system first processes the target video data to obtain multiple video segments and their corresponding sub-tasks. Then, based on these video segments and their corresponding sub-tasks, the system performs annotation tasks in at least one dimension to obtain annotated data for each video segment. This annotated data includes reference data for at least one dimension. Finally, based on the target video data, the sub-tasks for each video segment, and the annotated data for each video segment, the system obtains sample data.

[0087] In some embodiments, the data processing system can split the target video data to obtain multiple video segments and corresponding subtasks for each video segment. As an example, the data processing system can first obtain splitting reference information for the target video data, which includes multiple splitting nodes and descriptions of the subtasks corresponding to each splitting node.

[0088] Figure 6 A schematic diagram illustrating a method for splitting target video data according to an embodiment of this specification is shown.

[0089] As an example, refer to Figure 6 Assuming the splitting reference information includes four splitting nodes and their corresponding subtask descriptions, the data processing system can split the target video data into four video segments based on the four splitting nodes, and determine the subtask corresponding to each video segment based on the subtask description information corresponding to each splitting node. The splitting nodes can be time nodes or frame number nodes; for example, splitting node 1 could be the 4th second in the target video data, or it could be the 120th frame in the target video data.

[0090] In some embodiments, the splitting reference information may be generated by manually annotating the target video data and then uploading it to the data processing system. Alternatively, the splitting reference information may be generated by the data processing system using a pre-trained model to perform inference analysis on the target video data. This specification does not limit the method of obtaining the splitting reference information.

[0091] Figure 7 A schematic diagram of a task annotation task in the task planning dimension is shown according to an embodiment of this specification.

[0092] In some embodiments, the annotation task of at least one dimension includes the annotation task of the task planning dimension, and the reference data corresponding to the task planning dimension includes: task status information. Figure 7 The data processing system can first generate a first prompt instruction based on multiple video clips and their corresponding sub-tasks. This first prompt instruction guides the Visual Large Language Model (VLM) to perform content recognition on the multiple video clips, and then performs task planning dimension inference based on the content recognition results and sub-tasks for each video clip. Furthermore, the data processing system can input the first prompt instruction into the Visual Large Language Model to obtain the task status information corresponding to each video clip, output by the Visual Large Language Model.

[0093] In some embodiments, task status information may include the execution status of a subtask, the planning information of the subtask, and the status information of the target object in the subtask within the corresponding video segment. As an example, task status information can be recorded in the form of question-and-answer pairs.

[0094] For example, the task status information corresponding to each video segment may include at least one task status question-answer pair. The task status question-answer pair includes question text and answer text. Accordingly, the first prompt instruction is also used to guide the visual large language model to retrieve multiple question texts from a pre-set question library, and based on the content recognition results and sub-tasks corresponding to each video segment, to perform task planning dimension reasoning and generate the answer text corresponding to the question text.

[0095] As an example, in the preset question library, the question text corresponding to the execution status of a subtask may include: What is the task objective of the subtask? What needs to be done now to achieve the task objective? Has the task objective of the subtask been successfully executed? The question text corresponding to the planning information of a subtask may include: Which step of the production task is the subtask? What steps were completed before the subtask? What are the subsequent steps of the subtask? The question text corresponding to the status information of the target object in the subtask may include: Where is the target object located? Can the target object be grabbed? Does grabbing the target object require movement?

[0096] The data processing system can guide the VLM to randomly select a preset number of question texts from a pre-defined question database using a first prompt command. Then, the data processing system, using the first prompt command, performs inference based on the content recognition results (e.g., content description text) generated by the VLM from multiple video clips to generate corresponding answer texts for each question text.

[0097] In some embodiments, the data processing system can generate a first prompt instruction based on multiple video segments, sub-tasks corresponding to the multiple video segments, a preset question library, and a prompt instruction template.

[0098] The following is an example of a first prompt instruction template: "You are a task planning expert, responsible for task planning from multiple input video clips. The input consists of a set of video clips, each corresponding to a specific subtask. Your goal is:" 1. Generate detailed content description text for each video segment.

[0099] 2. Obtain the question text associated with each subtask.

[0100] 3. Using the content description text of the video clips, answer the text of each question, and generate the corresponding answer text.

[0101] enter: List of video clips: [Video Clip 1, Video Clip 2, Video Clip 3, Video Clip 4].

[0102] Subtask list: [Subtasks corresponding to video clip 1, video clip 2, video clip 3, and video clip 4].

[0103] Question text: [Preset question library].

[0104] Processing steps: For each video segment, perform the following operations in sequence, based on the subtasks corresponding to that video segment: Step 1: Generate a content description text for the video clip. The content description text should include at least the events that occur in the video clip, the objects that appear, the sequence of actions, and the environmental context. Avoid subjective inferences and only describe objective facts based on visual content. Output: [Content Description Text].

[0105] Step 2: Based on the subtask corresponding to the current video segment, randomly select [5] question texts related to the subtask from the preset question library. Output: [List of Question Texts] 3. Using the content description text generated in step 1, answer each question text in step 2 one by one. Ensure that the answers are based on the video content, accurate, concise, and directly address the questions. Output: [List of answer texts].

[0106] Output requirements: All outputs must be based on the visual content of the video clips and must not incorporate external knowledge or fictitious information. Answer text should be clear, structured, and easy to use later. If a question cannot be answered from the video content, please indicate "Insufficient Information" and briefly explain why.

[0107] Output format: Please output the following for each video clip in the order they appear (using video clip 1 as an example): Video clip index: [e.g., video clip 1] Corresponding subtask: [Subtask corresponding to video clip 1] Content description: [For example, the description text for video clip 1] [Video Clip 1] Corresponding Task Status Question and Answer: Question: [Question text 1]; Answer: [Answer text 1].

[0108] Question: [Question text 2]; Answer: [Answer text 2].

[0109] Question: [Question text 3]; Answer: [Answer text 3].

[0110] Question: [Question text 4]; Answer: [Answer text 4].

[0111] Question: [Question Text 5]; Answer: [Answer Text 5]. In this context, [ ] represents a placeholder, and the content within [ ] is filled in based on the actual application or generated using VLM.

[0112] As an example, refer to Figure 7 The data processing system can load multiple video clips, their corresponding subtasks, and a pre-defined question library into the first prompt instruction template shown above, generating a first prompt instruction. Then, the data processing system inputs the first prompt instruction into the VLM, guiding the VLM to output the task status information corresponding to each video clip (i.e., at least one task status question-and-answer pair corresponding to each video clip).

[0113] In this embodiment, the data processing system generates a first prompt instruction based on multiple video clips and their corresponding sub-tasks, guiding the visual large language model to perform content recognition on the video clips and execute reasoning in the task planning dimension. The data processing system uses the task state question-answer pairs output by the visual large language model as task state information. This task planning dimension annotation method, by introducing structured question text and utilizing the content recognition results of the video clips and the corresponding answer text generated by the sub-tasks, provides a natural language-driven task planning dimension annotation method. The reference data corresponding to the task planning dimension generated using this annotation method has better interpretability, and the obtained task planning results are more reliable. Furthermore, using sample data generated based on this reference data to train the embodied robot can improve the embodied robot's task planning ability and error recovery ability in a production environment.

[0114] Figure 8A schematic diagram of an operational region dimension annotation task is shown according to an embodiment of this specification.

[0115] In some embodiments, the annotation task in at least one dimension includes an annotation task in the operable region dimension, and the reference data corresponding to the operable region dimension includes: the operable region. The data processing system can utilize a target recognition model to perform target recognition on multiple video segments, obtaining candidate objects present in the multiple video segments and their bounding boxes. Furthermore, the data processing system can, based on the sub-tasks corresponding to the multiple video segments, determine the target object corresponding to the sub-task from the candidate objects, and determine the operable region based on the bounding box of the target object.

[0116] Candidate objects refer to items, human body parts (such as arms and hands), and components used for manipulation that appear in multiple video clips.

[0117] In some embodiments, the data processing system can use Segment Anything Model 2 (SAM2) and Grounding Distillation with NO labels (Grounding DINO) to identify candidate objects and obtain their bounding boxes.

[0118] As an example, for any video clip, the data processing system can use Grounding DINO to perform frame-by-frame recognition of the video clip, identifying candidate objects in the video clip, as well as the bounding boxes of the candidate objects.

[0119] For example, the data processing system can first use object recognition instructions to direct Grounding DINO to identify mobile phones (i.e., candidate objects) in a video clip. After receiving the object recognition instructions, Grounding DINO can perform frame-by-frame recognition of the video clip to determine the bounding box of the mobile phone in each frame. The bounding box can be the smallest rectangle that can cover the area where the mobile phone is located. Then, the data processing system can use SAM2 to identify the segmentation mask corresponding to the mobile phone based on its bounding box. The edge contour of the segmentation mask corresponding to the mobile phone is the bounding box of the mobile phone. Finally, the data processing system can use the segmentation mask generated by SAM2 and the category label (i.e., mobile phone) identified by Grounding DINO to fuse with the video clip to obtain a video clip containing reference data corresponding to the dimensions of the operable region.

[0120] In some embodiments, the data processing system can replace Grounding DINO with image recognition-capable models such as You Only Look Once (YOLO) and Large Language and Vision Assistant (LLaVA) to perform image recognition tasks. Additionally, the data processing system can replace SAM2 with image segmentation-capable models such as Fast Segment Anything Model (FastSAM) and Segment Everything Everywhere All at Once Model (SEEM) to perform image segmentation tasks. This specification does not limit the types of models used for image recognition and image segmentation.

[0121] In some embodiments, the target object includes a target article and a target operating part, wherein the target operating part is a component or hand for operation. Correspondingly, the operable area includes: the area corresponding to the operable article and the area corresponding to the operating position of the operable article.

[0122] The process of determining the operable area based on the target object's bounding box includes at least one of the following steps: If the target object corresponding to the subtask includes a target item, the area marked by the bounding box of the target item is determined as the area corresponding to the operable item. If the target object corresponding to the subtask includes both a target item and a target operating unit, the overlapping area between the bounding boxes of the target operating unit and the target item is determined as the area corresponding to the operating position of the operable item.

[0123] As an example, suppose that in a frame of a video clip, the target objects include a tablet computer (target object) and the operator's hands (target operating parts), with the operator's two hands grasping the two ends of the tablet computer. The data processing system can first use Grounding DINO and SAM2 to identify the tablet computer and hands in the frame image, obtaining the bounding boxes of the tablet computer and the two hands. Then, the data processing system can determine that the area where the bounding box of the tablet computer is located corresponds to the area of ​​the operable object. Next, the data processing system can obtain the areas where the bounding boxes of the two hands overlap with the bounding boxes of the tablet computer, and determine that these two overlapping areas are the areas corresponding to the operating positions of the operable object.

[0124] Figure 9 A schematic diagram of a navigation trajectory dimension annotation task provided according to an embodiment of this specification is shown.

[0125] In some embodiments, the annotation task for at least one dimension includes a navigation trajectory dimension annotation task, and the reference data corresponding to the navigation trajectory dimension includes: navigation guidance text. Figure 9 The data processing system can first identify at least one navigation video segment among multiple video segments based on the sub-tasks corresponding to each segment. Then, the data processing system uses a visual large language model to generate navigation guidance text corresponding to at least one navigation video segment.

[0126] In some embodiments, the navigation guidance text is described using natural language, describing the movement trajectory of the target object in the navigation video clip within a production environment. The navigation guidance text includes navigation direction, navigation distance, and navigation angle. The navigation direction is obtained by the visual large language model through inference of the navigation trajectory dimension based on multiple target images. The navigation distance and navigation angle are obtained by the visual large language model through inference of the navigation trajectory dimension based on the spatial coordinates corresponding to the target object in each of the multiple target images. As an example, the navigation direction may include going straight or backward; the navigation distance may include the distance traveled straight or backward; and the navigation angle may include the direction and angle of turning. For example, the navigation guidance text may include: "Confirm that the target object has been captured at the starting position, go straight forward for 5 meters, turn left 90 degrees, go straight for 10 meters, turn right 90 degrees, go straight for 8 meters, and arrive at the target location."

[0127] In some embodiments, reference Figure 9 The data processing system first identifies multiple target images within the navigation video clip. These target images are those that meet preset conditions. Then, the system uses a spatial reconstruction model to obtain the spatial coordinates of the target objects within each of the multiple target images. Next, based on the multiple target images and their corresponding spatial coordinates, the system generates a second prompt instruction. This second prompt instruction guides the visual big data language model to perform content recognition on the multiple target images and to infer the navigation trajectory dimension based on the spatial coordinates of the target objects. Finally, the system inputs the second prompt instruction into the visual big data language model to obtain the navigation guidance text output by the model.

[0128] The target image can be a frame in the navigation video clip where the direction of movement of the target object changes. For example, referring to the example of navigation guidance text above, multiple frames in the navigation video clip during the process of changing the navigation action from going straight to turning left can all be identified as the target object.

[0129] As an example, the data processing system can use an entropy-based sampling method to identify multiple target images that meet preset conditions in a navigation video clip. Specifically, when identifying multiple target images using the entropy-based sampling method, the data processing system can determine the preset condition as follows: the entropy of the probability distribution of the predicted navigation action corresponding to the image is greater than a preset threshold.

[0130] In some embodiments, the data processing system can utilize a pre-trained entropy prediction model to obtain the entropy corresponding to each frame of a navigation video segment. This entropy prediction model has the ability to determine the probability distribution of predicted navigation actions based on the image. Then, the data processing system can identify images among multiple images whose entropy is greater than a preset threshold as target images.

[0131] As an example, the data processing system can use an entropy prediction model to identify the probability distribution of navigation actions corresponding to each frame in a navigation video clip. For example, suppose the navigation actions include: going straight, turning left, and turning right. Referring to the example of navigation guidance text above, in a navigation video clip where the navigation action is to move forward 5 meters straight, the data processing system uses an entropy prediction model to predict the navigation action for the corresponding frame. For example, in this case, the predicted probability distribution of the navigation action might be: [going straight: 0.95, turning left: 0.02; turning right: 0.03]. The data processing system can then use the entropy prediction model to calculate the entropy of the predicted probability distribution of the navigation action (the calculated entropy is 0.34).

[0132] Correspondingly, during the transition from a straight-ahead to a left-turn navigation action, the data processing system uses an entropy prediction model to calculate the corresponding frame for navigation action prediction. For example, in this case, the predicted probability distribution of the navigation action might be: [Straight-ahead: 0.35, Left-turn: 0.4; Right-turn: 0.25]. The data processing system can then use the entropy prediction model to calculate the entropy of the predicted probability distribution of the navigation action (the calculated entropy is 1.56).

[0133] Based on the above example, if the preset threshold is set to 1, the entropy of the corresponding frame during the process of moving straight forward for 5 meters is less than the preset threshold; the entropy of the corresponding frame during the process of changing from straight forward to left turn is greater than the preset threshold. Therefore, the corresponding frame during the process of changing from straight forward to left turn can be identified as the target image.

[0134] In this embodiment, the data processing system uses a pre-trained entropy prediction model to obtain the entropy corresponding to each frame of the navigation video segment. This entropy reflects the uncertainty of the predicted navigation action probability distribution. Specifically, a smaller entropy indicates that the probability distribution among different navigation actions is more concentrated on a single navigation action, while a larger entropy indicates a more uniform probability distribution among different navigation actions. The data processing system identifies frames with entropy greater than a preset threshold as target images, ensuring that the target images correspond to key nodes with high uncertainty in the navigation process. When generating navigation guidance text based on the target images, the data processing system can analyze and process only the images corresponding to key nodes, thereby improving the efficiency of navigation guidance text generation and reducing the resource consumption of the data processing system.

[0135] In some embodiments, the spatial reconstruction model can be a Matching and Stereo 3D Reconstruction (MASt3R) model. MASt3R can identify the spatial coordinates of the target object in 3D space in each of multiple discrete planar images (i.e., multiple target images) including the target object. The spatial coordinates of the target object in the multiple target images constitute the movement trajectory of the target object, and the movement trajectory of the target object in 3D space can be the navigation trajectory of the target object.

[0136] In some embodiments, the data processing system can generate a second prompt instruction based on multiple target images and the spatial coordinates of the target objects in the multiple target images.

[0137] The following is an example of a second prompt instruction template: "You are a navigation annotation expert responsible for annotating input images. These images correspond to key nodes in navigation video clips. Each target image contains a target object. Your goal is:" 1. Perform content recognition on each target image to generate objective and detailed image description text.

[0138] 2. Based on the spatial coordinates of the target object in each target image, perform reasoning on the navigation trajectory dimension to determine the navigation route (including movement direction, distance changes, and direction turning points).

[0139] 3. Generate navigation guidance text based on the navigation route. The navigation guidance text needs to describe the movement trajectory of the target object. The navigation guidance text must be based solely on the provided image content and coordinate data, using objective facts and directional descriptions (such as "turn left 90°", "turn right 90°", "move forward x meters", etc.), avoiding subjective inferences or external knowledge.

[0140] enter: List of target images: [target image 1, target image 2, target image 3, target image 4].

[0141] Target object description: [Name or description of the target object].

[0142] List of spatial coordinates: [Spatial coordinates of the target object in target image 1, spatial coordinates of the target object in target image 2, spatial coordinates of the target object in target image 3, spatial coordinates of the target object in target image 4].

[0143] Processing steps: For multiple input target images, please perform the following operations in sequence: Step 1: Analyze each target image in the target image list and generate a concise and objective descriptive text. The description should include the target object's appearance in the image, its location context, and any relevant environmental details (such as background objects or scene layout). Avoid subjective interpretations and base your analysis solely on visual content.

[0144] Step 2: Calculate the navigation trajectory of the target object based on the list of target images and the list of spatial coordinates.

[0145] Step 3: Combine the content description text of the target images and the navigation trajectory of the target object to generate a coherent navigation guide text. The navigation guide text should describe the navigation trajectory of the target object from the first target image to the last target image, using objective facts and specified directions. The description should include: the start and end points; the direction and distance of movement, e.g., "Move forward 10 meters"; and the points of directional change and the angle of directional change, e.g., "Turn 90° left at point P". Output Requirements: All descriptions and inferences must be based on the provided image content and spatial coordinates, and must not introduce external knowledge or fictitious information. Navigation guidance text should be clear, concise, and focused on an objective description of the movement trajectory. Use directional terms such as "turn left 90°", "turn right 90°", "go straight", etc. If the coordinate data cannot infer a specific directional change, please indicate "insufficient information" and explain the reason. Ensure the output is structured for easy subsequent use.

[0146] Output format: Please output the following: [Navigation guide text]. In this context, [ ] represents a placeholder, and the content within [ ] is filled in based on the actual application or generated using VLM.

[0147] As an example, refer to Figure 9The data processing system can load multiple target images and the spatial coordinates of the target objects in each image into the second prompt instruction template shown above, generating a second prompt instruction. Then, the data processing system inputs the second prompt instruction into the VLM, guiding the VLM to output navigation guidance text.

[0148] Figure 10 A schematic diagram of an operation trajectory dimension annotation task provided according to an embodiment of this specification is shown.

[0149] In some embodiments, the annotation task for at least one dimension includes an annotation task for the operation trajectory dimension, and the reference data corresponding to the operation trajectory dimension includes the movement trajectory of the target operation unit. (Reference) Figure 10 The data processing system can first determine at least one operation video segment involving operations based on the sub-tasks corresponding to multiple video segments. Then, the data processing system can obtain the target operator's coordinate sequence as its spatial coordinates change over time within the at least one operation video segment, and generate the target operator's movement trajectory based on the target operator's coordinate sequence. The target operator's movement trajectory can be represented by the target operator's spatial coordinate sequence within the operation video segment (i.e., the target operator's coordinate sequence).

[0150] In some embodiments, the data processing system can utilize a spatial reconstruction model to identify the spatial coordinates of the target operating part. For example, assuming the target operating part includes a hand or a component for operation, the data processing system can first use a target recognition model to identify the target operating part in the operation video segment, as well as at least one key point corresponding to the target operating part. Then, the data processing system can further utilize the spatial reconstruction model to obtain the spatial coordinates of the key points in each frame of the operation video segment. Finally, the data processing system can generate a sequence of target operating part coordinates corresponding to the operation video segment based on the spatial coordinates of the key points in each frame.

[0151] In some embodiments, the data processing system can also utilize a positioning module to acquire the spatial coordinates of the target operating part. For example, assuming the target operating part includes a hand or a component for operation, the target operating part may include at least one positioning module. The positioning module may include a three-axis gyroscope, accelerometer, etc. For example, when the target operating part is a hand, the positioning module can be integrated into a wearable device and worn on the hand. The positioning module can be integrated into wearable devices such as smart gloves, smartwatches, smart bracelets, and smart rings. Alternatively, when the target operating part is a component for operation, the positioning module can be integrated into that component.

[0152] When acquiring target video data, the data processing system can also acquire the corresponding location data. For example, when acquiring target video data using a wearable video acquisition device, the device can also communicate with a positioning module to record the location data collected by the module. After determining the operation video segment, the data processing system can extract the location data segment corresponding to the operation segment from the location data of the target video data, using this segment as the location data for the operation video segment. Then, based on the operation video segment and its corresponding location data, the system can generate a target operation unit coordinate sequence.

[0153] In some embodiments, the data processing system marks the spatial movement trajectory of key points of the target operating unit in the operation video segment based on the target operating unit's coordinate sequence, thereby generating the target operating unit's movement trajectory. As an example, the data processing system can mark the operation video segment using highlighted, brightly colored lines, arrows, etc., and this specification does not limit the marking method.

[0154] In this embodiment, the data processing system can combine multi-dimensional annotation tasks to annotate the target video data, obtaining multi-dimensional annotated data. This annotated data includes reference data needed for production task planning and execution. The sample data generated using multi-dimensional annotated data can meet the training needs of various types of production tasks, and is a training sample with better applicability and high transferability.

[0155] In some embodiments, after obtaining the annotation data corresponding to multiple video segments, the data processing system can further perform accuracy verification on the obtained annotation data corresponding to the multiple video segments, and obtain accuracy verification results. The accuracy verification results are used to characterize whether the annotation data is accurate. Furthermore, when the accuracy verification result indicates that the annotation data is inaccurate, the annotation data is corrected.

[0156] As an example, the data processing system can display each video clip and its corresponding annotation data on a terminal device, instructing staff performing the data annotation work to verify the accuracy of the annotation data. The terminal device can be a computer, mobile phone, tablet, personal digital assistant, or other similar device used by the staff.

[0157] Then, the data processing system can receive the accuracy verification result from the terminal device. This result can indicate that the labeled data is accurate or inaccurate. When the accuracy verification result indicates that the labeled data is inaccurate, the terminal device can also prompt the staff to upload accurate labeled data and send both the accuracy verification result and the accurate labeled data to the data processing system. Correspondingly, when the data processing system receives the accuracy verification result and the accurate labeled data from the terminal device, it can replace the inaccurate labeled data with the accurate data.

[0158] In this embodiment, the data processing system can display video clips and corresponding annotation data to staff via a terminal device and receive accuracy verification results submitted by the staff. If the accuracy verification result indicates that the annotation data is inaccurate, the data processing system can also use the terminal device to obtain accurate annotation data uploaded by the staff, and then use the accurate annotation data to replace the erroneous annotations. Through this accuracy verification method, the data processing system can reduce the probability of errors in the annotation data and improve the accuracy and reliability of the sample data.

[0159] In some implementations, the embodied robot may include a motion model and a navigation model. Sample data can be used to train the motion and navigation models in the embodied robot, enabling it to perform actual production tasks performed by an operator. Specifically, the embodied robot can use the motion model to identify the scene, plan the task, and determine the actions to be performed, as well as the execution details of those actions. The embodied robot can use the navigation model to plan a navigation route to move itself to a target location.

[0160] As an example, the labeled data in the sample data can include: reference data corresponding to the task planning dimension, reference data corresponding to the operable area dimension, reference data corresponding to the navigation trajectory dimension, and reference data corresponding to the operation trajectory dimension. The reference data corresponding to the task planning dimension, the operable area dimension, and the operation trajectory dimension can be used to train the action model. The reference data corresponding to the navigation trajectory dimension can be used to train the navigation model.

[0161] In summary, the sample data acquisition method provided in this specification allows the data processing system to annotate target video data to obtain sample data. The target video data is production-accompanying video data collected during actual production processes. This data contains visual information about the real production environment and the complete production task execution process. By annotating the production-accompanying video data and obtaining sample data, the sample data can characterize the process of executing production tasks in the actual production environment. Using this sample data to train embodied robots can improve the matching degree and adaptability between the embodied robot and the actual production environment, making the robot's behavior more aligned with the actual production environment. This, in turn, can improve the embodied robot's execution capability and overall performance in actual production tasks. Specifically, when annotating the sample data, the data processing system can combine multiple dimensions of annotation tasks to annotate the target video data. The resulting sample data has better applicability and higher transferability, and can be applied to the training of various types of production tasks.

[0162] This specification, in another aspect, provides a computer-readable non-transitory storage medium storing at least one set of instructions for acquiring sample data. When the at least one set of instructions is executed by a processor, it instructs the processor to perform the steps of the sample data acquisition method described in this specification. In some possible embodiments, various aspects of this specification can also be implemented as a program product comprising program code. When the program product is run on a computing system 200, the program code causes the computing system 200 to perform the steps of the sample data acquisition method described in this specification. The program product for implementing the above method may employ a portable compact disc read-only memory (CD-ROM) containing program code and may run on the computing system 200. However, the program product of this specification is not limited thereto. In this specification, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system. The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. The computer-readable storage medium may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium that can send, propagate, or transmit programs for use by or in connection with an instruction execution system, apparatus, or device. Program code contained on a readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof. Program code for performing the operations described herein can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on computing system 200, partially on computing system 200, as a standalone software package, partially on computing system 200 and partially on a remote computing device, or entirely on a remote computing device.

[0163] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0164] In summary, after reading this detailed disclosure, those skilled in the art will understand that the foregoing detailed disclosure is presented by way of example only and is not restrictive. Although not explicitly stated herein, those skilled in the art will understand that this specification requires various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be made by this specification and are within the spirit and scope of the exemplary embodiments described herein.

[0165] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "an embodiment," "an embodiment," and / or "some embodiments" mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of this specification. Therefore, it is to be emphasized and understood that two or more references to "an embodiment" or "an embodiment" or "alternative embodiment" in various parts of this specification do not necessarily refer to the same embodiment. Moreover, specific features, structures, or characteristics may be suitably combined in one or more embodiments of this specification.

[0166] It should be understood that in the foregoing description of the embodiments in this specification, various features are combined in a single embodiment, drawing, or description for the purpose of simplifying the description and to aid in understanding a feature. However, this does not mean that the combination of these features is necessary, and those skilled in the art, upon reading this specification, may readily identify some of the devices as separate embodiments. That is, the embodiments in this specification can also be understood as an integration of multiple secondary embodiments. And the content of each secondary embodiment is valid even if it contains fewer than all the features of a single foregoing disclosed embodiment.

[0167] Every patent, patent application, publication of a patent application, and other material such as articles, books, specifications, publications, documents, articles, etc., cited herein, except for those inconsistent with or conflicting with this document, or those having a restrictive effect on the widest scope of the claims, may be incorporated herein by reference for all purposes now or hereafter associated with this document. Furthermore, in the event of any inconsistency or conflict between the description, definition, and / or use of relevant terms in any material and the description, definition, and / or use of relevant terms in this document, the terms in this document shall prevail.

[0168] Finally, it should be understood that the embodiments disclosed herein are illustrative of the principles of the embodiments described in this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can implement the applications described in this specification using alternative configurations based on the embodiments in this specification. Therefore, the embodiments in this specification are not limited to the embodiments precisely described in the applications.

Claims

1. A method for acquiring sample data, the method comprising: Acquire target video data to be labeled, wherein the target video data to be labeled is production-related video data collected by an operator from a first-person perspective using a wearable video acquisition device, and the production-related video data refers to data generated by the operator during the execution of actual production tasks; and The target video data is labeled to obtain sample data, which is used to train the android to perform the actual production tasks performed by the operator.

2. The method according to claim 1, wherein, The step of annotating the target video data to obtain sample data includes: The target video data is processed to obtain multiple video segments and sub-tasks corresponding to each of the multiple video segments; Based on the plurality of video segments and the sub-tasks corresponding to each of the plurality of video segments, an annotation task of at least one dimension is performed to obtain annotation data corresponding to each of the plurality of video segments. The annotation data includes: reference data corresponding to the at least one dimension, and The sample data is obtained based on the target video data, the sub-tasks corresponding to the multiple video segments, and the annotation data corresponding to the multiple video segments.

3. The method according to claim 2, wherein, The annotation task of at least one dimension includes the annotation task of the task planning dimension, and the reference data corresponding to the task planning dimension includes: task status information; The step of performing a labeling task of at least one dimension based on the plurality of video segments and the sub-tasks corresponding to the plurality of video segments includes: A first prompt instruction is generated based on the plurality of video segments and the sub-tasks corresponding to each of the plurality of video segments. This first prompt instruction guides the visual large language model to perform content recognition on the plurality of video segments, and to perform task planning dimension reasoning based on the content recognition results and sub-tasks corresponding to the plurality of video segments; and... The first prompt instruction is input into the visual big language model to obtain task status information corresponding to multiple video segments output by the visual big language model.

4. The method according to claim 3, wherein, The task status information includes at least one task status question-and-answer pair, which includes question text and answer text. Accordingly, the first prompt instruction is also used to guide the visual big language model to obtain multiple question texts from a preset question library, and to perform task planning dimension reasoning based on the content recognition results and sub-tasks corresponding to each video segment, thereby generating the answer text corresponding to the question text.

5. The method according to claim 2, wherein, The annotation task of at least one dimension includes the annotation task of the operable region dimension, and the reference data corresponding to the operable region dimension includes: operable region; The step of performing a labeling task of at least one dimension based on the plurality of video segments and the sub-tasks corresponding to the plurality of video segments includes: The target recognition model is used to identify targets in the multiple video segments, obtaining candidate objects present in the multiple video segments and bounding boxes of the candidate objects; and Based on the sub-tasks corresponding to the multiple video segments, the target object corresponding to the sub-task is determined from the candidate objects, and the operable area is determined according to the bounding box of the target object.

6. The method according to claim 5, wherein, The operable area includes: the area corresponding to the operable item and the area corresponding to the operation position of the operable item; the target object includes the target item and the target operation part; the target operation part includes components or a hand for operation. Determining the operable area based on the bounding box of the target object includes at least one of the following steps: If the target object corresponding to the subtask includes the target item, the area marked by the annotation box of the target item is determined to be the area corresponding to the operable item; If the target object corresponding to the subtask includes the target item and the target operating part, the overlapping area of ​​the annotation box of the target operating part and the annotation box of the target item is determined as the area corresponding to the operating position of the operable item.

7. The method according to claim 2, wherein, The annotation task for at least one dimension includes the annotation task for the navigation trajectory dimension, and the reference data corresponding to the navigation trajectory dimension includes: navigation guidance text; The step of performing a labeling task of at least one dimension based on the plurality of video segments and the sub-tasks corresponding to the plurality of video segments includes: Based on the sub-tasks corresponding to the plurality of video segments, at least one navigation video segment involving navigation is identified from among the plurality of video segments; and The navigation guidance text corresponding to the at least one navigation video segment is generated using a visual large language model.

8. The method according to claim 7, wherein, The step of generating navigation guidance text corresponding to the at least one navigation video segment using a visual large language model includes: Multiple target images are identified in the navigation video clip, and the target images are images in the navigation video clip that meet preset conditions; The spatial coordinates of the target objects in the multiple target images are obtained using a spatial reconstruction model. Based on the plurality of target images and the spatial coordinates corresponding to the target objects in the plurality of target images, a second prompt instruction is generated. This second prompt instruction guides the visual large language model to perform content recognition on the plurality of target images and to infer the navigation trajectory dimension based on the spatial coordinates corresponding to the target objects in the plurality of target images; and The second prompt instruction is input into the visual big language model to obtain the navigation guidance text output by the visual big language model.

9. The method according to claim 8, wherein, The preset conditions include: the entropy of the probability distribution of the predicted navigation action corresponding to the image is greater than a preset threshold. The step of determining multiple target images in the navigation video segment includes: The entropy corresponding to each frame of the navigation video segment is obtained using a pre-trained entropy prediction model, which has the ability to determine the probability distribution of predicted navigation actions based on the images; and The image whose entropy is greater than a preset threshold among multiple images is identified as the target image.

10. The method according to claim 9, wherein, The navigation guidance text includes navigation direction, navigation distance, and navigation angle. The navigation direction is obtained by the visual big language model based on the navigation trajectory dimension inference of the multiple target images. The navigation distance and navigation angle are obtained by the visual big language model based on the spatial coordinates of the target objects in the multiple target images.

11. The method according to claim 2, wherein, The annotation task of at least one dimension includes the annotation task of the operation trajectory dimension, and the reference data corresponding to the operation trajectory dimension includes: the movement trajectory of the target operation unit; The step of performing a labeling task of at least one dimension based on the plurality of video segments and the sub-tasks corresponding to the plurality of video segments includes: Based on the sub-tasks corresponding to the plurality of video segments, at least one operation video segment involving an operation is determined from among the plurality of video segments; and The target operator obtains a target operator coordinate sequence in the at least one operation video segment, which shows the spatial coordinates of the target operator changing over time, and generates the movement trajectory of the target operator based on the target operator coordinate sequence.

12. The method according to claim 11, wherein, The target operating unit includes a hand or a component for operation. The step of obtaining a sequence of target operating unit coordinates that change over time in the at least one operation video segment includes: The target operation part in the operation video segment is identified using a target recognition model, as well as at least one key point corresponding to the target operation part; The spatial coordinates of the key points in each frame of the operation video segment are obtained using a spatial reconstruction model; and Based on the spatial coordinates of the key points in each frame, a target operation unit coordinate sequence corresponding to the operation video segment is generated.

13. The method according to claim 11, wherein, The target operating part includes a hand or a component for operation, and the target operating part includes at least one positioning module; The step of obtaining the target operator coordinate sequence, in the at least one operation video segment, of spatial coordinates changing over time, includes: The positioning data corresponding to the operation video segment is obtained. The positioning data is collected by the positioning module in the target operation unit when collecting target video data. as well as Based on the operation video clip and the positioning data, a target operation unit coordinate sequence corresponding to the target operation unit is generated.

14. The method according to claim 1, wherein, Before annotating the target video data to obtain sample data, the process further includes: The target video data is preprocessed by performing at least one of the following preprocessing steps to obtain preprocessed target video data: Replace at least one candidate object or target operation unit in the target video data; The target video data is subjected to derivative generation processing to generate a derivative video of the target video data. This derivative video differs from the target video data in either the production task it corresponds to or the execution result of that production task. A portion of the target video data is edited.

15. The method of claim 14, wherein, The target operating part in the target video data is a hand. Replacing the target operating part in the target video data includes: Obtain the image of the component to be manipulated as the replacement object; and Using a large video editing model, the replacement object is placed at the position of the hand in the target video data.

16. The method of claim 14, wherein, The step of performing derivative generation processing on the target video data to generate a derivative video of the target video data includes: Using predefined derived content as the target, a large language model is used to generate descriptive text related to the predefined derived content. The predefined derived content includes: content that contradicts the execution result of the production task in the target video data, and content that differs from the production task in the target video data; and Based on the target video data and the content, descriptive text is generated, and the derived video is generated using a text-based video model.

17. The method according to any one of claims 2-16, wherein, After obtaining the annotation data corresponding to the plurality of video segments respectively, the method further includes: The accuracy of the annotation data corresponding to the multiple video segments is verified, and the accuracy verification results are obtained. These accuracy verification results characterize whether the annotation data is accurate. When the accuracy verification result indicates that the labeled data is inaccurate, the labeled data is corrected.

18. The method according to any one of claims 1-16, wherein, The avatar robot includes a motion model and a navigation model. The sample data is used to train the motion model and navigation model of the avatar robot so that the avatar robot can perform the actual production tasks performed by the operator.

19. A data processing system, comprising: At least one storage medium storing at least one instruction set for acquiring sample data; as well as At least one processor is communicatively connected to the at least one storage medium, wherein, when the data processing system is running, the at least one processor reads the at least one instruction set and implements the method as described in any one of claims 1-18 according to the instructions of the at least one instruction set.

20. A computer-readable non-volatile storage medium, wherein, The computer-readable non-volatile storage medium stores at least one instruction set, which, when executed by at least one processor, implements the method as described in any one of claims 1-18.

Citation Information

Cited By

  • Robot generalization control method, system and equipment

    CN122210628A