Model training data determination method, apparatus, device, medium, and product
By automatically parsing multimodal input data to generate simulation environment configuration files and executing target trajectories, the high cost of obtaining high-quality training data and the semantic gap in simulation data are solved, thus achieving efficient generation of robot learning training data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DEXFORCE TECH CO LTD
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-12
AI Technical Summary
Acquiring large-scale, high-quality visual-language-action demonstration data in robot learning presents challenges such as high cost, low efficiency, and questionable safety. Direct real-world data collection is costly, while purely simulated data faces the semantic and dynamic gap between simulation and reality.
By acquiring multimodal input data from the task scenario, including historical operation videos, historical operation action trajectories, and natural language task descriptions, the simulation environment configuration file is automatically parsed to generate a set of pose trajectories and execute the target trajectory in the physical simulation environment, thereby generating training data for the target processing model.
It achieves automated conversion from real-world teleoperated videos to simulation environments, generating diverse and physically reasonable large-scale training data, thereby improving the training efficiency and model performance of data-driven robot learning methods.
Smart Images

Figure CN122196543A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, medium, and product for determining model training data. Background Technology
[0002] In the fields of embodied intelligence and robot learning, acquiring large-scale, high-quality visual-language-motor demonstration data is a key bottleneck in training high-performance models. Directly collecting real-world data is costly, inefficient, and raises security concerns, while purely simulated data faces a semantic and dynamic gap between simulation and reality. Summary of the Invention
[0003] This disclosure provides a method, apparatus, device, medium, and product for determining model training data, which breaks through the limitation of the original data scale and realizes the automatic conversion of easily obtainable real-world teleoperated videos into large-scale training data that is programmable, scalable, and physically reasonable within a simulation environment.
[0004] Firstly, a method for determining model training data is provided, including: Acquire multimodal input data for the task scenario; the multimodal input data includes historical operation videos, historical operation action trajectories, and natural language task descriptions; The natural language task description is parsed to determine the simulation environment configuration file corresponding to the task scenario; the simulation environment configuration file includes the simulation object of each target object in the task scenario. A set of pose trajectories is determined based on the historical operation video and the historical operation action trajectory; the set of pose trajectories includes the pose trajectory corresponding to each target object in the task scene and the associated trajectory annotation information; In a physical simulation environment, a target trajectory set is generated based on the pose trajectory set; the physical simulation environment is determined based on the simulation environment configuration file. Each target trajectory in the target trajectory set is executed in the physical simulation environment to generate training data for the target processing model based on the execution results of each target trajectory; the target processing model is used to process tasks in the task scenario.
[0005] Secondly, a device for determining model training data is provided, comprising: The data acquisition module is used to acquire multimodal input data of the task scenario; the multimodal input data includes historical operation videos, historical operation action trajectories, and natural language task descriptions. The parsing module is used to parse the natural language task description to determine the simulation environment configuration file corresponding to the task scenario; the simulation environment configuration file includes the simulation object of each target object in the task scenario; The pose trajectory set determination module is used to determine a pose trajectory set based on the historical operation video and the historical operation action trajectory; the pose trajectory set includes the pose trajectory corresponding to each target object in the task scene and the associated trajectory annotation information; The target trajectory set determination module is used to generate a target trajectory set based on the pose trajectory set in a physical simulation environment; the physical simulation environment is determined based on the simulation environment configuration file. The training data determination module is used to execute each target trajectory in the target trajectory set in the physical simulation environment to generate training data for the target processing model based on the execution results of each target trajectory; the target processing model is used to process tasks in the task scenario.
[0006] Thirdly, an electronic device is provided, comprising: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the method for determining model training data as described in the first aspect above.
[0007] Fourthly, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method for determining model training data as described in the first aspect above.
[0008] Fifthly, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the method for determining model training data as described in the first aspect above.
[0009] This disclosure provides a method, apparatus, device, medium, and product for determining model training data. The method includes: acquiring multimodal input data of a task scenario; the multimodal input data includes historical operation videos, historical operation action trajectories, and natural language task descriptions; parsing the natural language task descriptions to determine a simulation environment configuration file corresponding to the task scenario; the simulation environment configuration file includes simulation objects for each target object in the task scenario; determining a set of pose trajectories based on the historical operation videos and historical operation action trajectories; the set of pose trajectories includes pose trajectories corresponding to each target object in the task scenario and associated trajectory annotation information; generating a set of target trajectories based on the set of pose trajectories in a physical simulation environment; the physical simulation environment is determined based on the simulation environment configuration file; executing each target trajectory in the set of target trajectories in the physical simulation environment to generate training data for a target processing model based on the execution results of each target trajectory; the target processing model is used to process tasks in the task scenario. This technical solution first acquires multimodal input data including historical operation videos, motion trajectories, and natural language descriptions. It then automatically configures the corresponding physical simulation environment by parsing the natural language descriptions. Next, it extracts pose trajectories and their annotation information by combining the video and trajectory data. The target trajectory is then generated and executed within the constructed simulation environment. Finally, based on the execution results, data for training the task processing model is generated. This solution automatically transforms readily available real-world teleoperation videos into large-scale, programmable, scalable, and physically plausible training data within a simulation environment. It fundamentally overcomes the limitations of the original data size, enabling the generation of diverse and physically plausible training data, and significantly improving the training efficiency and model performance of data-driven robot learning methods.
[0010] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this disclosure, nor is it intended to limit the scope of the embodiments of this disclosure. Other features of the embodiments of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart of a method for determining model training data provided in Embodiment 1 of this disclosure; Figure 2This is a schematic diagram illustrating the execution process of a method for determining model training data provided in Embodiment 1 of this disclosure; Figure 3 This is a schematic diagram of the structure of a device for determining model training data provided in Embodiment 2 of this disclosure; Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of this disclosure. Detailed Implementation
[0013] To enable those skilled in the art to better understand the solutions of the embodiments of this disclosure, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the protection scope of the embodiments of this disclosure.
[0014] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0015] Example 1 Figure 1 This is a flowchart illustrating a method for determining model training data according to Embodiment 1 of this disclosure. This embodiment is applicable to situations involving the determination of model training data. The method can be executed by a device for determining model training data. This device can be implemented in hardware and / or software and can be configured in an electronic device, including but not limited to computers, PCs, electronic devices, and servers, which are devices with data processing capabilities. Figure 1 As shown, the method includes: S110. Obtain multimodal input data for the task scenario; the multimodal input data includes historical operation videos, historical operation action trajectories, and natural language task descriptions.
[0016] In this embodiment, the task scenario can be a real-world scenario requiring data transformation. Multimodal input data for the task scenario can be acquired, which may include historical operation videos, historical operation action trajectories, and natural language task descriptions collected from the task scenario when humans complete a specific task.
[0017] Among them, historical operation videos can be visual image data of the task execution process, historical operation action trajectories can be action trajectory data of the operation end and the spatial motion state of the object, and natural language task descriptions can be text instruction data describing the task objectives, operation objects, and execution requirements.
[0018] S120. Parse the natural language task description to determine the simulation environment configuration file corresponding to the task scenario; the simulation environment configuration file includes the simulation object of each target object in the task scenario.
[0019] It is understood that after obtaining the natural language task, natural language processing (NLP) techniques can be used to parse the task description, extract the task objectives, core objects, and functional constraints, and use spatial reasoning to determine the layout relationships between objects, thereby generating corresponding 3D objects. Finally, a simulation environment configuration file is constructed, containing the simulation objects corresponding to all target physical objects and their simulation scene layout parameters. The simulation configuration file can be a configuration file that the physics simulation engine can recognize and load. The simulation environment configuration file can include which simulation objects the physics simulation engine needs to create, the geometry of the simulation objects, their physical properties (such as mass, friction coefficient, etc.), initial positions, relationships between them, and simulation environment parameters.
[0020] S130. Determine a set of pose trajectories based on historical operation videos and historical operation action trajectories; the set of pose trajectories includes the pose trajectory corresponding to each target object in the task scene and the associated trajectory annotation information.
[0021] It is known that after acquiring historical operation videos and historical operation motion trajectories, real-world operation data can be transformed into structured, reusable trajectory information. That is, historical operation videos and historical operation motion trajectories can be analyzed to determine a set of pose trajectories. This set of pose trajectories can include the pose trajectory corresponding to each target object and its associated trajectory annotation information.
[0022] S140. In the physical simulation environment, a target trajectory set is generated based on the pose trajectory set; the physical simulation environment is determined based on the simulation environment configuration file.
[0023] Specifically, after the simulation configuration file is determined, a physical simulation environment can be constructed based on it. For example, the scene layout and robot-related information set in the simulation configuration file can be used to load relevant simulation objects and set their physical properties, thus completing the construction of the physical simulation environment. Within the physical simulation environment, cross-track replacement can be performed on various pose trajectories in the pose trajectory set, and a target trajectory set can be generated based on these replaced pose trajectories. Cross-track replacement can involve selecting semantically identical keyframes from various pose trajectories, performing cross-track replacement, and exchanging and recombining them between different trajectories to generate the target trajectory set.
[0024] S150. Execute each target trajectory in the target trajectory set in the physical simulation environment to generate training data for the target processing model based on the execution results of each target trajectory; the target processing model is used to process tasks in the task scenario.
[0025] It is known that after obtaining the set of target trajectories, each target trajectory in the set can be run in the constructed physical simulation environment. Based on the execution results of each target trajectory, the trajectory state sequence and environmental observation information during the execution process are converted into a standardized format to form training data that can be used to train the target processing model.
[0026] Among them, the target processing model can be used to process tasks in the task scenario. The target processing model can be deployed in the robot. By learning the correspondence between vision, language and action in the task scenario, the target processing model can gain the ability to perceive the environment, understand instructions and execute operations in real task scenarios.
[0027] This embodiment provides a method for determining model training data, including: acquiring multimodal input data of a task scenario; the multimodal input data includes historical operation videos, historical operation action trajectories, and natural language task descriptions; parsing the natural language task descriptions to determine a simulation environment configuration file corresponding to the task scenario; the simulation environment configuration file includes simulation objects for each target object in the task scenario; determining a set of pose trajectories based on the historical operation videos and historical operation action trajectories; the set of pose trajectories includes pose trajectories corresponding to each target object in the task scenario and associated trajectory annotation information; generating a set of target trajectories based on the set of pose trajectories in a physical simulation environment; the physical simulation environment is determined based on the simulation environment configuration file; executing each target trajectory in the set of target trajectories in the physical simulation environment to generate training data for a target processing model based on the execution results of each target trajectory; the target processing model is used to process tasks in the task scenario. The above technical solution achieves automated conversion from real-world demonstrations to simulation training data. Compared with traditional methods of manual data collection or manual design of simulation tasks, this technical solution significantly reduces the cost of data digitization through automated parsing and reconstruction. Furthermore, by using a programmable action library and in-simulation intelligent synthesis technology, it fundamentally breaks through the limitations of the original data scale, enabling the generation of diverse and physically reasonable training data. This greatly improves the training efficiency and model performance of data-driven robot learning methods.
[0028] As an optional implementation of this embodiment, parsing the natural language task description to determine the simulation environment configuration file corresponding to the task scenario further includes: 1) Parse the natural language task description to determine the initial parsing result corresponding to the task scenario; the initial parsing result includes at least one target object.
[0029] In this embodiment, the natural language task description can be parsed to determine the initial parsing result corresponding to the task scenario. The initial parsing result may include at least one target object. The target object refers to the specific physical object in the task scenario that needs to be manipulated or interacted with; it is the object directly acted upon during task execution.
[0030] It should be explained that the initial parsing results may also include task objectives and functional constraint information. The task objective can describe the final state or result that the task is expected to achieve, while the functional constraint information specifies the restrictions and normative requirements that must be followed during task execution, reflecting the rules and boundaries of task execution.
[0031] 2) Determine the spatial relationships between the various target objects.
[0032] Specifically, spatial reasoning chains can be used to determine the spatial relationships between various target objects. A spatial reasoning chain is a computational method that simulates the human spatial cognition process. It constructs a complete understanding of a spatial scene from fragmented information by linking multiple reasoning steps. The core of this method lies in chain-like derivation. Spatial reasoning chains typically integrate multiple knowledge sources, breaking down complex spatial problems into a series of interdependent sub-reasonings. The output of one step becomes the input of the next, forming a coherent logical chain.
[0033] As described above, spatial positional relationships can include the relative positional relationships and relative attitude relationships between various target objects.
[0034] 3) Determine the analysis result based on the initial analysis result and the spatial positional relationship.
[0035] Specifically, the initial parsing results and spatial relationships can be fused to determine the fused parsing results. The parsing results can be a unified scene representation that includes both semantic task understanding and geometric spatial layout.
[0036] 4) Based on the analysis results, determine the simulation object corresponding to each target physical object from the preset asset database.
[0037] It is known that after determining the analysis results, the retrieval enhancement generation mechanism can be used to match the simulation object corresponding to each target physical object from the preset asset database, and automatically complete its physical and interactive attributes to ensure that the assets can be used for physical simulation and interactive training.
[0038] The preset asset database can be a pre-built structured digital resource repository specifically used to store and manage various virtual objects and their related data required in the simulation environment. The preset asset database can store pre-modeled 3D models, physical parameters, and material definitions. Each asset has corresponding attribute tags such as category, size range, and functional type. Based on the semantic category, functional requirements, and spatial size constraints of the target object in the analysis results, the most matching simulation object can be retrieved from the preset asset database. It can be used to describe simulation objects using Mesh or URDF. Mesh can be a 3D mesh model of an object, describing its geometric shape (e.g., .obj, .stl, .dae format); URDF (Unified Robot Description Format) can be a unified robot description format used to define the object's physical properties (mass, inertia, colliders), joint structure, materials, and other simulation-required information.
[0039] As described above, the retrieval-enhanced generation mechanism is a technical paradigm that combines information retrieval with content generation. The core idea is to retrieve relevant reference information from an external knowledge base before generating answers or content, and then input these retrieval results as context into the generation model, thereby improving the accuracy, timeliness, and interpretability of the output.
[0040] It's important to explain that when the pre-defined asset database lacks specific objects or has insufficient diversity, the 3D generation module can automatically complete the entire process from description to modeling, refinement, and attribute completion, continuously expanding the coverage of the asset database. The 3D generation module can receive natural language descriptions or conceptual definitions of the target object and directly generate a basic 3D model using text-to-3D generation technology. This model may originate from diffusion models, neural radiation fields, or other methods, capturing key geometric features in the description, such as overall outline, main structure, and stylistic characteristics. The modeling stage typically outputs a coarse or low-resolution initial model, thus requiring a refinement stage. This stage uses geometric processing algorithms to optimize mesh topology, refine surface details, and correct topological errors, bringing the model to a quality standard usable in a simulation environment. This process may involve a combination of traditional graphics techniques and deep learning methods. The final attribute completion stage endows the newly generated model with complete simulation attributes, including automatically estimating physical parameters such as mass and inertia tensors, inferring material properties such as roughness and reflectivity, and annotating functional semantics such as graspable areas and interaction methods. These attributes may be based on visual feature predictions or obtained from similar assets.
[0041] 5) Generate the simulation environment configuration file based on each simulation object and the analysis results.
[0042] It is known that the simulation objects and the analysis results can be defined as simulation environment configuration files. These configuration files can also include instantiation information of the simulation objects, their physical properties, and the hierarchical or constraint relationships between them. Furthermore, task-related metadata can be embedded in the configuration files, such as the state conditions of the task objectives, the detection rules for functional constraints, and the observation point settings for evaluation. The generation process must adhere to the format specifications of the specific simulation platform to ensure that the file can be correctly parsed and loaded. The final output simulation environment configuration file is a self-contained scene definition. Loading it into the physical simulator allows for the reproduction of the complete task scene described by the analysis results, providing a runtime environment for subsequent trajectory execution and data generation.
[0043] As described above, once the simulation environment configuration file is determined, the environment layout, lighting settings, camera configuration, and rapid rendering verification can be completed based on the simulation environment configuration file to ensure the structural rationality and physical consistency of the generated scene.
[0044] As an optional implementation of this embodiment, the step of determining the pose trajectory set based on the historical operation video and the historical operation action trajectory includes: 1) Based on the historical operation video and the historical operation action trajectory, determine the initial pose trajectory corresponding to each target object.
[0045] Specifically, advanced vision-based models can be used to analyze historical operation videos and historical operation motion trajectories, thereby determining the initial pose trajectory for each target object in the task scene. The initial pose trajectory can be a sequence of poses of the corresponding target object over time, typically represented as a time series, where each moment can include position coordinates and the target object's attitude angle. For example, the initial pose sequence can be a 6D pose sequence.
[0046] Advanced vision foundation models refer to deep learning models that are pre-trained on large-scale datasets, possess powerful general visual understanding capabilities, and can adapt to various downstream tasks. For example, advanced vision foundation models can be models such as Stereo Anything, SAM3, and FoudationPose. These advanced vision foundation models can be used to predict 6D pose (position and orientation) of video streams frame by frame, thereby achieving high-precision and robust tracking of target objects.
[0047] 2) Perform keyframe annotation on each of the initial pose trajectories to determine the trajectory annotation information corresponding to each target object.
[0048] It is known that after obtaining the trajectories at each initial position, a preset algorithm can be used to annotate each initial pose trajectory with keyframes. Each keyframe can be a significant event point in the corresponding task execution process, and each keyframe can include a semantic label and a timestamp to form structured trajectory annotation information. The preset annotation algorithm can be a pre-defined algorithm; for example, the preset algorithm can be a Vision-Language Model (VLM), which can be a deep learning model that simultaneously processes image (or video) and text information.
[0049] 3) Based on the trajectory annotation information, determine the logical dependencies between the various action states of each target object in the initial pose trajectory.
[0050] Specifically, after determining the annotation information for each trajectory, analysis of this information can identify which actions must be completed before others, which actions can be executed in parallel, and which actions have resource contention or conditional triggering relationships. This allows for the determination of the logical dependencies between the various action states of each target object within the initial pose trajectory. For example, VLM can be used to annotate the initial pose trajectory with keyframes. Users can review and correct the automatically parsed results and define the logical dependencies between action states using connections. These logical dependencies provide a temporal framework for action execution, defining which action states have sequential or conditional triggering relationships. Determining these logical dependencies connects discrete annotation information into a network-like task structure, revealing the correct temporal constraints and parallel possibilities of task execution. This provides a compliance judgment basis for subsequent trajectory optimization, ensuring that the optimized trajectory still logically constitutes a reasonable task execution plan.
[0051] 4) Based on the target constraints, the initial pose trajectories are optimized using a preset optimization algorithm to determine the pose trajectory corresponding to each target object; the target constraints are determined based on the logical dependencies and the trajectory annotation information.
[0052] In this embodiment, target constraints can be determined based on logical dependencies and trajectory annotation information. Logical dependencies can serve as hard constraints in the optimization process to ensure that the basic execution logic of the task is not disrupted during any trajectory adjustment. The trajectory annotation information is used to guide the optimized trajectory to reflect the semantic features of the annotations as much as possible at keyframes, while maintaining the smoothness of the pose trajectory during the transition phase.
[0053] Following the above description, the preset optimization algorithm can be a pre-defined algorithm for trajectory optimization. For example, the preset optimization algorithm could be a differential similarity algorithm, which is a method that measures similarity by comparing data change trends rather than absolute values. It can first convert the sequence into differences (changes) between adjacent time points, and then compare the similarity of these change patterns. Under target constraints, the preset optimization algorithm optimizes each initial pose trajectory. The optimization process can include adjusting the trajectory interpolation method, smoothing transition regions and intermediate states between keyframes, reducing noise in the original data, etc., thereby determining the pose trajectory corresponding to each target object. For example, the pose trajectory corresponding to each target object can be a trajectory in HDF5 format.
[0054] 5) Determine the pose trajectory set based on the trajectory annotation information and the pose trajectory corresponding to each target object.
[0055] It can be seen that a set of position trajectories can be obtained based on the trajectory annotation information and the pose trajectory corresponding to each target object.
[0056] As an optional implementation of this embodiment, the step of generating a target trajectory set based on the pose trajectory set in a physical simulation environment includes: 1) In the physical simulation environment, based on the set of pose trajectories, the pose trajectories of each target object are combined across trajectories to determine each candidate pose trajectory.
[0057] It is known that the pose trajectory set includes the pose trajectory corresponding to each target object in the task scene, as well as the associated trajectory annotation information. In the physical simulation environment, target keyframes can be determined from each pose trajectory in the pose trajectory set based on the trajectory annotation information. Here, the target keyframes can be semantically identical keyframes in each pose trajectory.
[0058] Following the above description, after determining the target keyframes in each pose trajectory, the semantic similarity of each target keyframe can be calculated. Then, semantically identical target keyframes can be selected from different pose trajectories for cross-trajectory combination, thereby determining each candidate pose trajectory. For example, cross-trajectory combination may include randomly selecting semantically identical keyframes from different trajectories for replacement, and adjusting the keyframes in the stitched trajectory according to the actual pose of the interactive object in the current simulation environment. For instance, the trajectory point positions corresponding to the grasping keyframes can be registered at the grasping point of the interactive object, thus determining each candidate pose trajectory.
[0059] 2) Optimize each candidate pose trajectory using an optimization algorithm to determine the target trajectory set.
[0060] In this embodiment, the optimization algorithm can be an algorithm for optimizing pose trajectories. For example, the optimization algorithm can be a trajectory residual optimization network trained using reinforcement learning. Specifically, the optimization algorithm can be used to calibrate and optimize each candidate pose trajectory to determine the target trajectory set.
[0061] For example, based on the adjusted trajectory combination and physical simulation environment, a trajectory residual optimization network trained by reinforcement learning is used to perform dynamic calibration and optimization on the trajectory, such as compensating for trajectory deviations under physical constraints, eliminating trajectory inoperability problems caused by dynamic factors such as collisions, inertia, and friction, and then outputting a set of calibrated and optimized target trajectories. Each target trajectory in the target trajectory set satisfies physical consistency and can be directly executed in the simulation environment.
[0062] As an optional implementation of this embodiment, the step of executing each target trajectory in the target trajectory set in the physical simulation environment to generate training data for the target processing model based on the execution results of each target trajectory includes: 1) For any target trajectory in the set of target trajectories, execute the target trajectory in the physical simulation environment to determine the execution result of the target trajectory.
[0063] It is known that, for any target trajectory in the set of target trajectories, the virtual robot is driven to execute the target trajectory in the physical simulation environment in order to determine the execution result of the target trajectory.
[0064] 2) If the execution result is successful, then the initial training data is generated based on the trajectory execution data and environmental observation data during the execution process.
[0065] Specifically, if the execution result of the current target trajectory is successful, the trajectory execution data and environmental observation data during the execution process of the current target trajectory can be obtained, and these data can be used as initial training data. The trajectory execution data can include control command sequences, virtual robot joint states, end effector poses, contact force information, etc., reflecting the details of the action execution process. The environmental observation data can be used to represent the environmental state perceived by the virtual robot, such as simulated rendered images, depth maps, point clouds, or ground truth object poses.
[0066] 3) Convert the format of each of the initial training data to generate training data for the target processing model.
[0067] It is known that after obtaining the initial training data, the format of the initial training data can be converted and the information organized, thereby generating training data for the target processing model. For example, the training data can be in HDF5 format.
[0068] As an optional implementation of this embodiment, the method for determining model training data provided in this embodiment further includes: If the execution result is an execution failure, the target trajectory is replanned and executed again in the physical simulation environment until the execution result of the replanned target trajectory is a successful execution.
[0069] It is known that if the execution result is a failure, the target trajectory can be replanned. For example, the cause of the failure can be analyzed, the failure location can be determined, and local optimization can be performed near the failure location, a smooth transition segment can be inserted, or a motion planning algorithm can be called to redetermine the target trajectory. For instance, the trajectory of the current segment can be replanned, returning to the trajectory point corresponding to the previous keyframe, and execution can continue according to the newly planned trajectory.
[0070] As described above, after replanning the target trajectory, the replanned target trajectory can be executed in the physical simulation environment. If it still fails, iterative optimization continues until the execution result of the replanned target trajectory is successful, or the preset maximum number of retries is reached.
[0071] Figure 2 This is a schematic diagram illustrating the execution process of a method for determining model training data provided in this embodiment, as shown below. Figure 2As shown, the overall execution process is divided into two stages. Stage 1 is Real2Sim's automated parsing (from the real world to the simulation environment). Stage 1 mainly includes two steps: scene reconstruction and trajectory processing. Scene reconstruction mainly involves extracting task-related target objects from real-world input data, then reproducing and reasonably expanding the physical simulation environment based on the real layout. Trajectory processing is responsible for optimizing pose trajectories, including extracting the pose trajectories corresponding to the target objects, annotating trajectory keyframes, optimizing pose trajectories, and determining the pose trajectory set. The pose trajectory set includes the pose trajectory corresponding to each target object in the task scene and the associated trajectory annotation information. Through the two steps of scene reconstruction and trajectory processing in Stage 1, the simulation environment configuration file and the pose trajectory set can be obtained. Phase Two allows for Sim expansion, specifically including: Step 1: Initial Simulation Environment Construction: A physical simulation environment can be constructed using a simulation environment configuration file; Step 2: Keyframe Combination and Replacement: Keyframes for each pose trajectory in the pose trajectory set are determined, and cross-trajectory keyframe replacement is performed to obtain the replaced pose trajectory; Step 3: Determining Candidate Pose Trajectories: Based on the actual pose trajectories of interactive objects in the current simulation environment, keyframes in the stitched trajectory are adjusted, such as registering the trajectory point positions corresponding to the grasping keyframes to the grasping points of the interactive object, thereby determining each candidate pose trajectory; Step 4: Dynamic Trajectory Calibration: Under the action of the physics engine, using... The trajectory residual optimization network trained by RL further calibrates and optimizes each candidate pose trajectory, thereby determining the target trajectory set; Step 5: Actual execution and data collection: For any target trajectory in the target trajectory set, each target trajectory is actually executed in the physical simulation environment, and the success or failure of the execution is detected segment by segment. If successful, the next segment of trajectory is executed until the entire task is completed and the data is collected; if it fails, the trajectory of the current segment is replanned, returning to the trajectory point corresponding to the previous keyframe, and the execution continues according to the newly planned trajectory; Step 6: Training data determination: The trajectory execution data and environmental observation data are converted into formats to determine the training data.
[0072] The technical solution provided in this embodiment constructs an automated Real2Sim data pipeline: First, computer vision technology is used to automatically parse the 3D model of the interactive object and its six-degree-of-freedom motion trajectory from real teleoperation videos; then, with the help of a visual annotation tool, users can review the parsing results and transform the original demonstration into a structured programmable "action library" by marking keyframes and defining action logic dependencies; finally, in the simulation environment, the system automatically randomizes and expands the scene based on the action library, and intelligently splices and synthesizes a large number of novel training trajectories optimized by the physics engine to ensure dynamic feasibility, thereby generating high-quality visual-language-action data (training data) in batches, realizing the automated and intelligent production of massive simulation data from sparse real demonstrations. Compared with traditional methods of manual data collection or manual design of simulation tasks, this invention significantly reduces the cost of data digitization through automated parsing and reconstruction, and fundamentally breaks through the limitations of the original data scale through a programmable action library and intelligent synthesis technology within the simulation, enabling the generation of diverse and physically reasonable training data, greatly improving the training efficiency and model performance of data-driven robot learning methods.
[0073] Example 2 Figure 3 This is a schematic diagram of the structure of a device for determining model training data provided in Embodiment 2 of this disclosure; as shown Figure 3 As shown, the device includes: a data acquisition module 210, a parsing module 220, a pose trajectory set determination module 230, a target trajectory set determination module 240, and a training data determination module 250.
[0074] The data acquisition module 210 is used to acquire multimodal input data of the task scenario; the multimodal input data includes historical operation videos, historical operation action trajectories, and natural language task descriptions. The parsing module 220 is used to parse the natural language task description to determine the simulation environment configuration file corresponding to the task scenario; the simulation environment configuration file includes the simulation object of each target object in the task scenario; The pose trajectory set determination module 230 is used to determine a pose trajectory set based on the historical operation video and the historical operation action trajectory; the pose trajectory set includes the pose trajectory corresponding to each target object in the task scene and the associated trajectory annotation information; The target trajectory set determination module 240 is used to generate a target trajectory set based on the pose trajectory set in a physical simulation environment; the physical simulation environment is determined based on the simulation environment configuration file. The training data determination module 250 is used to execute each target trajectory in the target trajectory set in the physical simulation environment to generate training data for the target processing model based on the execution results of each target trajectory; the target processing model is used to process tasks in the task scenario.
[0075] Embodiment 2 of this disclosure provides a device for determining model training data, which realizes the automatic conversion of easily obtainable real-world teleoperation videos into large-scale training data that is programmable, scalable, and physically reasonable within a simulation environment.
[0076] Furthermore, the parsing module 220 is also used for: The natural language task description is parsed to determine the initial parsing result corresponding to the task scenario; the initial parsing result includes at least one target object. Determine the spatial relationships between the various target objects; The analysis result is determined based on the initial analysis result and the spatial positional relationship; Based on the analysis results, the simulation object corresponding to each target physical object is determined from the preset asset database; The simulation environment configuration file is generated based on each of the simulation objects and the analysis results.
[0077] Furthermore, the pose trajectory set determination module 230 is also used for: Based on the historical operation video and the historical operation action trajectory, determine the initial pose trajectory corresponding to each target object; Keyframe annotations are performed on each of the initial pose trajectories to determine the trajectory annotation information corresponding to each target object; Based on the trajectory annotation information, the logical dependencies between the various action states of each target object in the initial pose trajectory are determined. Based on the target constraints, a preset optimization algorithm is used to optimize each of the initial pose trajectories to determine the pose trajectory corresponding to each target object; the target constraints are determined based on the logical dependencies and the trajectory annotation information. The pose trajectory set is determined based on the trajectory annotation information and the pose trajectory corresponding to each target object.
[0078] Furthermore, the target trajectory set determination module 240 is also used for: In the physical simulation environment, based on the pose trajectory set, the pose trajectories of each target object are combined across trajectories to determine each candidate pose trajectory; The target trajectory set is determined by optimizing each candidate pose trajectory using an optimization algorithm.
[0079] Furthermore, the training data determination module is used for: For any target trajectory in the set of target trajectories, the target trajectory is executed in the physical simulation environment to determine the execution result of the target trajectory; If the execution result is successful, initial training data is generated based on the trajectory execution data and environmental observation data during the execution process; The initial training data are converted into different formats to generate training data for the target processing model.
[0080] Furthermore, the device also includes: The planning module is used to replan the target trajectory if the execution result is failure, and then execute the replanned target trajectory again in the physical simulation environment until the execution result of the replanned target trajectory is success.
[0081] The device for determining model training data provided in this disclosure can execute the method for determining model training data provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.
[0082] Example 3 Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the embodiments of the present disclosure described and / or claimed herein.
[0083] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0084] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0085] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microprocessor, etc. Processor 11 performs the various methods and processes described above, such as methods for determining model training data.
[0086] In some embodiments, the method for determining model training data may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method for determining model training data described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the method for determining model training data by any other suitable means (e.g., by means of firmware).
[0087] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0088] Computer programs for implementing the methods of embodiments of this disclosure may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0089] In the context of embodiments of this disclosure, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0090] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0091] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0092] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0093] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the embodiments of this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of the embodiments of this disclosure can be achieved, and this document does not impose any limitations.
[0094] The specific embodiments described above do not constitute a limitation on the scope of protection of the embodiments disclosed herein. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the embodiments disclosed herein should be included within the scope of protection of the embodiments disclosed herein.
[0095] This disclosure also provides a computer program product, including a computer program and / or instructions, which, when executed by a processor, implements the method for determining model training data as provided in any embodiment of this application.
[0096] In implementing a computer program product, computer program code for performing the operations of the embodiments of this disclosure can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0097] Note that the above are merely preferred embodiments and the technical principles applied in this disclosure. Those skilled in the art will understand that this disclosure is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the protection scope of this disclosure. Therefore, although the embodiments of this disclosure have been described in detail above, this disclosure is not limited to the above embodiments. More other equivalent embodiments may be included without departing from the concept of this disclosure, and the scope of this disclosure is determined by the scope of the appended claims.
Claims
1. A method for determining model training data, characterized in that, include: Acquire multimodal input data for the task scenario; The multimodal input data includes historical operation videos, historical operation action trajectories, and natural language task descriptions; The natural language task description is parsed to determine the simulation environment configuration file corresponding to the task scenario; the simulation environment configuration file includes the simulation object of each target object in the task scenario. A set of pose trajectories is determined based on the historical operation video and the historical operation action trajectory; the set of pose trajectories includes the pose trajectory corresponding to each target object in the task scene and the associated trajectory annotation information; In a physical simulation environment, a target trajectory set is generated based on the pose trajectory set; the physical simulation environment is determined based on the simulation environment configuration file. Each target trajectory in the target trajectory set is executed in the physical simulation environment to generate training data for the target processing model based on the execution results of each target trajectory; the target processing model is used to process tasks in the task scenario.
2. The method according to claim 1, characterized in that, The step of parsing the natural language task description to determine the simulation environment configuration file corresponding to the task scenario includes: The natural language task description is parsed to determine the initial parsing result corresponding to the task scenario; the initial parsing result includes at least one target object. Determine the spatial relationships between the various target objects; The analysis result is determined based on the initial analysis result and the spatial positional relationship; Based on the analysis results, the simulation object corresponding to each target physical object is determined from the preset asset database; The simulation environment configuration file is generated based on each of the simulation objects and the analysis results.
3. The method according to claim 1, characterized in that, The determination of the pose trajectory set based on the historical operation video and the historical operation action trajectory includes: Based on the historical operation video and the historical operation action trajectory, determine the initial pose trajectory corresponding to each target object; Keyframe annotations are performed on each of the initial pose trajectories to determine the trajectory annotation information corresponding to each target object; Based on the trajectory annotation information, the logical dependencies between the various action states of each target object in the initial pose trajectory are determined. Based on the target constraints, a preset optimization algorithm is used to optimize each of the initial pose trajectories to determine the pose trajectory corresponding to each target object; the target constraints are determined based on the logical dependencies and the trajectory annotation information. The pose trajectory set is determined based on the trajectory annotation information and the pose trajectory corresponding to each target object.
4. The method according to claim 1, characterized in that, In the physical simulation environment, generating a target trajectory set based on the pose trajectory set includes: In the physical simulation environment, based on the pose trajectory set, the pose trajectories of each target object are combined across trajectories to determine each candidate pose trajectory; The target trajectory set is determined by optimizing each candidate pose trajectory using an optimization algorithm.
5. The method according to claim 1, characterized in that, The step of executing each target trajectory in the target trajectory set in the physical simulation environment to generate training data for the target processing model based on the execution results of each target trajectory includes: For any target trajectory in the set of target trajectories, the target trajectory is executed in the physical simulation environment to determine the execution result of the target trajectory; If the execution result is successful, initial training data is generated based on the trajectory execution data and environmental observation data during the execution process; The initial training data are converted into different formats to generate training data for the target processing model.
6. The method according to claim 5, characterized in that, The method further includes: If the execution result is an execution failure, the target trajectory is replanned and executed again in the physical simulation environment until the execution result of the replanned target trajectory is a successful execution.
7. A device for determining model training data, characterized in that, include: The data acquisition module is used to acquire multimodal input data for the task scenario; The multimodal input data includes historical operation videos, historical operation action trajectories, and natural language task descriptions; The parsing module is used to parse the natural language task description to determine the simulation environment configuration file corresponding to the task scenario; the simulation environment configuration file includes the simulation object of each target object in the task scenario; The pose trajectory set determination module is used to determine a pose trajectory set based on the historical operation video and the historical operation action trajectory; the pose trajectory set includes the pose trajectory corresponding to each target object in the task scene and the associated trajectory annotation information; The target trajectory set determination module is used to generate a target trajectory set based on the pose trajectory set in a physical simulation environment; the physical simulation environment is determined based on the simulation environment configuration file. The training data determination module is used to execute each target trajectory in the target trajectory set in the physical simulation environment to generate training data for the target processing model based on the execution results of each target trajectory. The target processing model is used to process the tasks in the task scenario.
8. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the method for determining model training data as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the method for determining model training data as described in any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method for determining model training data as described in any one of claims 1-6.