Robot operation training data generation method based on dual simulation engine and medium
Patent Information
- Application Number
- CN202611301350.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-26
- Publication Date
- 2026-09-22
AI Technical Summary
[0016]本发明的有益效果:通过响应于接收的操作任务,控制基于域随机化的第一仿真引擎生成包含场景生成与运动轨迹规划的第一仿真操作数据及起始场景数据,并利用资产处理模块将起始场景数据解耦拆分为多个独立物体的资产数据,进而将这些资产数据输入至负责渲染及视觉语义标注的初始第二仿真引擎中,结合起始场景数据恢复各独立物体在世界坐标系下的初始化位姿及存在状态以构建场景初始化后的第二仿真引擎,随后将第一仿真操作数据输入该中间引擎进行基于关节映射的运动重放以获取第二仿真操作数据,最后对两路仿真操作数据进行采用相同数据结构定义的同构化处理,生成字段名称和语义一致的第一训练数据和第二训练数据并汇入训练数据资源池,使得来自不同优势仿真引擎的数据能够在保持任务回合级状态一致性的前提下,被转换为结构同构的统一格式,从而有效解决了在现有的跨仿真引擎数据迁移与联合训练应用中,面临的因无法精确复现初始随机状态和物体存在情况,以及多源数据异构导致难以有效整合用于同一模型联合训练的技术问题,因此避免了因仿真环境差异导致的数据对齐失败和模型训练泛化能力受限的情况,提升了多源仿真数据在视觉语言动作模型训练中的融合效率与应用价值。
Smart Images

Figure CN122797359A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot simulation technology, and in particular to a method and medium for generating robot operation training data based on a dual simulation engine. Background Technology
[0002] With the rapid development of embodied intelligence technology, the demand for large-scale, high-quality multimodal data for training robot visual language action models is increasing. This type of training data typically needs to include rich image information, accurate semantic annotations, detailed robot state records, and diverse language commands. Because collecting such data in the real physical world presents challenges such as high cost, low efficiency, and limited scene coverage, generating synthetic training data using simulation environments has become the mainstream technical solution in the industry.
[0003] In current technical practices, a single simulation engine is typically used to complete the entire process from scene construction, domain randomization, trajectory planning to data recording. However, there are obvious technical trade-offs in the practical application of a single simulation engine: one type of simulation engine focuses more on efficient scene construction and large-scale domain randomization, which can quickly generate a large number of randomized task rounds, but is relatively limited in rendering realism and pixel-level semantic annotation accuracy; the other type of simulation engine focuses more on high-fidelity rendering capabilities and accurate annotation tools, but its scene randomization capabilities and task batch generation efficiency are relatively low, making it difficult to independently support the output needs of large-scale training data.
[0004] It is evident that existing single simulation engine solutions cannot simultaneously achieve high efficiency in task generation, flexibility in scene randomization, and high fidelity in data output, thus hindering the improvement of robot training data quality and large-scale production. Summary of the Invention
[0005] Therefore, it is necessary to address the existing problem of generating robot operation training data based on dual simulation engines by proposing a method and medium for generating robot operation training data based on dual simulation engines.
[0006] A method for generating robot operation training data based on a dual simulation engine, the method comprising: In response to the received operation task, the first simulation engine is controlled to perform simulation operation and acquire the first simulation operation data and the initial scene data; wherein, the first simulation engine is configured to perform scene generation and motion trajectory planning for robot operation tasks based on domain randomization. The initial scene data is split into asset data of multiple independent objects through a preset asset processing module; wherein, the preset asset processing module is an asset format conversion unit located between the two engines, used to decompose the initial scene data generated by the first engine into asset data that can be loaded independently by the second engine. The asset data of multiple independent objects are input into the initial second simulation engine to obtain the initial pose of each independent object in the world coordinate system and the existence state of each independent object, thereby completing the scene initialization of the second simulation engine; wherein, the initial second simulation engine is configured to load preset target robot assets and perform rendering and visual semantic annotation. The first simulation operation data is input into the second simulation engine after the scene is initialized to perform simulation operations, and the second simulation operation data is obtained. The first simulation operation data and the second simulation operation data are homogenized to obtain the first training data and the second training data, respectively. The first and second training data are input into a preset training data resource pool for training the target robot.
[0007] Further, the step of controlling the first simulation engine to perform simulation operations in response to the received operation task, and acquiring the first simulation operation data and the initial scene data, includes: In response to the received operation task, the first simulation engine is controlled to perform domain randomization, and a scene initialization snapshot is collected after domain randomization and before the execution of the first frame simulation action to obtain the initial scene data; wherein, the scene initialization snapshot includes at least the robot initial pose, the initial pose of each object, and the existence status identifier of each object. After randomization in the first simulation engine domain, simulation operations are performed, and data from the simulation operation process is collected to obtain the first simulation operation data.
[0008] Further, the step of performing simulation operations after randomization in the first simulation engine domain and collecting data from the simulation operation process to obtain the first simulation operation data includes: After domain randomization, the robot model in the first simulation engine is controlled to perform simulated motions corresponding to the operation task; The robot joint state values of each time frame during the simulation motion are recorded in sequence to obtain the robot joint state trajectory. Record the original motion command values of each time frame to obtain the original motion trajectory, and record the joint name sequence corresponding to each dimension in the robot joint state trajectory; The robot joint state trajectory, the original motion trajectory, and the joint name sequence are used as the first simulation operation data.
[0009] Furthermore, the step of splitting the initial scene data into asset data of multiple independent objects through a preset asset processing module includes: The asset processing module reads the complete static scene asset corresponding to the initial scene data and expands the complete static scene asset into a scene hierarchy structure. Traverse each object node in the scene hierarchy and export each object node as an independent object asset file. The material files that each object asset file depends on are copied, and the reference paths within each object asset file pointing to the material files and texture maps are rewritten to make each object asset file independent of each other, thus obtaining asset data for multiple independent objects.
[0010] Furthermore, the step of inputting asset data of multiple independent objects into the initial second simulation engine to obtain the initial pose of each independent object in the world coordinate system and the existence state of each independent object, thereby completing the scene initialization of the second simulation engine, includes: The asset data of each of the independent objects is loaded one by one through the initial second simulation engine; Read the robot initial pose, the initial pose of each object, and the existence status identifier of each object contained in the initial scene data; Place the preset target robot asset in the second simulation engine scene according to the robot's initial pose; Based on the initial pose of each object, the corresponding object is placed in the scene of the second simulation engine, and the corresponding object is loaded or hidden according to the existence status flag of each object, thereby completing the scene initialization of the second simulation engine.
[0011] Further, the step of inputting the first simulation operation data into the second simulation engine after scene initialization to perform simulation operations and obtain the second simulation operation data includes: Read the joint name sequence in the first simulation operation data, and obtain the runtime joint names of the preset target robot assets in the second simulation engine after scene initialization; Each joint name in the joint name sequence is matched one by one with the runtime joint name to establish a joint mapping relationship; Based on the joint mapping relationship, the joint state values of each time frame in the joint state trajectory of the first simulation operation data are converted frame by frame into the executable joint target values of the second simulation engine after scene initialization. The joint target values of each time frame are sent out sequentially according to the time frame order, driving the target robot asset in the second simulation engine after scene initialization to perform motion replay; During the motion replay process, multimodal data including images, semantic annotations, robot states, actions, and language commands are recorded synchronously to obtain the second simulation operation data.
[0012] Further, the step of isomorphizing the first simulation operation data and the second simulation operation data to obtain the first training data and the second training data respectively includes: The first simulation operation data is input into the first transformation function to process the first simulation operation data into first training data containing image features, state features, action features and language features; The second simulation operation data is input into the second transformation function to process the second simulation operation data into second training data containing image features, state features, action features and language features; The first conversion function and the second conversion function use the same data structure definition, and the output field names and field semantics correspond one-to-one.
[0013] A robot operation training data generation device based on a dual simulation engine, the device comprising: The control module is used to respond to the received operation task, control the first simulation engine to perform simulation operation, and acquire the first simulation operation data and the initial scene data; wherein, the first simulation engine is configured to perform scene generation and motion trajectory planning for robot operation tasks based on domain randomization. The splitting module is used to split the initial scene data into asset data of multiple independent objects through a preset asset processing module; wherein, the preset asset processing module is an asset format conversion unit located between the two engines, used to decompose the initial scene data generated by the first engine into asset data that can be loaded independently by the second engine; The first input module is used to input asset data of multiple independent objects into the initial second simulation engine to obtain the initial pose of each independent object in the world coordinate system and the existence state of each independent object, thereby completing the scene initialization of the second simulation engine; wherein, the initial second simulation engine is configured to load preset target robot assets and perform rendering and visual semantic annotation. The simulation module is used to input the first simulation operation data into the second simulation engine after the scene is initialized to perform simulation operations and obtain the second simulation operation data. The processing module is used to perform isomorphic processing on the first simulation operation data and the second simulation operation data to obtain the first training data and the second training data, respectively. The second input module is used to input the first training data and the second training data into a preset training data resource pool for training the target robot.
[0014] An electronic device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the following steps: In response to the received operation task, the first simulation engine is controlled to perform simulation operation and acquire the first simulation operation data and the initial scene data; wherein, the first simulation engine is configured to perform scene generation and motion trajectory planning for robot operation tasks based on domain randomization. The initial scene data is split into asset data of multiple independent objects through a preset asset processing module; wherein, the preset asset processing module is an asset format conversion unit located between the two engines, used to decompose the initial scene data generated by the first engine into asset data that can be loaded independently by the second engine. The asset data of multiple independent objects are input into the initial second simulation engine to obtain the initial pose of each independent object in the world coordinate system and the existence state of each independent object, thereby completing the scene initialization of the second simulation engine; wherein, the initial second simulation engine is configured to load preset target robot assets and perform rendering and visual semantic annotation. The first simulation operation data is input into the second simulation engine after the scene is initialized to perform simulation operations, and the second simulation operation data is obtained. The first simulation operation data and the second simulation operation data are homogenized to obtain the first training data and the second training data, respectively. The first and second training data are input into a preset training data resource pool for training the target robot.
[0015] A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the following steps: In response to the received operation task, the first simulation engine is controlled to perform simulation operation and acquire the first simulation operation data and the initial scene data; wherein, the first simulation engine is configured to perform scene generation and motion trajectory planning for robot operation tasks based on domain randomization. The initial scene data is split into asset data of multiple independent objects through a preset asset processing module; wherein, the preset asset processing module is an asset format conversion unit located between the two engines, used to decompose the initial scene data generated by the first engine into asset data that can be loaded independently by the second engine. The asset data of multiple independent objects are input into the initial second simulation engine to obtain the initial pose of each independent object in the world coordinate system and the existence state of each independent object, thereby completing the scene initialization of the second simulation engine; wherein, the initial second simulation engine is configured to load preset target robot assets and perform rendering and visual semantic annotation. The first simulation operation data is input into the second simulation engine after the scene is initialized to perform simulation operations, and the second simulation operation data is obtained. The first simulation operation data and the second simulation operation data are homogenized to obtain the first training data and the second training data, respectively. The first and second training data are input into a preset training data resource pool for training the target robot.
[0016] The beneficial effects of this invention are as follows: In response to a received operational task, a first simulation engine based on domain randomization is controlled to generate first simulation operation data and initial scene data, including scene generation and motion trajectory planning. An asset processing module decouples and splits the initial scene data into asset data for multiple independent objects. This asset data is then input into an initial second simulation engine responsible for rendering and visual semantic annotation. The initial scene data is combined with the initial scene data to restore the initial pose and existence state of each independent object in the world coordinate system, thus constructing a second simulation engine after scene initialization. Subsequently, the first simulation operation data is input into this intermediate engine for joint-mapping-based motion replay to obtain the second simulation operation data. Finally, the two simulation operation data streams are processed using the same data... The isomorphism processing of the structure definition generates first and second training data with consistent field names and semantics, and merges them into the training data resource pool. This allows data from different simulation engines to be converted into a unified format with isomorphic structure while maintaining the consistency of task round-level states. This effectively solves the technical problems faced in existing cross-simulation engine data migration and joint training applications, such as the inability to accurately reproduce the initial random state and object existence, and the difficulty in effectively integrating multi-source data for joint training of the same model due to heterogeneity. Therefore, it avoids data alignment failure and limited model training generalization ability caused by differences in simulation environment, and improves the fusion efficiency and application value of multi-source simulation data in visual language action model training. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] in: Figure 1 This is an application environment diagram of a robot operation training data generation method based on a dual simulation engine in one embodiment; Figure 2 This is a flowchart of a robot operation training data generation method based on a dual simulation engine in one embodiment; Figure 3This is a structural block diagram of a robot operation training data generation device based on a dual simulation engine in one embodiment; Figure 4 This is a structural block diagram of an electronic device in one embodiment. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Figure 1 An application environment diagram is generated for robot operation training data based on a dual simulation engine in one embodiment. (Refer to...) Figure 1 This robot operation training data generation method based on a dual simulation engine is applied to a robot operation training data generation system based on a dual simulation engine. The system includes a terminal 110 and a server 120. The terminal 110 and server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal; the mobile terminal can be at least one of a mobile phone, tablet, or laptop. The server 120 can be a standalone server or a server cluster consisting of multiple servers.
[0021] like Figure 2 As shown, in one embodiment, a method for generating robot operation training data based on a dual simulation engine is provided. This method can be applied to both terminals and servers; this embodiment uses terminal application as an example. The specific steps of this method for generating robot operation training data based on a dual simulation engine are as follows: S1: In response to the received operation task, control the first simulation engine to perform simulation operation and acquire the first simulation operation data and the initial scene data; wherein, the first simulation engine is configured to perform scene generation and motion trajectory planning for robot operation tasks based on domain randomization; S2: The initial scene data is split into asset data of multiple independent objects through a preset asset processing module; wherein, the preset asset processing module is an asset format conversion unit located between the two engines, used to decompose the initial scene data generated by the first engine into asset data that can be loaded independently by the second engine; S3: Input the asset data of multiple independent objects into the initial second simulation engine to obtain the initial pose of each independent object in the world coordinate system and the existence state of each independent object, thereby completing the scene initialization of the second simulation engine; wherein, the initial second simulation engine is configured to load preset target robot assets and perform rendering and visual semantic annotation. S4: Input the first simulation operation data into the second simulation engine after scene initialization to perform simulation operations and obtain the second simulation operation data; S5: Perform isomorphic processing on the first simulation operation data and the second simulation operation data to obtain the first training data and the second training data, respectively. S6: Input the first training data and the second training data into a preset training data resource pool for training the target robot.
[0022] It should be noted that both the first and second simulation engines in this invention possess basic functions such as scene construction, physical simulation, motion control, and data recording. The core difference lies in their respective design focuses: the first simulation engine is deeply optimized for large-scale randomized tasks, while the second simulation engine has significant advantages in high-fidelity rendering and semantic annotation. This invention aims to fully utilize the strengths of both engines through a specific data transformation process.
[0023] As described in step S1 above, in response to the received operation task, the first simulation engine is controlled to perform simulation operations and acquire first simulation operation data and initial scene data. The first simulation engine is configured to perform scene generation and motion trajectory planning for robot operation tasks based on domain randomization. In this embodiment, the first simulation engine is used for task planning and trajectory generation, responsible for constructing robot operation scenes in a virtual environment and generating corresponding motion control commands. When a specific operation task is received, the first simulation engine immediately starts the simulation process. This first simulation engine is configured to perform scene generation based on domain randomization, that is, during the simulation initialization phase, parameters such as object positions, postures, textures, lighting conditions, and physical properties in the scene are randomly adjusted to increase the diversity and robustness of the training data. During the simulation process, the first simulation engine uses a built-in motion planning algorithm to calculate and generate the robot's motion trajectory according to the operation task objective, and collects key data during the simulation process in real time to form the first simulation operation data. This first simulation operation data may, for example, include the time series of robot joint states, the pose changes of the end effector, and the applied torque information; or it may include the joint state trajectory recording the angle values of each joint of the robot at each moment and the corresponding original motion command values. Furthermore, to ensure accurate reproduction of the current simulation scene in the subsequent second simulation engine, the system collects a scene initialization snapshot as initial scene data after domain randomization and before formally executing the first frame of simulation action. This initial scene data includes at least the robot's initial pose in the world coordinate system, the initial pose of each independent object in the scene, and state flags indicating the presence of each object. By acquiring the aforementioned first simulation operation data and initial scene data, the dynamic control information and static scene configuration information for this task round are completely saved.
[0024] As described in step S2 above, the initial scene data is split into asset data for multiple independent objects using a pre-defined asset processing module. To achieve cross-engine scene reproduction, the scene data generated by the first simulation engine needs to be converted into a format recognizable and loadable by the second simulation engine. This is executed by the pre-defined asset processing module, which is responsible for parsing the complete static scene assets corresponding to the initial scene data. This module is an asset format conversion unit located between the two engines. It decomposes the complete scene (i.e., the initial scene data) generated by the first engine into object asset packages (i.e., asset data for multiple independent objects) that can be independently loaded by the second engine. Through material rewriting and integrity verification, it ensures that the assets are correctly rendered and physically simulated in the second engine, thus achieving lossless migration and cross-task reuse of static assets between different simulation platforms. A complete static scene asset is typically a collection file containing scene hierarchy, object geometry models, material textures, and physical property information. The asset processing module first expands the complete static scene asset into a scene hierarchy and traverses each object node in this hierarchy. For each traversed object node, the asset processing module exports it as an independent object asset file. During the export process, to ensure that each object asset file can be loaded independently in the second simulation engine without missing dependencies, the asset processing module performs a dependency copy operation on each object asset file. This involves copying its dependent material files, texture files, etc., to the directory or specified resource path where the object asset file is located, and rewriting the reference paths of these dependent files within the object asset file. After this processing, each object asset file contains all the resources required for its own rendering and physical interaction, and there are no longer any external dependencies between them, thus forming asset data for multiple independent objects. For example, the asset processing module can export tables, cups, robotic arms, etc., in the scene as independent OBJ (Object File, a standard 3D model format) or URDF (Unified Robot Description Format) files, along with their respective MTL material files; or it can use the USD (Universal Scene Description) format to encapsulate each object as an independent USD primal and modify the material reference paths to relative paths. Through this separation, the static scene assets are decoupled, allowing the second simulation engine to load any subset of objects as needed. The originally tightly coupled overall scene is transformed into a modular asset set that can be independently distributed and flexibly reused, thereby improving the compatibility and processing efficiency of cross-simulation engine data migration.
[0025] As described in step S3 above, after receiving the asset data of the multiple independent objects, the initial second simulation engine searches for and associates the corresponding initial pose information from the initial scene data based on the unique identifier (such as name or ID) of each object, and then loads the asset data of each independent object one by one. By inputting the asset data of multiple independent objects into the initial second simulation engine, the initial pose of each independent object in the world coordinate system and the existence state of each independent object are obtained, thereby completing the scene initialization of the second simulation engine. The initial second simulation engine is configured to load preset target robot assets and perform rendering and visual semantic annotation. The initial second simulation engine can refer to a simulation engine instance that has not yet loaded the specific task scene state and has only completed the basic environment configuration. This initial second simulation engine is configured to load preset target robot assets, that is, a pre-prepared 3D model with the same kinematic and geometric parameters as the entity robot to be trained, and has the ability to perform high-fidelity rendering and visual semantic annotation. After receiving the generated asset data of multiple independent objects, the initial second simulation engine loads these asset data one by one. During loading, the system synchronously reads the initialization information contained in the initial scene data, specifically including the robot's initial pose, the initial poses of each object, and the existence status identifiers of each object. Based on the robot's initial pose, the system precisely places the preset target robot assets in the world coordinate system of the initial second simulation engine, setting their position and orientation. For each independent object, the system places it at a designated position in the scene according to its corresponding initial pose, and performs differentiated processing based on the existence status identifiers of each object: if the identifier indicates that an object exists in the current task round, the object is loaded and displayed normally; if the identifier indicates that the object does not exist, the system masks, hides, or disables its physical colliders. Through the above operations, the scene layout and object states in the second simulation engine are completely consistent with the results of domain randomization in the first simulation engine, thus obtaining the configured scene-initialized second simulation engine. This scene-initialized second simulation engine not only restores the initial scene but also has the ability to generate visual data based on real robot assets. For example, the scene-initialized second simulation engine can be a high-fidelity visual simulation environment built on Unity or Unreal Engine; or it can be a robot simulation platform that supports physical simulation and semantic segmentation, such as Isaac Sim. The preset target robot asset refers to a pre-built three-dimensional digital model file that has the same kinematic, dynamic parameters and geometric appearance as the final physical robot to be trained, and contains a complete joint hierarchy, colliders and visual mesh.
[0026] As described in step S4 above, the first simulation operation data is input into the second simulation engine after scene initialization for simulation operation, resulting in the second simulation operation data. After the second simulation engine is ready after scene initialization, the system inputs the first simulation operation data into it for simulation replay. The first simulation operation data contains the robot motion trajectory and control commands generated in the first simulation engine. Since the robot models used by the first and second simulation engines may have differences in joint definition order, number of joints, or degree of freedom configuration, directly replaying by data column index will lead to motion errors. Therefore, during the simulation operation, the system first extracts the joint name sequence carried in the first simulation operation data and obtains the runtime joint names of the preset target robot assets in the second simulation engine after scene initialization. By matching each joint name in the joint name sequence with the runtime joint names one by one, a precise joint mapping relationship is established. Based on this joint mapping relationship, the system converts the joint state values of each time frame in the joint state trajectory of the first simulation operation data into joint target values that can be executed by the second simulation engine after scene initialization. Subsequently, the joint target values of each time frame are sequentially distributed according to the time frame order, driving the target robot asset in the second simulation engine after scene initialization to perform motion replay. During motion replay, the second simulation engine after scene initialization utilizes its rendering pipeline and sensor model to synchronously record second simulation operation data containing multimodal information. This second simulation operation data may include, for example, high-resolution RGB images, depth images, instance segmentation masks, and semantic annotation maps, or it may include robot body state, joint torques, end effector speeds, and accompanying task language commands. Through this step, the trajectory generated by the first simulation engine is reproduced in the high-fidelity environment of the second simulation engine, generating training data with rich semantic annotations.
[0027] As described in step S5 above, the first simulation operation data and the second simulation operation data are homogenized to obtain the first training data and the second training data, respectively. Since the first and second simulation operation data originate from two different simulation engines, their original data formats, field names, and storage structures often differ, making direct merging for the same training pipeline impossible. To address this issue, homogenization is required to ensure consistency in data structure and semantics. This homogenization is achieved through a preset conversion function. Specifically, the system inputs the first simulation operation data into the first conversion function, processing it into first training data conforming to a preset standard format; simultaneously, it inputs the second simulation operation data into the second conversion function, processing it into second training data conforming to the same preset standard format. The first and second conversion functions use identical data structure definitions to ensure a one-to-one correspondence between output field names and field semantics. For example, both the first and second training data can contain fields such as image features, state features, action features, and language features, with image feature fields corresponding to visual observations at the same time in both datasets, and action feature fields corresponding to control commands within the same control cycle. By homogenizing the data, the data format barriers caused by heterogeneous simulation engines are eliminated. For example, the transformation function can uniformly convert coordinate system transformation data from different engines to the robot base coordinate system; or it can uniformly convert data types of different precision to 32-bit floating-point format. The resulting first and second training data have a unified interface specification. Eliminating the data format barriers caused by heterogeneous simulation engines allows subsequent training pipelines to directly read and process data without distinguishing the data source. In one specific implementation, the core logic of the first and second transformation functions is as follows: according to a preset field mapping table (for example, mapping joint_positions in the source data to the standard field state, and camera_rgb to the standard field image), the corresponding data is extracted from the source data, and dimension checks and filling are performed, finally assembling it into a tensor dictionary that conforms to the target data structure definition. This mapping table is pre-configured along with the data structure definition file, thereby ensuring strict consistency in the output format of the two transformation functions.
[0028] As described in step S6 above, the first training data and the second training data are input into a preset training data resource pool for training the target robot. Finally, the homogenized first and second training data are merged and stored in the preset training data resource pool. This training data resource pool provides data support for training machine learning models. Since the first and second training data have the same field structure and semantics, the same training pipeline can read these data from the resource pool and perform batch loading using the exact same parsing logic and input format. During training, the data in the training data resource pool is used to train the model corresponding to the target robot, such as jointly training a visual-language-action model. The visual-language-action model can simultaneously process image observation, language commands, and action output. By mixing data from the first simulation engine (emphasizing large-scale domain randomization and task diversity) and the second simulation engine (emphasizing high-fidelity visual rendering and realistic semantic annotation), the model can learn more generalizable feature representations. For example, the model can be a multimodal large model based on the Transformer architecture; or it can be a deep reinforcement learning strategy model based on a combination of convolutional neural networks and long short-term memory networks. This joint training effectively improves the model's transferability and robustness across different physical and visual domains.
[0029] In one embodiment, step S1, which responds to a received operation task by controlling the first simulation engine to perform simulation operations and acquiring first simulation operation data and initial scene data, includes: S101: In response to the received operation task, control the first simulation engine to perform domain randomization, and collect a scene initialization snapshot after domain randomization and before the execution of the first frame simulation action to obtain the initial scene data; wherein, the scene initialization snapshot includes at least the robot initial pose, the initial pose of each object, and the existence status identifier of each object. S102: After randomizing the first simulation engine domain, perform simulation operations and collect data from the simulation operation process to obtain the first simulation operation data.
[0030] As described in step S101 above, in response to the received operation task, the first simulation engine is controlled to perform domain randomization, and a scene initialization snapshot is collected after domain randomization and before the execution of the first frame of simulation action to obtain the initial scene data. The scene initialization snapshot includes at least the robot's initial pose, the initial poses of each object, and the existence status identifiers of each object. Upon receiving an operation task for the robot, the first simulation engine is controlled to execute the domain randomization process. Domain randomization increases the diversity of training data by randomly changing parameters such as visual texture, lighting conditions, object physical properties, camera viewpoint, and object position and pose in the simulation environment, thereby generating a simulation scene with broad distribution characteristics and improving the robustness and generalization ability of the trained model in the face of different environmental changes in the real world. After completing the domain randomization settings but before starting to execute the first frame of specific simulation action, a scene initialization snapshot is collected at that moment. This snapshot is a structured record of the static configuration of all key entities in the current simulation world, rather than a simple image screenshot. The scene initialization snapshot includes at least the robot's initial pose, the initial poses of each object, and the existence status identifiers of each object. The robot's initial pose describes its initial position and orientation in the world coordinate system, typically represented by a homogeneous transformation matrix or quaternions plus coordinate vectors. Each object's initial pose records the spatial coordinates and rotation angles of every interactive or background object in the scene. Each object's existence status is indicated by a binary or enumerated marker, specifying whether the corresponding object is active or visible in the current task round. For example, in a grasping task, domain randomization might randomly determine the presence of certain interfering objects; the existence status marker accurately reflects this randomization result. By taking snapshots at specific time points before the execution of the first frame of simulation actions, it is ensured that the acquired initial scene data is completely consistent with the physical configuration upon which subsequent simulation operations are based.
[0031] As described in step S102 above, simulation operations are performed after the first simulation engine domain is randomized, and data from the simulation operation process is collected to obtain the first simulation operation data. Based on the randomized scene configuration generated by domain randomization, a specific simulation operation process is executed in the first simulation engine. This process drives the simulation robot to execute a series of predefined motion trajectory planning or strategy control actions according to the operation task requirements. During the robot's execution of actions, data is simultaneously collected throughout the entire simulation operation process to obtain the first simulation operation data. The first simulation operation data constitutes a temporal sequence describing the dynamic changes of the entire interaction process. This data not only includes the robot's own motion information such as joint states and end effector pose and velocity at each moment, but also covers the physical response data of various objects in the scene caused by robot interaction, such as changes in object position, force conditions, changes in contact state, and whether the task success or failure judgment signal is triggered. This acquisition process records data frame by frame to ensure the complete temporal sequence of the data. By acquiring the first simulation operation data containing complete dynamic evolution information, subsequent data conversion and model training can utilize the causal relationship information before and after the action execution. These raw data, together with the aforementioned initial scenario data, constitute a complete closed-loop record from task initialization to execution completion.
[0032] In one embodiment, step S102, which involves performing simulation operations after randomization of the first simulation engine domain and collecting data from the simulation operation process to obtain the first simulation operation data, includes: S1021: After domain randomization, control the robot model in the first simulation engine to perform simulated motions corresponding to the operation task; S1022: Record the robot joint state values of each time frame during the simulation motion in the order of the time frames to obtain the robot joint state trajectory; S1023: Record the original motion command values of each time frame to obtain the original motion trajectory, and record the joint name sequence corresponding to each dimension in the robot joint state trajectory; S1024: Use the robot joint state trajectory, the original motion trajectory, and the joint name sequence as the first simulation operation data.
[0033] As described in steps S1021-S1024 above, after domain randomization, the robot model in the first simulation engine is controlled to execute the simulation motion corresponding to the operation task; the robot joint state values of each time frame during the execution of the simulation motion are recorded in the order of the time frames to obtain the robot joint state trajectory; the original motion command values of each time frame are recorded to obtain the original motion trajectory, and the joint name sequence corresponding to each dimension in the robot joint state trajectory is recorded; the robot joint state trajectory, the original motion trajectory and the joint name sequence are used as the first simulation operation data.
[0034] Within the physical environment of the first simulation engine, based on the target requirements of the current operational task, the robot model's joints and end effectors are driven to produce a series of actions using built-in motion planning algorithms or control strategies. This simulated motion process mimics real physical interactions, including collision detection, gravity response, and friction, ensuring that the robot model exhibits dynamic characteristics consistent with physical laws. The solver of the first simulation engine calculates the robot's spatial position, velocity, and acceleration information at each moment in real time, thereby verifying the feasibility of the operational task and generating corresponding motion trajectory data. During the execution of the above simulated motion, the robot joint state values of each time frame are continuously collected in sequence, forming a time-arranged sequence, i.e., the robot joint state trajectory. The robot joint state values characterize the physical state parameters of all movable joints (such as rotary or translational joints) of the robot at a specific simulation frame moment. These state values constitute a high-dimensional state vector reflecting the robot's configuration, with its dimension corresponding to the number of degrees of freedom of the robot. Taking a seven-DOF robotic arm as an example, the robot joint state values can be represented as a one-dimensional array containing seven elements, each element corresponding to the angular position (in radians) or displacement (in meters) of a joint. This trajectory fully describes the robot's motion changes during task execution. The original motion command values for each time frame are recorded synchronously, forming the original motion trajectory. The original motion command values are the low-level control instructions issued by the first simulation engine to the robot model's actuators during simulated motion execution, such as target position, target velocity, or target torque for joint motors. Unlike joint state values, which reflect the robot's actual state, the original motion command values reflect the state the control system expects the robot to achieve or the control quantities applied. The original motion trajectory consists of original motion command values arranged in chronological order, recording the complete input history of the control strategy. Structurally, the original motion trajectory typically maintains the same frame length and timing alignment as the robot's joint state trajectory, but its data dimensions depend on the definition of the control interface and may include opening / closing commands from the end effector or additional task space control parameters, thus differing from the joint state dimensions. Further, the physical entity identifiers corresponding to each data dimension in the robot's joint state trajectory are determined, generating a joint name sequence. This sequence is a list of strings containing specific identifiers, used to indicate the physical joint name or controller name corresponding to each data dimension in the robot's joint state trajectory. Given that the joint index order is often inconsistent in different environments in multi-robot or multi-simulation engine collaborative scenarios (for example, the first element of the array in the first simulation engine represents the base rotation joint, while in other engines it may represent the shoulder pitch joint), the joint name sequence establishes a mapping relationship from data index to physical entity by explicitly recording the semantic label of each dimension, giving the data self-descriptive ability.The sequence, for example, is in the form of [joint_base_yaw, joint_shoulder_pitch, ..., gripper_open], and its length is strictly equal to the dimension of the robot joint state trajectory. The robot joint state trajectory, the original motion trajectory, and the joint name sequence are aggregated to generate the first simulation operation data. The robot joint state trajectory provides a state-space description of the motion process, the original motion trajectory provides a motion-space description of the decision-making process, and the joint name sequence provides a semantic benchmark to ensure that the data is correctly parsed and mapped in different systems. This combined data structure not only supports playback and analysis within a single environment but also lays the data foundation for subsequent accurate cross-domain replay and training data isomorphism processing across different simulation engines.
[0035] In one embodiment, step S2, which involves splitting the initial scene data into asset data of multiple independent objects using a preset asset processing module, includes: S201: Read the complete static scene asset corresponding to the initial scene data through the asset processing module, and expand the complete static scene asset into a scene hierarchy structure; S202: Traverse each object node in the scene hierarchy and export each object node as an independent object asset file; S203: Copy the material files that each object asset file depends on, and rewrite the reference paths within each object asset file that point to the material files and texture maps, so that each object asset file is independent of each other, and obtain asset data for multiple independent objects.
[0036] As described in steps S201-S203 above, a complete static scene asset can refer to a combined data structure generated by the first simulation engine, containing all geometry, materials, textures, and physical attributes in the scene. Essentially, it is not a single file but a collection of resources organized based on a tree structure. This complete static scene asset is read through the data path or memory object corresponding to the scene initialization snapshot in the aforementioned steps. The asset processing module is configured to parse the internal binary or text format of the complete static scene asset, expanding it into a visual scene hierarchy structure. This hierarchy structure uses a nested parent-child node format, containing, from top to bottom, a scene root node, an environment node, and several object nodes. The scene hierarchy structure consists of multiple interconnected nodes, where object nodes are the basic units constituting the scene. Each object node corresponds to an independent entity in the scene, such as the robot body, the control panel, or the object to be grasped. During the traversal of the scene hierarchy structure by the asset processing module, for example, depth-first search or breadth-first search algorithms can be used to access each node in the tree structure. When the current node type is identified as an object node, an export operation is triggered, separating the object node and its bound geometric mesh data from the overall scene and serializing them into independent object asset files. These can be stored using general 3D exchange formats (such as OBJ, FBX, GLTF, etc.) or proprietary formats specific to certain simulation engines. Initially, these independent object asset files may still retain reference paths to the common material library in the original scene; directly moving or loading them separately can lead to material loss. Therefore, it is necessary to further analyze the dependencies of each object asset file and identify the referenced material files and texture resources. This process includes: creating an independent material subfolder for each object in the export directory, copying the required material files from their original paths to this subfolder; subsequently, modifying the metadata fields within the object asset files, rewriting the material reference paths to point to the newly copied local files. By rewriting the reference paths within each object asset file pointing to the material files and texture maps, it is ensured that each object asset file carries complete dependent resources, thus achieving independence between the object asset files. The resulting asset data for multiple independent objects possesses self-contained characteristics, enabling the second simulation engine, after scene initialization, to load any object on demand without relying on the complete original scene package. This asset decoupling and path rewriting mechanism transforms the originally tightly coupled overall scene into a modular asset collection that can be independently distributed and flexibly reused, thereby improving the compatibility and processing efficiency of cross-simulation engine data migration. Furthermore, the asset decoupling and path rewriting mechanism ensures the independence of each object's asset file, allowing the second simulation engine, after scene initialization, to load any object on demand.
[0037] In one embodiment, step S3, which involves inputting asset data of multiple independent objects into the initial second simulation engine to obtain the initial pose of each independent object in the world coordinate system and the existence state of each independent object, thereby completing the scene initialization of the second simulation engine, includes: S301: Load the asset data of each of the independent objects one by one through the initial second simulation engine; S302: Read the robot initial pose, the initial pose of each object, and the existence status identifier of each object contained in the initial scene data. S303: Place the preset target robot asset in the second simulation engine scene according to the robot's initial pose; S304: Place the corresponding objects in the second simulation engine scene according to the initial pose of each object, and load or hide the corresponding objects according to the existence status flag of each object, thereby completing the scene initialization of the second simulation engine.
[0038] As described in step S301 above, the asset data of each independent object is loaded one by one through the initial second simulation engine. The initial second simulation engine serves as the execution container for scene reconstruction and initiates the asset loading process. The multiple independent object asset files generated in the preceding steps serve as input data. These files contain information such as the geometric models, material textures, and physical properties of the objects. By introducing the scattered asset files one by one into the engine, virtual object instances that can be used for physical calculations and rendering are constructed, thereby providing the material basis for the second simulation engine to reproduce the scene required by the first simulation engine.
[0039] As described in step S302 above, the previously acquired initial scene data is extracted from the storage medium. This initial scene data records the static configuration information of the first simulation engine at the moment of domain randomization completion. This data is parsed to locate the robot's initial pose, the initial poses of each object, and the existence status identifiers of each object. The pose information consists of position coordinates and rotation quaternions or Euler angles in the world coordinate system, and the existence status identifiers indicate whether a specific object should be rendered or participate in physical interaction in the current task round. By capturing the randomization results of the first simulation engine, accurate parameter basis is provided for subsequent state synchronization.
[0040] As described in step S303 above, a preset target robot asset is placed in the second simulation engine scene according to the robot's initial pose. Based on the read robot initial pose, the preset target robot asset is deployed in the virtual space of the second simulation engine. This preset target robot asset is a high-fidelity robot model specifically configured for the final training target, and its structure may differ from the simplified model used in the first simulation engine. A coordinate transformation operation is performed to precisely align the origin of the local coordinate system of the preset target robot asset to the world coordinate position and attitude specified by the robot's initial pose through translation and rotation. Thus, the robot in the second simulation engine is given an initial spatial state that is completely consistent with that of the first simulation engine, ensuring the consistency of the starting conditions for subsequent motion replay.
[0041] As described in step S304 above, corresponding objects are placed in the second simulation engine scene according to their initial poses, and the corresponding objects are loaded or hidden according to their existence status identifiers, thereby completing the scene initialization of the second simulation engine. After the robot assets are placed, spatial layout and visibility configuration are performed on each independent object. For each loaded independent object, it is positioned to a specified coordinate in the second simulation engine scene according to its corresponding object initial pose in the initial scene data. At the same time, differentiated processing is performed based on the existence status identifiers of each object: if the identifier indicates that the object exists in the current task round, the loading state of the object is maintained and it participates in subsequent physical simulation and rendering; if the identifier indicates that the object does not exist, the object is ensured not to affect the simulation environment through a hiding operation (such as turning off the rendering layer, disabling physical colliders, or removing it directly from the scene tree). Through the precise placement and state filtering of all objects, the second simulation engine is configured to be a completely equivalent running environment to the first simulation engine at the initial moment, that is, the second simulation engine after scene initialization. The second simulation engine, after initialization, not only reproduces the spatial distribution of objects but also accurately restores the randomized scene configuration of the first simulation engine, providing the necessary environmental support for trajectory replay based on the same starting line. The second simulation engine, constructed through the above steps, achieves complete replication of the initialization state across simulation engines by precisely aligning the pose parameters in the world coordinate system and strictly executing the masking logic of the existence state identifier. This state consistency reconstruction mechanism allows the second simulation engine to execute subsequent simulation replays in a deterministic initial state consistent with the source engine, thereby eliminating simulation errors caused by initial condition deviations and laying the environmental foundation for generating high-quality robot operation training data. By accurately placing objects and differentiating them according to the existence state identifier, the second simulation engine is configured to operate in a completely equivalent environment to the first simulation engine at the initial moment.
[0042] In one embodiment, step S4, which involves inputting the first simulation operation data into a second simulation engine after scene initialization to perform simulation operations and obtain the second simulation operation data, includes: S401: Read the joint name sequence in the first simulation operation data, and obtain the runtime joint name of the preset target robot asset in the second simulation engine after scene initialization; S402: Match each joint name in the joint name sequence with the runtime joint name to establish a joint mapping relationship; S403: Based on the joint mapping relationship, the joint state values of each time frame in the joint state trajectory of the first simulation operation data are converted frame by frame into the executable joint target values of the second simulation engine after scene initialization. S404: The joint target values of each time frame are sent out sequentially according to the time frame order, driving the target robot asset in the second simulation engine after scene initialization to perform motion replay; S405: During the motion replay process, multimodal data including images, semantic annotations, robot state, actions, and language commands are recorded synchronously to obtain the second simulation operation data.
[0043] As described in step S401 above, the joint name sequence in the first simulation operation data is read. A pre-stored joint name sequence is extracted from the first simulation operation data. This first simulation operation data contains key state information recorded by the first simulation engine for driving robot motion. The joint name sequence is a list of identifiers that corresponds one-to-one with each dimension of the robot's state trajectory, used to identify the physical joint entity corresponding to each column of values in the state trajectory. Compared to conventional techniques that only store numerical matrices and ignore semantic labels, retaining the joint name sequence ensures traceability and semantic consistency of data when migrating between different simulation environments. In a specific implementation, the joint name sequence can be represented as a string array, for example, containing elements such as joint_1, joint_2, etc., with each element uniquely indexing a specific dimension of the state trajectory. By parsing this sequence, the necessary semantic indexing foundation is built for subsequent cross-engine data mapping.
[0044] The runtime joint names of the preset target robot assets in the second simulation engine after scene initialization are obtained. After determining the joint definitions at the source data end, the joint definition information of the preset target robot assets in the current running environment within the second simulation engine after scene initialization is further obtained. This process is achieved by calling the application programming interface (API) of the second simulation engine or parsing its robot model description file. The runtime joint names reflect the actual kinematic structure of the robot model currently loaded in the second simulation engine. This structure may differ from the model in the first simulation engine in terms of joint arrangement order, number of joints, or naming conventions. For example, the first simulation engine may use a hierarchical naming method (such as base_link->link1->link2), while the second simulation engine may use a hardware-driven naming method (such as motor_1, motor_2). Dynamically reading these runtime joint names can accurately capture the internal logical structure of the target environment, thus providing a basis for solving the model heterogeneity problem between different simulation engines.
[0045] As described in step S402 above, each joint name in the joint name sequence is matched one-to-one with the runtime joint name to establish a joint mapping relationship. Priority is given to finding names that are completely identical to establish a mapping relationship. If a joint name cannot be matched successfully, it is converted according to a preset joint semantic mapping table. This mapping table records the correspondence between joint names between the first simulation engine and the second simulation engine (e.g., "arm_1_joint" corresponds to "joint_2"), thus establishing a complete joint mapping relationship. Each name in the joint name sequence is traversed, and a completely identical corresponding item is retrieved from the runtime joint name set. When a joint name recorded by the first simulation engine is found to be identical to a runtime joint name of the second simulation engine, a mapping association is established between the two. This mapping process can be formally described as follows: Let the joint name sequence recorded by the first simulation engine be... The runtime joint name sequence of the second simulation engine is as follows: For the second simulation engine, the first... A joint, if it exists Then establish a mapping. ,in This indicates the column index of the joint in the state trajectory of the first simulation engine. Through this name-based matching mechanism, even if the joint definition order or the number of joints differs between the two engines, the data position of the same physical joint in different systems can be accurately identified, thereby constructing the correct data channel and avoiding the risk of robot motion disorder or loss of control caused by replaying according to fixed column numbers.
[0046] As described in step S403 above, based on the joint mapping relationship, the joint state values of each time frame in the joint state trajectory of the first simulation operation data are converted frame by frame into the executable joint target values of the second simulation engine after scene initialization. The established joint mapping relationship is used to perform format conversion on the time-series state data in the first simulation operation data. The robot joint state trajectory in the first simulation operation data records the robot state changing over time, typically presented as a matrix-like data sequence. According to the mapping relationship, the joint state values of each time frame are extracted from the state trajectory of the first simulation engine and rearranged according to the joint order of the target engine to generate a sequence of executable joint target values that the second simulation engine can recognize and execute. This conversion process achieves precise routing from numerical values to target physical joints. For example, if the first simulation engine's... The column data corresponds to the joint left_arm, and the mapping relationship indicates that the joint is located in the second simulation engine. The column will then be the first of all frames in the trajectory. The column data is assigned sequentially to the first column of the target sequence. By rearranging data by name, heterogeneous source data is homogenized into standard control commands for the target engine, eliminating data biases introduced by differences in coordinate system definitions or joint indexes, and achieving precise migration of control signals across engines.
[0047] As described in step S404 above, the joint target values of each time frame are sequentially sent out according to the time frame order, driving the target robot asset in the second simulation engine after scene initialization to perform motion replay. The converted executable joint target values are input into the second simulation engine after scene initialization to drive the target robot asset to perform motion. Specifically, according to the original recorded time frame order, starting from the first frame, the joint target values of each frame are sequentially sent to the controller of the second simulation engine. After receiving these target values, the second simulation engine uses them as motion control commands to update the joint state of the target robot asset in real time. This process is a precise reproduction of the physical motion process that occurs in the first simulation engine, i.e., motion replay. During the replay process, time synchronization is strictly controlled to ensure that the time interval between the sending of each frame of data is consistent with the original record, thereby restoring a coherent and smooth robot motion trajectory. This replay mechanism based on precise data-driven operation allows the second simulation engine to reproduce the complex operation actions generated by the first simulation engine without relying on complex real-time planning algorithms.
[0048] As described in step S405 above, multimodal data including images, semantic annotations, robot states, actions, and language commands are simultaneously recorded during motion replay to obtain the second simulation operation data. While driving the target robot asset to perform motion replay, the advantages of the second simulation engine in rendering and annotation are utilized to simultaneously record multimodal data. The second simulation engine is configured to trigger multiple sensors or recording modules to work in parallel during each frame of simulation step. These modules include: a camera sensor for acquiring color and depth images of the scene; a semantic segmentation module for generating semantically annotated images containing object categories and instance masks; a state listener for recording the robot's current joint angles, end-effector poses, and velocity information; and a command recorder for capturing the corresponding natural language command or raw action command at the current moment. All the above data streams are strictly aligned in timestamps to form a record containing rich information. Finally, all data recorded throughout the motion replay process is summarized to obtain the second simulation operation data. This data not only includes the actions themselves but also includes high-fidelity visual perception data and semantic understanding data, providing high-quality, multimodal supervision signals for subsequent training of the robot's visual-language-action model.
[0049] In addition, the method also includes: verifying whether the format version of the first simulation operation data is compatible with the second simulation engine after scene initialization; verifying whether the total number of frames of the joint state trajectory is greater than zero; verifying whether there are duplicate joint names in the joint name sequence; and verifying whether the dimension of the joint state trajectory is consistent with the length of the joint name sequence.
[0050] The compatibility check between the format version of the first simulation operation data and the second simulation engine after scene initialization aims to ensure the consistency of the data interface and protocol. This is achieved by reading the version identifier field carried in the header or metadata of the first simulation operation data and comparing it with the list of versions supported by the second simulation engine after scene initialization. If the version identifier is within the supported range, it is considered compatible; if the version identifier is too high, too low, or undefined, it indicates that there may be incompatible field additions, deletions, or type changes in the data structure, and in this case, it is considered incompatible. This format version compatibility check prevents parsing errors caused by data structure evolution, ensuring that the second simulation engine after scene initialization can correctly read and process the incoming data. The validity check of the total number of frames in the joint state trajectory is a fundamental verification of data existence. The joint state trajectory records the sequence of joint states of the robot over time, and the total number of frames corresponds to the time length of this sequence. The total number of frames is obtained by parsing the frame counter in the data file or calculating the row number of the state matrix. If the total number of frames is less than or equal to zero, it means that the trajectory is empty or does not contain any valid motion information. Only when the total number of frames is greater than zero does it indicate the existence of a replayable motion process, thus avoiding invalid replay operations on empty data. Uniqueness verification of the joint name sequence ensures the determinism of joint mapping relationships. The joint name sequence is the key basis for establishing the joint correspondence between the first simulation engine and the second simulation engine after scene initialization. By traversing the joint name sequence, the frequency of each name is counted using data structures such as hash tables or sets. If any joint name is detected to appear more than once, it is determined to be a duplicate, and the system will terminate the subsequent process or report an error. This mechanism eliminates the many-to-one mapping ambiguity caused by duplicate joint names in the sequence, ensuring that the target state of a specific joint can be uniquely determined, thereby guaranteeing the accuracy of joint control during replay. Consistency verification between the joint state trajectory dimension and the joint name sequence length aims to ensure the logical self-consistency of the data structure. The dimension of the joint state trajectory is usually represented by the number of columns in the state matrix, and each column should correspond to a specific joint in the joint name sequence. The shape information of the joint state trajectory data is extracted to obtain its column number or feature dimension value, and this is numerically compared with the number of elements in the joint name sequence. If the two values are equal, it indicates that each joint state value has a corresponding name label; if they are not equal, it means that there are missing or redundant columns in the data. This dimensional consistency check effectively filters out incomplete datasets caused by data truncation or splicing errors, ensuring the integrity and reliability of the replay data. The above verification steps together constitute the data quality defense line before replay. Through multiple checks on format version, frame validity, name uniqueness, and dimensional consistency, abnormal data can be identified and blocked in advance before the second simulation engine executes the core replay logic after the data enters the scene initialization stage.This pre-verification mechanism avoids simulation engine crashes, wasted computing resources, or the generation of incorrect training samples due to data anomalies, thereby improving the robustness and stability of the entire dual-simulation engine data generation system.
[0051] In one embodiment, step S5, which involves isomorphizing the first simulation operation data and the second simulation operation data to obtain the first training data and the second training data respectively, includes: S501: Input the first simulation operation data into the first conversion function to process the first simulation operation data into first training data containing image features, state features, action features and language features; S502: Input the second simulation operation data into the second conversion function to process the second simulation operation data into second training data containing image features, state features, action features and language features; wherein, the first conversion function and the second conversion function adopt the same data structure definition, and the output field names and field semantics correspond one-to-one.
[0052] As described in steps S501-S502 above, the first simulation operation data is input into the first transformation function, and the first simulation operation data is processed into first training data containing image features, state features, action features, and language features; the second simulation operation data is input into the second transformation function, and the second simulation operation data is processed into second training data containing image features, state features, action features, and language features. The first and second transformation functions use the same data structure definition, and the output field names and field semantics correspond one-to-one. The first simulation operation data obtained from the first simulation engine typically includes temporally sequenced joint states, raw action commands, and visual observations. These data often directly correspond to the memory structure of the simulation engine in their internal storage format. Through the first transformation function, these raw data from the first simulation engine are mapped to a unified training data pattern, extracting key representations for model training. In this process, image features are constructed as structured tensor sequences representing the visual environment state during the operation task, rather than simple pixel matrix stacks, to reflect the spatiotemporal changes in object pose, texture, and lighting conditions. State features are used to describe the robot's physical configuration at the current moment, such as joint angles, end-effector pose, and velocity. Motion features represent various control commands that drive changes in the robot's state, such as joint torques or target position values. Language features are vectorized representations of natural language commands or task descriptions associated with the current operation task. Specifically, the original motion trajectories and robot joint state trajectories in the first simulation operation data are read, and continuous motions are smoothed or normalized to generate standardized motion feature vectors. Simultaneously, based on a predefined semantic mapping table, the task description text is converted into fixed-dimensional embedding vectors as language features, which are then integrated to form the first training data. For the second simulation operation data recorded during motion replay, which contains high-fidelity images, semantic annotations, and other information, and whose data organization differs from the first simulation engine, a second transformation function is used to convert heterogeneous data to a standard format, processing it into second training data containing image features, state features, motion features, and language features. During processing, high-fidelity image sequences are extracted and processed using Convolutional Neural Networks (CNNs) or directly used as image feature input. For data containing rich semantic segmentation information, corresponding masks or category labels are extracted as auxiliary channels for image features. The robot's state and action records are parsed and rearranged into vectors with the same dimensions as the first training data. The same vocabulary mapping and encoding operations are performed on language commands. Through the above transformations, the high-quality rendering data from the second simulation engine is reorganized into structured samples, ensuring that the generated second training data maintains a strict logical correspondence with the first training data in terms of feature composition.To ensure consistency across different data sources, both the first and second transformation functions use the same data structure definition, ensuring a one-to-one correspondence between output field names and their semantics. This "same data structure definition" means that both transformation functions follow the same pre-defined data pattern or architectural specifications, such as using Python dictionaries, C++ structs, or specific data serialization formats (e.g., JSON, Protobuf). Under this definition, the output field names are completely identical; for example, both use `image`, `state`, `action`, and `language` as keys. A one-to-one correspondence in field semantics means that the physical meaning, data type, and numerical range represented by the same field name are strictly identical across different data sources. For example, if the field name `state` represents a 7-dimensional robot end-effector pose vector in the first training data, then in the second training data, this field must also represent a 7-dimensional end-effector pose vector, not a joint angle. The pre-defined data structure is determined based on the input interface requirements of the target model: Input interface specifications of the target training model (e.g., a visual-language-action model) are collected, and the required field list and tensor dimension requirements for each field are extracted. Based on this, a data loading script and validation logic are written, forming a standard data structure definition file. Both the first and second transformation functions load this definition file during implementation, organizing the output data according to a unified key-value pair rule. If a new feature type (e.g., a depth map) needs to be added, only the data structure definition file needs to be updated, and the two transformation functions adjusted synchronously. Through the standardization processing of the first and second transformation functions, the first and second simulation operation data are mapped to the same semantic space. This isomorphism eliminates the differences in data storage format, field naming conventions, and numerical encoding methods between different simulation engines, allowing subsequent training pipelines to directly read and process data without distinguishing its source. The first training data retains the large-scale domain randomization advantage of the first simulation engine, while the second training data retains the high-fidelity rendering and semantic annotation advantages of the second simulation engine. Using both together effectively improves the model's generalization ability to different visual and physical domains. By standardizing data, data from different sources is mapped to the same semantic space, eliminating data format differences and eliminating the need for subsequent training pipelines to distinguish data sources.
[0053] In one embodiment, the step of inputting data into a training data resource pool for training includes: S601: Merge the first training data and the second training data and store them in the training data resource pool; when the same training pipeline reads training data from the training data resource pool, the same parsing logic and input format are used for the first training data and the second training data; use the merged and stored training data to jointly train the visual language action model.
[0054] Specifically, the first and second training data are merged and stored in a training data resource pool. This pool is configured to centrally manage structured data produced by different simulation engines. At the physical storage level, the pool can be constructed using a distributed file system or object storage service, writing the first training data from the first simulation engine and the second training data from the second simulation engine into the same logical storage space. Despite the differences in their data sources, these data are uniformly treated as a sequence of training samples within the pool. Merging the storage breaks down data silos, ensuring that subsequent training processes can access the full set of samples without discrimination, thus comprehensively utilizing the breadth advantage of the first simulation engine in domain randomization and the accuracy advantage of the second simulation engine in visual realism. When the same training pipeline reads training data from the training data resource pool, the same parsing logic and input format are used for both the first and second training data. To eliminate access barriers caused by heterogeneous data sources, the training pipeline is configured with a unified interface specification at the data reading layer. Regardless of whether the data samples specifically come from the first or second simulation engine, the parsing logic calls the same set of data loaders and field parsing functions. The uniformity of the input format is reflected in the dimensional definitions, data types, and key-value pair mappings of the data tensors. For example, for key fields such as image features, state features, action features, and language features, the data from both sides are aligned to a completely consistent tensor shape and semantic space before being input into the model. Through this unified parsing process, the training pipeline can treat data from mixed sources as a homogeneous flow, eliminating the need to write specific branch logic for different data sources in the model training loop. The visual-language-action model is jointly trained using the merged and stored training data. The visual-language-action model is a multimodal neural network model capable of simultaneously processing visual image information, natural language commands, and robot action sequences. During the model training phase, training algorithms (such as stochastic gradient descent and its variants) randomly sample small batches of data from the training data resource pool, where each batch may contain first and second training data from mixed sources. The model's forward propagation process receives these uniformly formatted inputs, calculates the loss value (such as mean squared error or cross-entropy loss) between the predicted action and the actual action, and updates the model parameters through the backpropagation algorithm. The joint training mechanism forces the model to learn general operational strategies across different rendering styles, physical feedback, and asset characteristics, rather than overfitting to the specific features of a single simulation environment. Through this cross-domain joint optimization, the model can exhibit stronger environmental adaptability and operational robustness in actual deployment.
[0055] Reference Figure 3 The present invention also provides a robot operation training data generation device based on a dual simulation engine, the device comprising: The control module 902 is used to respond to the received operation task, control the first simulation engine to perform simulation operation, and acquire the first simulation operation data and the initial scene data; wherein, the first simulation engine is configured to perform scene generation and motion trajectory planning for robot operation tasks based on domain randomization. The splitting module 904 is used to split the initial scene data into asset data of multiple independent objects through a preset asset processing module; wherein, the preset asset processing module is an asset format conversion unit located between the two engines, used to decompose the initial scene data generated by the first engine into asset data that can be loaded independently by the second engine; The first input module 906 is used to input asset data of multiple independent objects into the initial second simulation engine to obtain the initial pose of each independent object in the world coordinate system and the existence state of each independent object, thereby completing the scene initialization of the second simulation engine; wherein, the initial second simulation engine is configured to load preset target robot assets and perform rendering and visual semantic annotation. The simulation module 908 is used to input the first simulation operation data into the second simulation engine after the scene is initialized to perform simulation operations and obtain the second simulation operation data. Processing module 910 is used to perform isomorphic processing on the first simulation operation data and the second simulation operation data to obtain the first training data and the second training data, respectively. The second input module 912 is used to input the first training data and the second training data into a preset training data resource pool for training the target robot.
[0056] In one embodiment, the control module 902 includes: The control submodule is used to respond to the received operation task, control the first simulation engine to perform domain randomization, and collect a scene initialization snapshot after domain randomization and before the execution of the first frame simulation action to obtain the initial scene data; wherein, the scene initialization snapshot includes at least the robot initial pose, the initial pose of each object, and the existence status identifier of each object. The acquisition submodule is used to perform simulation operations after randomization in the first simulation engine domain, and to acquire data during the simulation operation process to obtain the first simulation operation data.
[0057] In one embodiment, the acquisition submodule includes: An execution unit is used to control the robot model in the first simulation engine to perform simulated motions corresponding to the operation task after domain randomization. The first recording unit is used to record the robot joint state values of each time frame during the execution of the simulation motion in the order of the time frames, so as to obtain the robot joint state trajectory; The second recording unit is used to record the original motion command values of each time frame to obtain the original motion trajectory, and to record the joint name sequence corresponding to each dimension in the robot joint state trajectory. As a unit, it is used to use the robot joint state trajectory, the original motion trajectory, and the joint name sequence as the first simulation operation data.
[0058] In one embodiment, the splitting module 904 includes: The expansion submodule is used to read the complete static scene asset corresponding to the initial scene data through the asset processing module, and expand the complete static scene asset into a scene hierarchy structure. The traversal submodule is used to traverse each object node in the scene hierarchy and export each object node as an independent object asset file. The rewrite submodule is used to copy the material files that each object asset file depends on, and rewrite the reference paths inside each object asset file that point to the material files and texture maps, so that each object asset file is independent of each other, resulting in asset data for multiple independent objects.
[0059] In one embodiment, the first input module 906 includes: The loading submodule is used to load the asset data of each of the independent objects one by one through the initial second simulation engine; The reading submodule is used to read the robot's initial pose, the initial pose of each object, and the existence status identifier of each object contained in the initial scene data. The first placement submodule is used to place the preset target robot asset in the second simulation engine scene according to the robot's initial pose. The second placement submodule is used to place the corresponding objects in the scene of the second simulation engine according to the initial pose of each object, and to load or hide the corresponding objects according to the existence status identifier of each object, thereby completing the scene initialization of the second simulation engine.
[0060] In one embodiment, the simulation module 908 includes: The reading submodule is used to read the joint name sequence in the first simulation operation data and to obtain the runtime joint names of the preset target robot assets in the second simulation engine after scene initialization. The matching submodule is used to match each joint name in the joint name sequence with the runtime joint name one by one to establish a joint mapping relationship; The conversion submodule is used to convert the joint state values of each time frame in the joint state trajectory of the first simulation operation data into the executable joint target values of the second simulation engine after scene initialization, according to the joint mapping relationship. The driving submodule is used to sequentially send out the joint target values of each time frame according to the time frame order, and drive the target robot asset in the second simulation engine after scene initialization to perform motion replay; The recording submodule is used to synchronously record multimodal data including images, semantic annotations, robot states, actions, and language commands during the motion replay process to obtain the second simulation operation data.
[0061] In one embodiment, the processing module 910 includes: The first input submodule is used to input the first simulation operation data into the first transformation function to process the first simulation operation data into first training data containing image features, state features, action features and language features; The second input submodule is used to input the second simulation operation data into the second conversion function to process the second simulation operation data into second training data containing image features, state features, action features and language features; wherein, the first conversion function and the second conversion function adopt the same data structure definition, and the output field names and field semantics correspond one-to-one.
[0062] Figure 4 An internal structural diagram of an electronic device in one embodiment is shown. This electronic device can specifically be a terminal or a server, and more specifically, a computer device. Figure 4 As shown, the electronic device includes a processor, a memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and may also store a computer program. When executed by the processor, this computer program enables the processor to implement a robot operation training data generation method based on a dual-simulation engine. The internal memory may also store a computer program, which, when executed by the processor, enables the processor to execute the robot operation training data generation method based on a dual-simulation engine. Those skilled in the art will understand that... Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0063] In one embodiment, an electronic device is provided, including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the following steps: In response to the received operation task, the first simulation engine is controlled to perform simulation operation and acquire the first simulation operation data and the initial scene data; wherein, the first simulation engine is configured to perform scene generation and motion trajectory planning for robot operation tasks based on domain randomization. The initial scene data is split into asset data of multiple independent objects through a preset asset processing module; wherein, the preset asset processing module is an asset format conversion unit located between the two engines, used to decompose the initial scene data generated by the first engine into asset data that can be loaded independently by the second engine. The asset data of multiple independent objects are input into the initial second simulation engine to obtain the initial pose of each independent object in the world coordinate system and the existence state of each independent object, thereby completing the scene initialization of the second simulation engine; wherein, the initial second simulation engine is configured to load preset target robot assets and perform rendering and visual semantic annotation. The first simulation operation data is input into the second simulation engine after the scene is initialized to perform simulation operations, and the second simulation operation data is obtained. The first simulation operation data and the second simulation operation data are homogenized to obtain the first training data and the second training data, respectively. The first and second training data are input into a preset training data resource pool for training the target robot.
[0064] It effectively solves the technical problems faced in existing cross-simulation engine data migration and joint training applications, such as the inability to accurately reproduce the initial random state and the existence of objects, and the difficulty in effectively integrating multi-source data for joint training of the same model due to heterogeneity. Therefore, it avoids the situation of data alignment failure and limited model training generalization ability caused by differences in simulation environment, and improves the fusion efficiency and application value of multi-source simulation data in visual language action model training.
[0065] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, causes the processor to perform the following steps: In response to the received operation task, the first simulation engine is controlled to perform simulation operation and acquire the first simulation operation data and the initial scene data; wherein, the first simulation engine is configured to perform scene generation and motion trajectory planning for robot operation tasks based on domain randomization. The initial scene data is split into asset data of multiple independent objects through a preset asset processing module; wherein, the preset asset processing module is an asset format conversion unit located between the two engines, used to decompose the initial scene data generated by the first engine into asset data that can be loaded independently by the second engine. The asset data of multiple independent objects are input into the initial second simulation engine to obtain the initial pose of each independent object in the world coordinate system and the existence state of each independent object, thereby completing the scene initialization of the second simulation engine; wherein, the initial second simulation engine is configured to load preset target robot assets and perform rendering and visual semantic annotation. The first simulation operation data is input into the second simulation engine after the scene is initialized to perform simulation operations, and the second simulation operation data is obtained. The first simulation operation data and the second simulation operation data are homogenized to obtain the first training data and the second training data, respectively. The first and second training data are input into a preset training data resource pool for training the target robot.
[0066] It effectively solves the technical problems faced in existing cross-simulation engine data migration and joint training applications, such as the inability to accurately reproduce the initial random state and the existence of objects, and the difficulty in effectively integrating multi-source data for joint training of the same model due to heterogeneity. Therefore, it avoids the situation of data alignment failure and limited model training generalization ability caused by differences in simulation environment, and improves the fusion efficiency and application value of multi-source simulation data in visual language action model training.
[0067] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0068] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0069] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for generating robot operation training data based on a dual simulation engine, characterized in that, The method includes: In response to the received operation task, the first simulation engine is controlled to perform simulation operation and acquire the first simulation operation data and the initial scene data; wherein, the first simulation engine is configured to perform scene generation and motion trajectory planning for robot operation tasks based on domain randomization. The initial scene data is split into asset data of multiple independent objects through a preset asset processing module; wherein, the preset asset processing module is an asset format conversion unit located between the two engines, used to decompose the initial scene data generated by the first engine into asset data that can be loaded independently by the second engine. The asset data of multiple independent objects are input into the initial second simulation engine to obtain the initial pose of each independent object in the world coordinate system and the existence state of each independent object, thereby completing the scene initialization of the second simulation engine; wherein, the initial second simulation engine is configured to load preset target robot assets and perform rendering and visual semantic annotation. The first simulation operation data is input into the second simulation engine after the scene is initialized to perform simulation operations, and the second simulation operation data is obtained. The first simulation operation data and the second simulation operation data are homogenized to obtain the first training data and the second training data, respectively. The first and second training data are input into a preset training data resource pool for training the target robot.
2. The method for generating robot operation training data based on a dual simulation engine according to claim 1, characterized in that, The steps of controlling the first simulation engine to perform simulation operations in response to the received operation task, and acquiring the first simulation operation data and the initial scene data, include: In response to the received operation task, the first simulation engine is controlled to perform domain randomization, and a scene initialization snapshot is collected after domain randomization and before the execution of the first frame simulation action to obtain the initial scene data; wherein, the scene initialization snapshot includes at least the robot initial pose, the initial pose of each object, and the existence status identifier of each object. After randomization in the first simulation engine domain, simulation operations are performed, and data from the simulation operation process is collected to obtain the first simulation operation data.
3. The method for generating robot operation training data based on a dual simulation engine according to claim 2, characterized in that, The step of performing simulation operations after randomization in the first simulation engine domain and collecting data from the simulation operation process to obtain the first simulation operation data includes: After domain randomization, the robot model in the first simulation engine is controlled to perform simulated motions corresponding to the operation task; The robot joint state values of each time frame during the simulation motion are recorded in sequence to obtain the robot joint state trajectory. Record the original motion command values of each time frame to obtain the original motion trajectory, and record the joint name sequence corresponding to each dimension in the robot joint state trajectory; The robot joint state trajectory, the original motion trajectory, and the joint name sequence are used as the first simulation operation data.
4. The method for generating robot operation training data based on a dual simulation engine according to claim 1, characterized in that, The step of splitting the initial scene data into asset data of multiple independent objects through a preset asset processing module includes: The asset processing module reads the complete static scene asset corresponding to the initial scene data and expands the complete static scene asset into a scene hierarchy structure. Traverse each object node in the scene hierarchy and export each object node as an independent object asset file. The material files that each object asset file depends on are copied, and the reference paths within each object asset file pointing to the material files and texture maps are rewritten to make each object asset file independent of each other, thus obtaining asset data for multiple independent objects.
5. The method for generating robot operation training data based on a dual simulation engine according to claim 1, characterized in that, The step of inputting asset data of multiple independent objects into the initial second simulation engine to obtain the initial pose and existence state of each independent object in the world coordinate system, thereby completing the scene initialization of the second simulation engine, includes: The asset data of each of the independent objects is loaded one by one through the initial second simulation engine; Read the robot initial pose, the initial pose of each object, and the existence status identifier of each object contained in the initial scene data; Place the preset target robot asset in the second simulation engine scene according to the robot's initial pose; Based on the initial pose of each object, the corresponding object is placed in the scene of the second simulation engine, and the corresponding object is loaded or hidden according to the existence status flag of each object, thereby completing the scene initialization of the second simulation engine.
6. The method for generating robot operation training data based on a dual simulation engine according to claim 1, characterized in that, The step of inputting the first simulation operation data into the second simulation engine after scene initialization to perform simulation operations and obtain the second simulation operation data includes: Read the joint name sequence in the first simulation operation data, and obtain the runtime joint names of the preset target robot assets in the second simulation engine after scene initialization; Each joint name in the joint name sequence is matched one by one with the runtime joint name to establish a joint mapping relationship; Based on the joint mapping relationship, the joint state values of each time frame in the joint state trajectory of the first simulation operation data are converted frame by frame into the executable joint target values of the second simulation engine after scene initialization. The joint target values of each time frame are sent out sequentially according to the time frame order, driving the target robot asset in the second simulation engine after scene initialization to perform motion replay; During the motion replay process, multimodal data including images, semantic annotations, robot states, actions, and language commands are recorded synchronously to obtain the second simulation operation data.
7. The method for generating robot operation training data based on a dual simulation engine according to claim 1, characterized in that, The step of isomorphizing the first simulation operation data and the second simulation operation data to obtain the first training data and the second training data respectively includes: The first simulation operation data is input into the first transformation function to process the first simulation operation data into first training data containing image features, state features, action features and language features; The second simulation operation data is input into the second transformation function to process the second simulation operation data into second training data containing image features, state features, action features and language features; The first conversion function and the second conversion function use the same data structure definition, and the output field names and field semantics correspond one-to-one.
8. A robot operation training data generation device based on a dual simulation engine, characterized in that, The device includes: The control module is used to respond to the received operation task, control the first simulation engine to perform simulation operation, and acquire the first simulation operation data and the initial scene data; wherein, the first simulation engine is configured to perform scene generation and motion trajectory planning for robot operation tasks based on domain randomization. The splitting module is used to split the initial scene data into asset data of multiple independent objects through a preset asset processing module; wherein, the preset asset processing module is an asset format conversion unit located between the two engines, used to decompose the initial scene data generated by the first engine into asset data that can be loaded independently by the second engine; The first input module is used to input asset data of multiple independent objects into the initial second simulation engine to obtain the initial pose of each independent object in the world coordinate system and the existence state of each independent object, thereby completing the scene initialization of the second simulation engine; wherein, the initial second simulation engine is configured to load preset target robot assets and perform rendering and visual semantic annotation. The simulation module is used to input the first simulation operation data into the second simulation engine after the scene is initialized to perform simulation operations and obtain the second simulation operation data. The processing module is used to perform isomorphic processing on the first simulation operation data and the second simulation operation data to obtain the first training data and the second training data, respectively. The second input module is used to input the first training data and the second training data into a preset training data resource pool for training the target robot.
9. A computer-readable storage medium, characterized in that, The system contains a computer program that, when executed by a processor, causes the processor to perform the steps of the robot operation training data generation method based on a dual simulation engine as described in any one of claims 1 to 7.
10. An electronic device, characterized in that, The device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the robot operation training data generation method based on a dual simulation engine as described in any one of claims 1 to 7.