Body-aware multi-modal data synthesis method, device, medium and product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]针对现有的多源数据利用不充分、生成结果跨模态不一致、缺乏基于任务结构的定向扩增机制及训练价值评估缺失等问题,现提供一种旨在实现任务语义引导下关键节点定向扩增、多模态轨迹与画面同步生成以及训练增益自动筛选的具身智能多模态数据合成方法、设备、介质及产品
[0029]This application's embodied intelligence multimodal data synthesis method achieves precise localization of decisive moments in the task execution structure through multi-level task phase decomposition and key node extraction. Based on this, a task causal skeleton is constructed, clearly defining the range of immutable constraints and variable factors, providing structured task semantic guidance for subsequent targeted augmentation and significantly improving the relevance of data augmentation. By applying controlled counterfactual perturbations to key nodes under causal skeleton constraints, new robot motion trajectories, state trajectories, and multi-view color and depth videos are generated simultaneously, ensuring strict temporal and 3D geometric alignment between visual and motion information, solving the cross-modal inconsistency problem in traditional methods. A training gain evaluation mechanism based on sensitivity, novelty, boundary coverage, and visual robustness is introduced, combined with cross-modal consistency verification for sample screening, automatically eliminating low-quality samples and ensuring that the output data has a real gain for embodied intelligence model training. The final synthesized data maintains the same organizational form as the real-world data and can be directly used for training models such as imitation learning and policy learning, improving the utilization efficiency and practical value of multimodal data.
Smart Images

Figure CN122548656A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to methods, devices, media and products for synthesizing embodied intelligence multimodal data. Background Technology
[0002] Embodied intelligent model training relies on multimodal aligned data, including vision, depth, robot state, motion trajectory, and stage semantic annotation, emphasizing temporal continuity, spatial geometric consistency, and semantic interpretability. However, existing data generation methods suffer from four shortcomings in this scenario: First, multi-source data is not fully utilized, typically using only RGB video or robot trajectory alone, without effectively integrating depth, joint state, gripper state, and task metadata; second, the generated results only change a single modality, with image augmentation and trajectory generation being independent, leading to temporal misalignment and cross-modal inconsistency between vision and motion; third, there is a lack of task-structure-based targeted augmentation mechanisms, and random perturbations cannot focus on critical moments that determine task success or failure, such as grasp establishment, load transfer, and release completion; fourth, there is a lack of quantitative evaluation of the training value of generated samples, only verifying "whether generation is possible" while ignoring "whether generation truly benefits model training." Therefore, existing technologies have not yet achieved a complete closed loop of task semantic-guided targeted augmentation of key nodes, cross-modal synchronous generation, and training gain selection. Summary of the Invention
[0003] To address the existing problems of insufficient utilization of multi-source data, inconsistent generation results across modalities, lack of task-structure-based targeted augmentation mechanisms, and lack of training value evaluation, this paper presents an embodied intelligent multimodal data synthesis method, device, medium, and product that aims to achieve task semantic-guided targeted augmentation of key nodes, synchronous generation of multimodal trajectories and images, and automatic selection of training gains.
[0004] This application provides a method for synthesizing embodied intelligence multimodal data, including:
[0005] Based on task semantics and multimodal dynamic features, a multi-level task phase decomposition is performed on the task execution process to extract key nodes that have a decisive impact on the success or failure of the task, and a task causal skeleton is constructed; the task causal skeleton defines the range of immutable constraints and variable factors in the task execution process.
[0006] Under the constraints of the task causal framework, controlled counterfactual perturbations are applied to the key nodes to generate smooth and executable new robot motion trajectories and state trajectories.
[0007] Based on the new motion trajectory and historical context information, a new multi-view color video and a new multi-view depth video are simultaneously generated that are aligned with the new motion trajectory in terms of time sequence and three-dimensional geometry.
[0008] The generated samples are subjected to cross-modal consistency verification and training gain evaluation, and high-value synthetic multimodal data are selected and output.
[0009] Optionally, the multi-level task phase decomposition of the task execution process includes:
[0010] The task execution process is divided into first-level task phases with macro-level semantics;
[0011] Based on the relative pose of the target object and the end effector, the rate of change of the gripper state, and the characteristics of the robot load change, the primary task phase is further divided into secondary micro-phases with fine-grained motion semantics.
[0012] Optionally, the immutable constraints include: the target object and the end effector remain relatively stationary after the grasping is completed, and the target object is located within the effective support area after the placement is completed;
[0013] The variable factors include: the approach direction of the end effector, the spatial offset of the target placement point, the timing of the gripper closing or opening, and the change in the local observation perspective.
[0014] Optionally, under the constraints of the task causal framework, applying controlled counterfactual perturbations to the key nodes to generate smooth and executable new robot motion trajectories and state trajectories includes:
[0015] Define a counterfactual parameter vector that includes end-effector position perturbation, end-effector attitude perturbation, gripper timing offset, target point offset, and velocity scaling factor;
[0016] Within the local time window corresponding to the key node, a new local end trajectory is generated based on the counterfactual parameter vector. A spline interpolation method based on minimum jerk constraints is used to smoothly connect the new local motion trajectory with the preceding and following original trajectory segments. The new joint trajectory is obtained by solving the inverse kinematics mapping.
[0017] Optionally, the step of simultaneously generating a new multi-view color video and a new multi-view depth video that are temporally and geometrically aligned with the new motion trajectory based on the new motion trajectory and historical context information includes:
[0018] The projection positions of 3D scene points on the image planes of each viewpoint are constrained to satisfy the multi-view geometric projection consistency constraint.
[0019] During the generation process, the content consistency of the static background area is maintained, and the dynamic area affected by the action is migrated and repaired based on structure preservation.
[0020] Optionally, the training gain evaluation is based on key node sensitivity, sample novelty, boundary coverage capability, and visual robustness, and introduces a penalty term for inconsistency between action and image.
[0021] The sensitivity of key nodes is obtained by measuring the marginal impact of local action perturbations on the task results;
[0022] Samples are filtered and retained as synthetic output only when their gain score is higher than a preset threshold.
[0023] Optionally, the synthetic multimodal data further includes structured annotation information, which includes at least: the start and end frame range of the action phase, the location of key nodes, the configuration of applied counterfactual parameters, and the training gain score of the samples.
[0024] This application provides an electronic device, the electronic device comprising:
[0025] One or more processors; and a memory storing computer program instructions that, when executed, cause the processors to perform the steps of the method described above.
[0026] This application also provides a computer-readable medium having a computer program / instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.
[0027] This application also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0028] The beneficial effects of the above technical solution are as follows:
[0029] This application's embodied intelligence multimodal data synthesis method achieves precise localization of decisive moments in the task execution structure through multi-level task phase decomposition and key node extraction. Based on this, a task causal skeleton is constructed, clearly defining the range of immutable constraints and variable factors, providing structured task semantic guidance for subsequent targeted augmentation and significantly improving the relevance of data augmentation. By applying controlled counterfactual perturbations to key nodes under causal skeleton constraints, new robot motion trajectories, state trajectories, and multi-view color and depth videos are generated simultaneously, ensuring strict temporal and 3D geometric alignment between visual and motion information, solving the cross-modal inconsistency problem in traditional methods. A training gain evaluation mechanism based on sensitivity, novelty, boundary coverage, and visual robustness is introduced, combined with cross-modal consistency verification for sample screening, automatically eliminating low-quality samples and ensuring that the output data has a real gain for embodied intelligence model training. The final synthesized data maintains the same organizational form as the real-world data and can be directly used for training models such as imitation learning and policy learning, improving the utilization efficiency and practical value of multimodal data. Attached Figure Description
[0030] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0031] Figure 1 This is a schematic flowchart of one embodiment of the embodied intelligence multimodal data synthesis method described in this application;
[0032] Figure 2 This is a schematic flowchart illustrating an embodiment of the method for multi-level task phase decomposition in the task execution process according to this application.
[0033] Figure 3 A schematic flowchart illustrating a method for generating new multi-view color video and new multi-view depth video according to an embodiment of this application;
[0034] Figure 4 This is an exemplary structural diagram of the electronic device of this application. Detailed Implementation
[0035] The advantages of this application are further illustrated below with reference to the accompanying drawings and specific embodiments.
[0036] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0037] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0038] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0039] In the description of this application, it should be understood that the numerical labels before the steps do not indicate the order of the steps, but are only used to facilitate the description of this application and to distinguish each step, and therefore should not be construed as a limitation of this application.
[0040] The embodied intelligent multimodal data synthesis method of this application embodiment can run on robot platforms equipped with multi-view image acquisition devices, simulation environment workstations, cloud training servers, edge computing devices, desktop computers, industrial control computers, and other computing terminals with data processing capabilities. The devices transmit data and work collaboratively with each other through wired or wireless networks.
[0041] The method described in this application can be applied not only to desktop robotic tasks such as folding clothes, pouring coffee, and assembling parts, but also to any embodied intelligence application scenario that requires expanding multimodal training samples from a small amount of real-world demonstration data. For example, it can be applied to logistics sorting scenarios, medical surgical robot operations, and the daily household chores performed by home service robots. This application uses the application of the method to robot grasping-moving-placing tasks as an example for illustration, but it is not limited to this.
[0042] In this embodiment, image acquisition devices such as cameras on the robot's head, left wrist, and right wrist transmit their acquired color and depth videos, along with robot motion and state data, to a data processing server via a network. After receiving the multimodal raw data, the data processing server performs a series of processes, including task phase decomposition, key node extraction, counterfactual trajectory generation, multi-view image generation, and gain filtering, ultimately outputting a synthetic dataset. This synthetic dataset can be directly distributed to various training nodes for downstream embodied intelligent model imitation learning or policy learning training. In this embodiment, the data processing server can be a cloud server or a local server, deployed either inside or outside the robot platform. This example uses a single data processing server; in actual deployment, it can also include a computing cluster composed of multiple communicating servers to process multimodal data synthesis requests for multiple tasks in parallel.
[0043] This application provides an embodied intelligent multimodal data synthesis method, realizing an automated closed-loop generation process from multimodal raw data of a real robot's single task execution to new RGB video, new depth video, new robot motion trajectory and state trajectory data, and new annotation information output. The overall process of the method is as follows: input multimodal raw data → unified encoding and spatiotemporal alignment of raw data → task phase decomposition → key node extraction and causal skeleton construction → local counterfactual trajectory generation → multi-view RGB-D image prediction and migration generation → motion-image consistency optimization and gain filtering → output synthesis results. The input data includes: three color videos and corresponding depth videos acquired by the head camera, left wrist camera, and right wrist camera; timestamp information for each video; intrinsic and extrinsic parameter information of the three cameras; robot motion data and state data and their timestamp information; and a task annotation file containing task text description, scene information, and motion stage division information. The output data includes: the synthesized three color videos, three depth videos aligned frame-by-frame with the color videos, the synthesized robot motion trajectory file and state trajectory file, and structured stage annotation file and key node annotation file.
[0044] This application proposes an embodied intelligent multimodal data synthesis method to address shortcomings in existing technologies, such as insufficient utilization of multi-source data, inconsistencies in generated results across modalities, lack of task-structure-based targeted augmentation mechanisms, and lack of quantitative evaluation of the training value of generated samples. (See attached specification.) Figure 1 This is a flowchart illustrating a preferred embodiment of the embodied intelligent multimodal data synthesis method according to this application. As can be seen from the figure, the embodied intelligent multimodal data synthesis method provided in this embodiment mainly includes the following steps:
[0045] S1. Based on task semantics and multimodal dynamic features, perform multi-level task phase decomposition on the task execution process, extract key nodes that have a decisive impact on the success or failure of the task, and construct a task causal skeleton; the task causal skeleton defines the range of immutable constraints and variable factors in the task execution process.
[0046] S2. Under the constraints of the task causal framework, a controlled counterfactual perturbation is applied to the key nodes to generate a smooth and executable new robot motion trajectory and state trajectory;
[0047] S3. Based on the new motion trajectory and historical context information, simultaneously generate a new multi-view color video and a new multi-view depth video that are aligned with the new motion trajectory in terms of time sequence and three-dimensional geometry;
[0048] S4. Perform cross-modal consistency verification and training gain evaluation on the generated samples, and screen and output high-value synthetic multimodal data.
[0049] In this embodiment, through multi-level task phase decomposition and key node extraction, precise localization of decisive moments in the task execution structure is achieved. Based on this, a task causal skeleton is constructed, clearly defining the range of immutable constraints and variable factors, providing structured task semantic guidance for subsequent targeted augmentation and significantly improving the relevance of data augmentation. By applying controlled counterfactual perturbations to key nodes under the constraints of the causal skeleton, new robot motion trajectories, state trajectories, and multi-view color and depth videos are generated simultaneously, ensuring strict temporal and 3D geometric alignment between visual and motion information, solving the cross-modal inconsistency problem in traditional methods. A training gain evaluation mechanism based on sensitivity, novelty, boundary coverage, and visual robustness is introduced, combined with cross-modal consistency verification for sample screening, automatically eliminating low-quality samples and ensuring that the output data has a real gain for training embodied intelligence models. The final output synthetic data maintains the same organizational form as the real-world data and can be directly used for training models such as imitation learning and policy learning, improving the utilization efficiency and practical value of multimodal data.
[0050] In an alternative embodiment, such as Figure 2 As shown, step S1, which performs multi-level task phase decomposition on the task execution process, may include the following steps:
[0051] S11. Divide the task execution process into first-level task phases with macro-semantic meaning; for example, decompose the complete grab-move-place task into coarse-grained stages such as approach, grab, lift, transport, descent, release, and evacuation.
[0052] S12. Based on the relative pose of the target object and the end effector, the rate of change of the gripper state, and the characteristics of the robot load change, the primary task phase is further divided into secondary micro-phases with fine-grained motion semantics. For example, the grasping phase is further subdivided into sub-phases such as pre-alignment, slow approach, contact establishment, gripper closure, and stabilization confirmation, and the release phase is further subdivided into sub-phases such as approaching the support surface, contact confirmation, gripper opening, and disengagement.
[0053] In this embodiment, the continuous task execution process is transformed into a hierarchical semantic boundary structure through the aforementioned two-level phase decomposition, making the task logic originally implicit in the time-series data explicit and operable. Compared to traditional methods that indiscriminately perturb the entire trajectory, this multi-level decomposition mechanism provides a more refined time boundary for the precise positioning of subsequent key nodes. This allows data augmentation to accurately focus on local intervals that have a decisive impact on the success or failure of the task, such as capture establishment, carrier transfer, and release completion. It effectively avoids ineffective augmentation in irrelevant or redundant segments of task execution, thereby significantly improving the training value and augmentation efficiency of the synthetic data.
[0054] In this embodiment, step S12 performs secondary micro-phase partitioning based on the relative pose of the target object and the end effector, the rate of change of the gripper state, and the characteristics of the robot load change. Specifically, this is achieved by constructing a multimodal feature vector. Achieve as shown in formula (1):
[0055] (1),
[0056] in, , and Representing time respectively The distance between the target and the end effector, the end effector velocity, and the end effector acceleration; and Representing time respectively The opening and closing amount of the lower gripper and its rate of change; , and Representing time respectively The height of the target object relative to the support surface, the contact state between the target object and the support surface, and the change in the robot's drive load.
[0057] The above features are obtained as follows: Based on multi-view color images, depth images, and camera calibration parameters, the spatial positions of the target object and the end effector in a unified coordinate system are reconstructed, and the center position of the target object is denoted as... The position of the end effector is denoted as The distance between the target and the end effector is as shown in formula (2):
[0058] (2),
[0059] in, Let L2 represent the vector norm. The end effector velocity and end effector acceleration are obtained from the first and second differences of the end effector position with respect to time, as shown in equations (3) and (4):
[0060] (3),
[0061] (4),
[0062] Gripper opening / closing amount The gripper change rate is read directly from the gripper status data. It is calculated from the difference in the opening and closing amount of the grippers at adjacent moments, as shown in formula (5):
[0063] (5),
[0064] The height of the target object relative to the supporting surface It is calculated from the normal distance between the center of the target object and the supporting plane. If the equation of the supporting plane is expressed as... Then we have formula (6):
[0065] (6),
[0066] in, This represents the normal vector of the supporting plane. This indicates the offset of the supporting plane. It also represents the contact state between the target object and the supporting surface. The change in robot-driven load is determined by comparing the distance from the bottom region of the target object to the supporting plane with a preset threshold. It is obtained from the difference in driving load values at adjacent times, as shown in formula (7):
[0067] (7),
[0068] in, Indicates time The system calculates the robot drive load value. Based on the aforementioned feature vectors, it performs abrupt change detection and boundary delineation on the temporal features within each primary task phase to obtain a set of secondary micro-phases.
[0069] In an optional embodiment, step S1 involves extracting key nodes that have a decisive impact on the success or failure of the task and constructing a task causal skeleton, including: firstly, obtaining a set of candidate key nodes based on the second-order micro-phase boundary and multimodal change characteristics. The candidate key nodes include the grab-and-establish node, the detachment from support node, the stable transport node, the load transfer node, and the release completion node.
[0070] Then for each candidate key node Perform local sensitivity assessment: Set nodes Nearby local action segments are Within this local action segment, small perturbations are applied to the end position, end pose, and gripper action timing, and the changes in the task result score are observed. The node sensitivity is defined as in formula (8):
[0071] (8),
[0072] in, Indicates candidate key nodes Sensitivity; Represents a node Nearby local action segments; Indicates the first Sub-local perturbation; This represents a function representing the result of a local task. This represents the total number of local perturbations. When... If the value exceeds a preset threshold, the node is retained as the final key node, resulting in a set of key nodes. .
[0073] Subsequently, using the key nodes as graph nodes and the local action segments between the key nodes as graph edges, a task causal skeleton graph is constructed, as shown in formula (9):
[0074] (9),
[0075] in, Represents the causal skeleton diagram of the task; Represents the set of graph nodes and ; Represents the set of transition edges between key nodes; Represents the set of immutable constraints; This represents the set of variable factors. The set of immutable constraints describes the conditions that must be maintained during task execution, while the set of variable factors describes the factors that are allowed to change without altering the task semantics and objectives.
[0076] Through the aforementioned multi-level task phase decomposition, key node extraction, and task causal skeleton construction, the system obtains a structured semantic representation of task execution and clear constraint boundaries, providing a precise operational range and safe perturbation scope for subsequent targeted counterfactual augmentation. Based on this, the system, through the explicit definition of immutable constraints and variable factors, restricts counterfactual augmentation to a safe perturbation space with physical meaning and task semantic guarantees.
[0077] In an optional embodiment, the immutable constraints include: the target object and the end effector remain relatively stationary after grasping (i.e., maintain a stable follow-up relationship), and the target object is located within the effective support area after placement.
[0078] The variable factors include: the approach direction of the end effector, the spatial offset of the target placement point, the timing of the gripper closing or opening, and the change in the local observation perspective.
[0079] Among them, immutable constraints define the physical and logical conditions that must always be met during task execution, and violations will directly lead to task failure; variable factors define the dimensions of operational parameters that can be adjusted while ensuring that the semantics of the task remain unchanged.
[0080] In this embodiment, by clearly defining immutable constraints and variable factors, counterfactual amplification is confined to a safe perturbation space with physical meaning and task semantic guarantees. When generating new motion trajectories, the system only applies perturbations to variable factors such as approach direction, placement point offset, gripper timing, and velocity scaling. Simultaneously, constraint checks ensure that the stable follow-up relationship after grasping and the effective support relationship after placement are not disrupted. This bounded perturbation mechanism allows the generated synthetic samples to cover diverse execution methods without producing physically infeasible or semantically distorted invalid data. This effectively solves the problems of inconsistent sample quality and a large number of samples not conforming to task logic in traditional random perturbation methods, thereby ensuring the effectiveness of the output data for downstream policy learning.
[0081] In an optional embodiment, step S2, under the constraints of the task causal framework, applies controlled counterfactual perturbations to the key nodes to generate smooth and executable new robot motion trajectories and state trajectories, including:
[0082] For each key node The corresponding local time window is defined as shown in formula (10):
[0083] (10)
[0084] in, Indicates the first Local time windows corresponding to key nodes; Indicates the first A key node; and These represent the lengths of the time windows before and after the critical node, respectively.
[0085] Definition includes end position perturbation Terminal attitude perturbation Grip timing offset (i.e., the timing offset of gripper closure or opening), target point offset) (i.e., target placement point or proximity point offset) and (local) velocity scaling factor The counterfactual parameter vector is shown in formula (11):
[0086] (11),
[0087] Within the local time window corresponding to the key node, a new local end trajectory is generated based on the counterfactual parameter vector. A spline interpolation method based on minimum jerk constraints is used to smoothly connect the new local motion trajectory with the preceding and following original trajectory segments. The new joint trajectory is obtained by solving the inverse kinematics mapping.
[0088] Specifically, within the local time window corresponding to the key node, a new local terminal trajectory is generated based on the counterfactual parameter vector, and the new local terminal trajectory is represented as shown in formula (12):
[0089] (12),
[0090] in, Indicates the first The local new terminal trajectory corresponding to each key node; This represents the original local end trajectory near the critical node; Represented by the counterfactual parameter vector The driving continuous perturbation function.
[0091] A spline interpolation method based on minimum jerk constraints is used to smoothly connect the local new motion trajectory with the preceding and following original trajectory segments to obtain a new complete end trajectory. The new joint trajectory is obtained by solving the inverse kinematics mapping as shown in formula (13):
[0092] (13)
[0093] in, Indicates time The new joint trajectory below; Indicates time The new terminal trajectory below; This represents the inverse kinematics solution operation. Based on the new joint trajectories and the new end-effector trajectories, new robot motion trajectories are generated synchronously. With state trajectory .
[0094] In this embodiment, through the multidimensional definition of the counterfactual parameter vector, the local perturbation around the key node is precisely decomposed into independent and controllable dimensions such as end-effector pose, gripper timing, target point position, and execution speed. This allows each augmentation to be directed to specific operational features under clear task semantic guidance, rather than blindly and randomly perturbing the entire trajectory globally. A spline interpolation method based on minimum jerk constraints is used to smoothly connect the local new motion trajectory with the original trajectory segment, effectively ensuring the continuity of velocity and acceleration at the connection point of the generated trajectory. This avoids the velocity jumps and joint impacts common in traditional trajectory perturbation methods, thus ensuring that the generated new motion trajectory is physically executable and the motion process is smooth and fluid. Based on this, the end-effector trajectory is transformed into a joint space trajectory through inverse kinematic mapping, enabling the generated result to be directly deployed to the actual robot for execution, achieving a complete connection from task-level semantic perturbation to joint-level executable instructions.
[0095] In an optional embodiment, step S3, based on the new motion trajectory and historical context information, synchronously generates a new multi-view color video and a new multi-view depth video that are temporally and geometrically aligned with the new motion trajectory, such as... Figure 3 As shown, the following steps may be included:
[0096] S31. Constrain the projection positions of 3D scene points on the image planes of each viewpoint to satisfy the multi-view geometric projection consistency constraint;
[0097] S32. Maintain the consistency of content in static background areas during the generation process, and perform structure-preserving migration and repair on dynamic areas affected by actions.
[0098] In this embodiment, the multi-view geometric projection consistency constraint in step S31 ensures that the projection position of the same 3D scene point in the images generated by the head camera, left wrist camera, and right wrist camera strictly follows the actual intrinsic and extrinsic parameter relationships between each camera. This ensures that the three generated videos maintain spatial consistency in geometric structure, avoiding problems such as object misalignment, depth inconsistencies, or inconsistent occlusion relationships between multi-view images, effectively guaranteeing the visual realism and cross-view spatial coherence of the generated data. Through the coordinated processing of static background preservation and dynamic region structure preservation migration repair in step S32, only the local regions that have changed due to the new motion trajectory are directionally generated and texture migrated, while the unrelated static background regions directly retain the original image content. This significantly reduces the computational cost of image generation and avoids artifacts such as background distortion, texture blurring, and temporal flicker that are prone to occur in traditional full-frame generation methods.
[0099] In this embodiment, step S3 simultaneously generates new multi-view color video and new multi-view depth video. Specifically, it uses the three-view RGB-D observations within the previous time window, the unified 3D scene point cloud, the new local trajectories within the key node time window, and the task stage labels as input conditions, and utilizes a multi-view video prediction and transfer generation model. Predicting the future after key milestones The new frame sequence is as shown in formula (14):
[0100] (14)
[0101] in, Indicates time to The predicted sequence of images; Indicates time to Historical RGB-D observation sequence; Indicates time to A sequence of historical 3D scene point clouds; Indicates time to New motion trajectory conditions; Indicates the current time Task stage tags; Indicates the length of the historical time window; This indicates the number of frames predicted forward. The generated results include new color video clips from the head camera, the left wrist camera, and the right wrist camera, as well as three new depth video clips corresponding to each of the three color video clips frame by frame.
[0102] The multi-view geometric projection consistency constraint is specifically defined as follows: for any three-dimensional point in a unified three-dimensional scene... Its projection coordinates on the image planes of the head camera, left wrist camera, and right wrist camera satisfy the following projection equation (15):
[0103] (15)
[0104] in, Representing a three-dimensional point In the Pixel coordinates on the camera image plane; Indicates the first The intrinsic parameter matrix of each camera; Indicates the distance from the robot's base coordinate system to the... The extrinsic transformation matrix of each camera coordinate system; symbol This indicates that both sides are equal in the sense of homogeneous coordinates.
[0105] The combination of these two mechanisms ensures that the generated multi-view images are not only strictly aligned with the new action trajectories in terms of time sequence, but also meet the high-quality requirements of embodied intelligence models for training data in terms of geometric consistency, texture realism, and temporal coherence. Meanwhile, structure-preserving local migration and restoration ensure that dynamic changes in the image have realistic textures and lighting variations, avoiding artifacts such as background distortion and temporal flicker common in traditional methods.
[0106] In an optional embodiment, the training gain evaluation is based on key node sensitivity, sample novelty, boundary coverage, and visual robustness, and introduces a penalty term for inconsistency between action and image.
[0107] The sensitivity of key nodes is obtained by measuring the marginal impact of local action perturbations on the task results;
[0108] Samples are filtered and retained as synthetic output only when their gain score is higher than a preset threshold.
[0109] In this embodiment, a multi-dimensional sample gain evaluation system is constructed, expanding the data selection criteria from a single visual quality indicator in traditional methods to a comprehensive evaluation mechanism that integrates task sensitivity, distribution novelty, boundary coverage capability, visual robustness, and cross-modal consistency penalty. Specifically, key node sensitivity quantifies the impact of the operational interval corresponding to the sample on task success or failure, prioritizing samples involved in critical decision-making moments; sample novelty and boundary coverage capability ensure that the synthetic data can effectively supplement sparse regions and unseen boundary situations in the original data distribution; visual robustness measures the generation quality of the image under different viewpoints and lighting conditions; and the action-image inconsistency penalty automatically eliminates low-quality samples with acceptable visual generation quality but geometrical inconsistencies between the trajectory and the image. Through the above multi-dimensional evaluation and threshold selection, this embodiment achieves a qualitative leap from "whether it can be generated" to "whether the generated data truly has a gain for model training," ensuring that every sample in the final output synthetic dataset has quantifiable positive training value for downstream policy learning, effectively avoiding the occupation of training resources by invalid or low-quality data.
[0110] In this embodiment, the cross-modal consistency verification also includes semantic consistency verification: the system combines the task name and action phase configuration in the task annotation file to check the consistency between the execution phase order and semantics of the generated samples. The generated samples are required to maintain the same action phase structure and phase order as the original task to ensure that counterfactual amplification does not destroy the semantics of the original task. Samples that pass the above verification proceed to the subsequent training gain evaluation stage.
[0111] Wherein, the sample gain score The calculation method is as shown in formula (16):
[0112] (16)
[0113] in, Indicates the sensitivity of critical nodes; Indicates the sample novelty score; This indicates the ability to cover situations where no boundaries are visible; Indicates visual robustness gain; This indicates a penalty for inconsistency between action and visuals; to These are the weighting coefficients for each item. The action-screen inconsistency penalty item... It is composed of a weighted average of RGB consistency loss, depth consistency loss, and trajectory consistency loss, as shown in formula (17):
[0114] (17)
[0115] in, Indicates the generation of RGB images Geometric image obtained by projecting action state Consistency loss between them; Indicates the generation of depth image Geometric depth map obtained by projecting action state Consistency loss between them; Indicates the generation of motion trajectory Trajectory obtained by visual reconstruction Consistency loss between them; , and These are the weighting coefficients for each type of loss.
[0116] Only when the sample gain score Higher than the preset threshold At that time, samples are selected and retained as the synthetic output.
[0117] In an optional embodiment, the synthetic multimodal data further includes structured annotation information, which includes at least: the start and end frame range of the action phase, the location of key nodes, the configuration of applied counterfactual parameters, and the training gain score of the samples.
[0118] In this embodiment, by synchronously generating structured annotation information in the synthetic data output, each synthetic sample not only contains visual images and action trajectories but also carries complete task semantic metadata. Specifically, the start and end frame ranges of the action phases accurately map the timeline of the task execution process to semantic phase labels, facilitating downstream models to perform phased course learning or phased loss weighting during training. Key node positions clearly identify the frame indices of decisive moments such as grasping establishment, payload transfer, and release completion within the sample, enabling the model to specifically strengthen the supervisory signal intensity of these key frames. The applied counterfactual parameter configuration fully records the specific parameters used when generating the sample, such as perturbation dimension, perturbation amplitude, and velocity scaling, providing a traceable data foundation for subsequent analysis of the impact of different amplification strategies on model performance. The training gain score of the sample serves as a quantitative label of the sample's expected training value, supporting weighted sampling or course sorting based on the score during training, prioritizing the use of high-value samples for learning. This method of encapsulating and outputting visual information, action information, and structured semantic annotation in an integrated manner makes the synthetic data completely consistent with the carefully manually annotated real demonstration data in terms of organization. It can be directly connected to existing training pipelines such as imitation learning and policy learning, which greatly reduces the preprocessing cost for data users and improves the usability and engineering practicality of synthetic datasets.
[0119] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed in this application; however, this does not mean that other units are absent in this embodiment.
[0120] Furthermore, some embodiments of this application also provide an electronic device. The electronic device can be various forms of digital computer, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device can also be various forms of mobile devices, such as cellular phones, smartphones, wearable devices, and other similar computing devices.
[0121] The electronic device includes: one or more processors; and a memory storing computer program instructions that, when executed, cause the processor to perform the steps of the methods provided in any one or more of the above embodiments. Figure 4 An exemplary structural diagram of the electronic device is disclosed. The electronic device includes one or more processors 1101, a memory 1102, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0122] The electronic device may further include an input device 1103 and an output device 1104. The processor 1101, memory 1102, input device 1103 and output device 1104 may be connected by a bus or other means, as shown in the figure, which is connected by a bus.
[0123] Input device 1103 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 1104 may include a display device, auxiliary lighting device (e.g., LED), and haptic feedback device (e.g., vibration motor). The display device may include, but is not limited to, a liquid crystal display, a light-emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.
[0124] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device (e.g., a cathode ray tube or LCD monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback); and input from the user can be received in any form (e.g., voice input or tactile input).
[0125] In this embodiment, a computer-readable medium stores a computer program / instructions that, when executed by a processor, implement the steps of the methods provided in any one or more of the above embodiments. This computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into that device. The aforementioned computer-readable medium carries one or more computer-readable instructions.
[0126] The memory 1102 can serve as a non-transitory computer-readable storage medium, used to store non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, thereby implementing the program instructions / modules corresponding to the methods provided in any one or more of the embodiments described above in this application.
[0127] The memory 1102 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 1102 may optionally include memory remotely located relative to the processor 1101, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0128] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0129] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technologies, read-only optical discs, digital versatile optical discs or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0130] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0131] In the above embodiments, all or part of the implementation can be achieved through software, hardware, firmware, or any combination thereof. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drives, floppy disks, and similar devices. In addition, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.
[0132] The computer program product provided in this application includes one or more computer programs / instructions. When executed by a processor, these computer programs / instructions generate, in whole or in part, the processes or functions described in this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0133] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0134] The scope of this application is defined by the appended claims rather than the foregoing description, and is therefore intended to encompass all variations falling within the meaning and scope of equivalents of the claims. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device in software or hardware. Terms such as "first," "second," etc., are used only for distinguishing descriptions and do not indicate any particular order, nor should they be construed as indicating or implying relative importance.
[0135] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily made by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.
Claims
1. A method for synthesizing embodied intelligent multimodal data, characterized in that, include: Based on task semantics and multimodal dynamic features, a multi-level task phase decomposition is performed on the task execution process to extract key nodes that have a decisive impact on the success or failure of the task, and a task causal skeleton is constructed; the task causal skeleton defines the range of immutable constraints and variable factors in the task execution process. Under the constraints of the task causal framework, controlled counterfactual perturbations are applied to the key nodes to generate smooth and executable new robot motion trajectories and state trajectories. Based on the new motion trajectory and historical context information, a new multi-view color video and a new multi-view depth video are simultaneously generated that are aligned with the new motion trajectory in terms of time sequence and three-dimensional geometry. The generated samples are subjected to cross-modal consistency verification and training gain evaluation, and high-value synthetic multimodal data are selected and output.
2. The embodied intelligent multimodal data synthesis method according to claim 1, characterized in that, The multi-level task phase decomposition of the task execution process includes: The task execution process is divided into first-level task phases with macro-level semantics; Based on the relative pose of the target object and the end effector, the rate of change of the gripper state, and the characteristics of the robot load change, the primary task phase is further divided into secondary micro-phases with fine-grained motion semantics.
3. The embodied intelligent multimodal data synthesis method according to claim 1, characterized in that, The immutable constraints include: after the grasping is completed, the target object and the end effector remain relatively stationary; after the placement is completed, the target object is located within the effective support area. The variable factors include: the approach direction of the end effector, the spatial offset of the target placement point, the timing of the gripper closing or opening, and the change in the local observation perspective.
4. The embodied intelligent multimodal data synthesis method according to claim 1, characterized in that, Under the constraints of the task causal framework, controlled counterfactual perturbations are applied to the key nodes to generate smooth and executable new robot motion trajectories and state trajectories, including: Define a counterfactual parameter vector that includes end-effector position perturbation, end-effector attitude perturbation, gripper timing offset, target point offset, and velocity scaling factor; Within the local time window corresponding to the key node, a new local end trajectory is generated based on the counterfactual parameter vector. A spline interpolation method based on minimum jerk constraints is used to smoothly connect the new local motion trajectory with the preceding and following original trajectory segments. The new joint trajectory is obtained by solving the inverse kinematics mapping.
5. The embodied intelligent multimodal data synthesis method according to claim 1, characterized in that, The process of simultaneously generating new multi-view color videos and new multi-view depth videos that are temporally and geometrically aligned with the new motion trajectory based on the new motion trajectory and historical context information includes: The projection positions of 3D scene points on the image planes of each viewpoint are constrained to satisfy the multi-view geometric projection consistency constraint. During the generation process, the content consistency of the static background area is maintained, and the dynamic area affected by the action is migrated and repaired based on structure preservation.
6. The embodied intelligent multimodal data synthesis method according to claim 1, characterized in that, The training gain evaluation is based on key node sensitivity, sample novelty, boundary coverage, and visual robustness, and introduces a penalty term for inconsistency between action and image. The sensitivity of key nodes is obtained by measuring the marginal impact of local action perturbations on the task results; Samples are filtered and retained as synthetic output only when their gain score is higher than a preset threshold.
7. The embodied intelligent multimodal data synthesis method according to claim 1, characterized in that, The synthetic multimodal data also includes structured annotation information, which includes at least: the start and end frame range of the action phase, the location of key nodes, the configuration of applied counterfactual parameters, and the training gain score of the samples.
8. An electronic device, characterized in that, The electronic device includes: One or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method as described in any one of claims 1 to 7.
9. A computer-readable medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 7.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 7.