Intelligent data synthesis and labeling method, equipment, medium and product
By decomposing embodied intelligence tasks into stage semantic sequences through a multi-model collaborative mechanism, diverse scene images are generated and the robot control trajectory is inferred. This solves the problems of single modality, lack of scene diversity, and inconsistency in time sequence in embodied intelligence data generation, realizes an efficient and structured data generation process, and improves the model's generalization ability and training efficiency.
Patent Information
- Application Number
- CN202511756105.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-11-27
AI Technical Summary
Existing embodied intelligence data generation methods suffer from problems such as single input modality, lack of data scene diversity, lack of automatic semantic annotation, and inconsistency between video and action temporal/physical sequence, which limits the generalization ability and robustness of trained models in complex environments.
A multi-model collaborative mechanism is adopted. The first major model decomposes the embodied intelligence task into a semantic sequence of task stages with action type labels. The second major model generates a set of scene images with different appearances. The third major model generates a color video. The robot's control trajectory data is inferred through inverse dynamics to generate a structured semantic annotation file.
It achieves highly diverse and precisely spatially constrained embodied intelligent data generation, ensuring strict physical and temporal consistency between videos and motion trajectories, reducing the cost of manual data collection and fine annotation, and improving the model's environmental generalization ability and training efficiency.
Smart Images

Figure CN121217907A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information technology, in particular to an embodied intelligent data synthesis and labeling method, device, medium and product. BACKGROUND
[0002] Embodied AI is a cutting-edge direction in the field of artificial intelligence, aiming to endow agents with the ability to perceive, understand and execute complex tasks in real or simulated environments. The training of such agents has very high requirements for the quality and scale of data, and its core lies in the need for large-scale, multi-modal demonstration data, which usually includes visual information (RGB), language instructions and precise robot action trajectories. However, the primary challenge faced by current technology is the high cost of data collection and labeling, especially when precise spatial action positioning, interaction phase division and fine-grained semantic labeling are required, traditional manual methods are inefficient and difficult to scale.
[0003] However, the existing embodied intelligent data of the prior art mainly relies on real-world collection or traditional simulation generation. Real-world collection is costly, inefficient, and difficult to cover diverse scenarios and extreme cases. Although traditional simulators can generate large-scale data, the generated videos often lack realism and diversity, and the consistency and physical reproducibility of control trajectory data and video timing are difficult to guarantee, resulting in a huge "inter-domain gap" when the agent migrates from simulation to reality.
[0004] However, the existing embodied data generation methods can generate action videos and trajectories based on language or images, but still have the following shortcomings:
[0005] First, the limitation of data modalities. Most generation methods rely only on RGB video as input or output, which makes the generated actions lack perception and constraints on scene depth information and three-dimensional structure. Due to the lack of accurate spatial information, the action trajectories derived by the model are often inaccurate in the actual physical world, making it difficult to ensure the reliability of the operation.
[0006] Second, the singularity of the scene and environment. Existing methods are usually limited to a single or a few environmental samples when generating videos and trajectories, making it difficult to effectively cover the complex and variable lighting conditions and background changes in the real world. This lack of scene diversity severely limits the generalization ability and robustness of the trained model when facing unknown environments.
[0007] Thirdly, the lack of automated annotation. Existing data generation pipelines usually output videos and control trajectories as the main outputs, but lack a mechanism to automatically and structurally align high-level task semantic sequences (e.g., task stages, action types) with the underlying time axis and control signals. This results in data that cannot be directly used for high-level embodied intelligence model training that requires fine-grained semantic structure.
[0008] Finally, the lack of physical and timing consistency. This is a critical flaw. Whether the data is obtained through simulation or generative models, there is often a timing deviation or physical mismatch between the output video and the control instructions derived through inverse dynamics. This inconsistency makes the trajectory data lack physical reproducibility, making it difficult to be directly used for high-performance simulation or robot control, greatly limiting the practical value of the data.
[0009] Therefore, there is an urgent need in the industry for a new and efficient method that can generate structured embodied intelligence data with high diversity, precise spatial constraints, and strict physical and timing consistency between videos and action trajectories using multi-modal input. SUMMARY
[0010] An object of the present application is to provide an embodied intelligence data synthesis and annotation method, device, medium and product, at least to solve the technical defects of the existing embodied intelligence data synthesis method, such as single input modality, lack of diversity in data scenarios, lack of automatic semantic annotation, and inconsistency between video and action timing / physics. To achieve the above object, some embodiments of the present application provide the following aspects:
[0011] The present application provides an embodied intelligence data synthesis and annotation method, which comprises:
[0012] Receiving a natural language description of an embodied intelligence task, and corresponding initial images and depth images, decomposing the embodied intelligence task into a task stage semantic sequence with action type labels through a first large model;
[0013] Performing perturbation transformation on the initial images, and generating a set of scene images with appearance differences from the initial images through a second large model;
[0014] Based on the task stage semantic sequence and the set of scene images, generating a color video corresponding to the task stage semantic sequence through a third large model;
[0015] Performing frame-by-frame monocular depth estimation on the color video to generate a depth video aligned with the color video;
[0016] Combining the color video and the corresponding depth video, and inversely deriving robot control trajectory data corresponding to the video timing through inverse dynamics;
[0017] mapping the task stage semantic sequence to a time axis of the color video according to a duration of each stage segment of the color video, to generate a structured semantic annotation file;
[0018] outputting the color video, the depth video, the robot control trajectory data, and the structured semantic annotation file.
[0019] In a second aspect, some embodiments of the present application further provide an electronic device, comprising: one or more processors; and a memory storing computer program instructions which, when executed, cause the processor to perform the steps of the method described above.
[0020] In a third aspect, some embodiments of the present application further provide a computer readable medium having stored thereon computer program instructions executable by a processor to implement the method described above.
[0021] In a fourth aspect, some embodiments of the present application further provide a computer program product comprising computer program / instructions which, when executed by a processor, implement the steps of the method described above.
[0022] Compared with related technologies, in the scheme provided by the embodiments of the present application, the multi-model collaborative mechanism is introduced at the initial stage of data generation, which greatly enhances the diversity and generalization ability of the data. On the one hand, the strong constraint semantic structuring ability of the first large model is used to accurately decompose the complex natural language task into an executable stage sequence, providing accurate action guidance and structured basis for the entire data generation process. On the other hand, by adjusting the constrained appearance attributes of the initial scene image, a large number of scene sets are generated which maintain the original three-dimensional structure unchanged but have highly diversified appearances, which fundamentally solves the problem of single scene in traditional methods and ensures that the trained model can have strong environmental generalization ability. Then, through the depth video strictly aligned with the color video, the robot control trajectory data corresponding to the video time sequence is obtained through inverse dynamics backstepping, so that a complete four-modal data package containing color video, depth video, high-fidelity control trajectory and structured semantic annotation file can be finally automatically output. Such data has reached a very high quality standard in structure, physics and time sequence, realizing the closed loop and high efficiency of the embodied data generation process, greatly reducing the cost of manual collection and fine annotation, and providing key technical support for promoting the development of embodied intelligent models. BRIEF DESCRIPTION OF DRAWINGS
[0023] One or more embodiments are illustrated by way of example in the drawings and are described herein in connection with the embodiments described. These embodiments are described in connection with the pertinent drawing figures with like numerals indicating like elements throughout the several figures, one or more embodiments are not limited to the embodiments described, but instead can be implemented in any desired environment. The drawings are not intended to be to scale.
[0024] Figure 1 A flowchart of a method for embodied intelligent data synthesis and annotation is provided for an exemplary embodiment of the present disclosure, the method comprising:
[0025] Figure 2 An exemplary structural diagram of the electronic device is provided for some embodiments of the present application. DETAILED DESCRIPTION
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0027] Figure 1 A flowchart of a method for embodied intelligent data synthesis and annotation is provided for an exemplary embodiment of the present disclosure, the method comprising:
[0028] S101, receiving a natural language described embodied intelligent task, and a corresponding initial image and depth image, and decomposing the embodied intelligent task into a task stage semantic sequence with an action type label by a first large model;
[0029] Specifically, the purpose of this step is to decompose a complex embodied task described in natural language (for example: "put the red water cup into the cabinet") into 2-6 executable stage texts to form a structured semantic sequence. A first large model (for example, Qwen2.5-72B) can be used as an embodied task planning assistant. After receiving the task description, the model can decompose the task into 2-6 executable stages according to the given Chinese task description, and ensure the continuity of the stages. The goal of the decomposition is to enable the robot or video generation to execute the task step by step, and each stage should represent a clear action or stage goal. For example: put the red water cup into the cabinet, decompose it into approach the red water cup → grab the red water cup → move the water cup in front of the cabinet → put the water cup into the cabinet.
[0030] S102, performing perturbation transformation on the initial image, and generating a set of scene images with appearance difference from the initial image by a second large model.
[0031] Specifically, this step is used to perturb the initial image (RGB image) to generate a set of enhanced scene image collection with appearance difference from the original image. A second large model (such as Stable Diffusion and other diffusion models) can be used. The image is perturbed by the prompt word to disturb the light, color distribution, environment background, etc. The core purpose is to greatly increase the visual diversity of the data set and enhance the generalization robustness of the model while ensuring the geometric structure and object spatial relationship of the initial scene remain unchanged.
[0032] S103, based on the task phase semantic sequence and the scene image set, a color video corresponding to the task phase semantic sequence is generated by a third large model.
[0033] Specifically, this step generates a color video corresponding to the task semantic based on the semantic sequence of S101 and the scene image set of S102 using a third large model. The third large model can use a video world model (such as COSMOS) to generate, taking the task phase semantic sequence as the semantic constraint and the initial picture as the appearance starting point, to independently generate a corresponding color video for each initial picture.
[0034] S104, frame-by-frame monocular depth estimation is performed on the color video to generate a depth video aligned with the color video.
[0035] Specifically, this step is used to perform frame-by-frame depth prediction on the color video generated by S103 to generate a depth video strictly aligned with the color video. A monocular depth estimation-based model (such as Depth Anything) can be used. The key lies in the reality recovery: first, the pixels of the active object region are extracted as the alignment reference using the predicted depth video and the corresponding real depth map of the initial picture; then, the linear mapping relationship of the predicted depth and the real depth in the reference area in terms of statistical characteristics (such as mean / variance) is calculated; finally, this mapping relationship is applied to the entire predicted depth video to generate a depth video with real scale.
[0036] S105, combining the color video and the corresponding depth video, the robot control trajectory data corresponding to the video time sequence is inversely deduced by inverse dynamics.
[0037] Specifically, combining the color video of S103 and the depth video of S104, the robot control trajectory data corresponding to the video time sequence is inversely deduced by inverse dynamics. First, the color video and the depth video are processed to accurately obtain the pose changes of the target object and the end effector in the time dimension, and the pose changes are input into the inverse dynamics network (IDM-A) for control parameter inference, thereby obtaining the robot control trajectory data corresponding to the video time sequence.
[0038] S106. Mapping the task stage semantic sequence to the timeline of the color video according to the duration of each stage segment in the color video, to generate a structured semantic annotation file.
[0039] Specifically, this step is used to automatically align the stage semantic information obtained by step S101 with the video timeline generated in step S103, to form the final semantic annotation file. Since the first large model (such as Qwen2.5-72B) has output a skill sequence text containing explicit stage and action descriptions in step S101, this step does not need to regenerate semantic content, but achieves precise correspondence between semantics and video frames through time mapping and consistency correction. First, according to the actual frame number or duration of each stage video, the start and end frame interval [start_frame, end_frame] of each stage is calculated, and on this basis, the stage name, description, etc. are merged and integrated to generate a unified structured annotation.
[0040] S107. Outputting the color video, the depth video, the robot control trajectory data, and the structured semantic annotation file.
[0041] Specifically, the above data is finally output, including: task execution color video (MP4 format); depth video strictly aligned with color video (DepthAnything generated frame by frame, saved in MKV format); inverse dynamics trajectory data (timestamp, joint / end sequence, saved in.h5 format); structured semantic annotation (action / stage, bilingual description, saved in json format).
[0042] In this embodiment, by introducing a multi-model collaborative mechanism at the initial stage of data generation, the diversity and generalization ability of the data are greatly enhanced. On the one hand, the strong constraint semantic structuring ability of the first large model is used to accurately decompose complex natural language tasks into an executable stage sequence, providing precise action guidance and structured basis for the entire data generation process. On the other hand, by adjusting the constrained appearance attributes of the initial scene image, a large number of scene sets are generated that maintain the original three-dimensional structure but have highly diversified appearances, which fundamentally solves the problem of single scene in traditional methods and ensures that the trained model has strong environmental generalization ability. Then, through the depth video strictly aligned with the color video, the robot control trajectory data corresponding to the video time sequence is obtained through inverse dynamics backstepping, so that the data package containing color video, depth video, high-fidelity control trajectory and structured semantic annotation file can be finally automatically output. This kind of data has reached a very high quality standard in structure, physics and time sequence, realizing the closed loop and high efficiency of embodied data generation process, greatly reducing the cost of manual collection and fine annotation, and providing key technical support for promoting the development of embodied intelligent models.
[0043] In one embodiment, the decomposing the embodied intelligent task into a task stage semantic sequence with action type labels by the first large model specifically comprises:
[0044] The first large model decomposes the task into at least two executable stages based on the natural language description;
[0045] The task stage semantic sequence is composed of the decomposed executable stages arranged in time sequence;
[0046] The task stage semantic sequence is in a strongly constrained structured format, and the structured format includes action type labels, mother language instructions and target language instructions.
[0047] Specifically, the present embodiment aims to convert the natural language task description input by the user (for example: "Please move the green object on the table to the basket on the right") into strongly constrained structured data recognizable by machines and usable by subsequent models. The first large model first decomposes complex and continuous tasks into a series of logically clear and sequentially executable stages based on the understanding and reasoning of the natural language description. These decomposed stages, for example, may include [reach_green_object], [move_to_basket], [release], etc., arranged in time sequence according to the task, which together constitute the task stage semantic sequence.
[0048] The task stage semantic sequence is defined as a strongly constrained structured format (e.g. JSON or XML format). This structured format ensures the rigor of the data and the reliability of the automated processing. In the structured format, each stage element contains at least three key information:
[0049] Action type label (skill): selected from a predefined limited set (e.g. [reach, move, place, release]) to guide the subsequent video generation model and trajectory analysis.
[0050] Native language instruction (instruction_zh): original language or user's native language task stage description.
[0051] Target language instruction (instruction_en): translated target language (usually English) task stage description.
[0052] Further, in one embodiment, the task stage semantic sequence only outputs a JSON array and cannot contain any explanations, annotations or additional text.
[0053] Each array element is a dictionary containing the following three fields:
[0054] "skill": stage name, taken from the preset action type label set;
[0055] "instruction_zh": Chinese step-by-step instructions for this stage;
[0056] "instruction_en": English instructions equivalent to Chinese content.
[0057] All the above fields are non-empty strings, and the output must be strictly valid JSON format.
[0058] Further, in one embodiment, the first large model disassembly should keep the task semantics consistent and cannot change the original task object or target location, nor can it create new objects or new actions.
[0059] Further, in one embodiment, the skill of each stage must be selected from the following set:
[0060] ["reach", "grasp", "move", "place", "align", "insert", "release", "pick"].
[0061] For example: input task: put the red water cup in the cabinet.
[0062] Output example: [
[0064] {"skill":"reach","instruction_zh":"Approach the red cup", "instruction_en":"approach the red cup"}
[0065] {"skill":"grasp","instruction_zh":"Grasp the red cup", "instruction_en":"grasp the red cup"}
[0066] {"skill":"move","instruction_zh":"Move the cup to the cabinet", "instruction_en":"move the cup to the cabinet"}
[0067] {"skill":"place","instruction_zh":"Place the cup into the cabinet", "instruction_en":"Place the cup into the cabinet"} ]
[0069] In the above embodiments, by limiting the first major model to decompose the task based on a strongly constrained structured format, efficient, automated, and high-precision semantic annotation of embodied intelligence tasks is achieved. By decomposing the task into multiple stages arranged in a time sequence and forcing the output to be in a strict JSON format containing action type labels (skills), native language instructions, and target language instructions, the problems of high cost and low structure in semantic annotation in traditional methods are completely solved. In particular, the strict restrictions on the output format (e.g., only outputting JSON arrays, with non-empty fields) and the constraints on the semantic consistency of the task ensure that the semantic sequences relied upon by subsequent video generation and trajectory analysis have extremely high rigor, accuracy, and multilingual availability, significantly improving the automated processing efficiency of data packets and the training value of the semantic layer.
[0070] In one embodiment, the step of perturbating the initial image and generating a set of scene images with appearance differences from the initial image using a second large model specifically includes:
[0071] The second large model uses positive prompts to perform perturbation transformations on the lighting and color distribution of the initial image.
[0072] In one embodiment, the second large model is prompted by negative information to generate the set of scene images consistent with the initial image geometry and semantic information.
[0073] Specifically, this step is used to enhance the lighting and color of the task scene image without changing the original image geometry, including object category, position, shape, size, and spatial relationship, to generate a set of scene images with slightly different appearances , to improve the robustness of subsequent models to environmental changes. The enhancement process is based on the diffusion-based generation model (Stable Diffusion image-to-image mode img2img).
[0074] The generation direction of the second large model can be controlled by positive and negative prompt words:
[0075] The positive prompt can be defined as adjusting only the lighting, shadows, exposure, and color tone, while maintaining the composition, style, and geometry of the original scene consistent;
[0076] The negative prompt can be used to suppress object deformation, relocation, style transfer, and material change, to prevent unintended geometric or semantic deviation.
[0077] Through the above combination of prompts, natural lighting changes and color disturbances can be achieved while maintaining the object and spatial layout unchanged. For example, each picture and the initial picture maintain consistency in object shape, position, and composition, and only natural changes exist in lighting, shadows, and color distribution.
[0078] In one embodiment, the generated set of scene images is verified by structural similarity and key point drift detection to ensure geometric consistency; if the deviation is beyond the limit, it is automatically rolled back or resampled.
[0079] Specifically, the purpose of this embodiment is to ensure that the set of scene images, although there are differences in appearance attributes such as lighting and color from the initial image, still maintain precise consistency in geometric structure and semantic information, preventing unexpected object deformation or position deviation during the generation process. The verification process is as follows:
[0080] For each image in the generated set of scene images, a structural similarity (SSIM) verification and a key point drift detection are performed. The structural similarity verification is used to assess the overall similarity in brightness, contrast, and structure between the generated image and the initial image. The key point drift detection accurately calculates the spatial displacement (i.e., "drift") of key feature points (such as object edges, corners, etc.) between the generated image and the initial image by identifying these key points in the image. The detection results are compared with a pre-set geometric consistency threshold. If the detected offset, especially the key point drift, exceeds the threshold limit, it is determined that the image does not meet the constraint requirements. At this time, an automatic rollback or resampling mechanism is triggered, i.e., the unqualified image is discarded, and the second large model is instructed to generate an image using a new random seed or slightly adjusting the prompt information until the generated set of scene images passes the verification, ensuring that the geometric structure and the initial image remain accurately consistent.
[0081] In the above embodiment, by introducing the verification mechanism of positive and negative prompt words, structural similarity verification, and key point drift detection, it is ensured that the enhanced set of scene images obtained by the generation model can strictly maintain its original geometric structure and object spatial position unchanged while having rich illumination and color changes. The problem of unintended geometric deformation or position offset easily introduced by generation techniques such as diffusion model when performing appearance perturbation is effectively solved. By automatically detecting and rolling back or resampling unqualified images, the geometric consistency of all scene images and the reliability of three-dimensional constraints are guaranteed, providing a high-quality and robust visual basis for subsequent video generation and accurate robot pose extraction, ensuring the physical accuracy of the entire data package from the source.
[0082] In one embodiment, the step of generating a color video corresponding to the task phase semantic sequence based on the task phase semantic sequence and the set of scene images by a third large model specifically includes:
[0083] The color video is sequentially spliced from phase video segments, and the number of phase video segments is consistent with the number of phases in the task phase semantic sequence.
[0084] The third large model takes the task phase semantic sequence as a generation constraint and takes the images in the set of scene images as the starting frames of the first phase video segment. For subsequent phase video segments, the final frame of the previous phase video segment is taken as the starting frame of itself to generate the color video.
[0085] Specifically, the embodiment aims to generate a complete and coherent long-time color video that precisely follows the task execution flow decomposed by the first large model. The color video is spliced by multiple independent stage video segments in sequence, and the number of stage video segments is strictly consistent with the number of stages decomposed in the task stage semantic sequence.
[0086] The generation process is implemented by the third large model, in which the task stage semantic sequence is used as the core generation constraint to ensure that the video content precisely matches the semantic instructions. For example, video world model (COSMOS) can be used for generation, with instruction_zh and instruction_en in the task stage semantic sequence as semantic constraints, and the initial picture as the appearance starting point.
[0087] The coherence of the video sequence is ensured by the following segment linking logic:
[0088] The starting frame of the first stage: when generating the first stage video segment, the third large model takes one of the scene image sets generated by the second large model as the starting frame. This ensures that the starting scene of the video has high diversity given by step S102.
[0089] The starting frame of the subsequent stage: when generating the subsequent stage video segment, the model takes the final frame of its previous stage video segment as its starting frame.
[0090] In this way, the action result of the previous stage directly serves as the starting state of the next stage, ensuring high smoothness and coherence in content, lighting, object position, and action timing of the entire video sequence, effectively avoiding the abrupt jump phenomenon commonly seen in traditional segmented generated videos.
[0091] Further, in an embodiment, the transition between frames is kept natural through optical flow alignment and color matching, and a flow-guided crossfade method is used to smoothly link at the segment boundary, avoiding picture jumps.
[0092] Further, in an embodiment, the negative prompt consistent with step S102 is followed throughout the entire generation process of the color video, suppressing geometric structure and style drift, and ensuring that objects, scenes, and actions remain consistent in vision and semantics.
[0093] Specifically, to ensure that the complete color video spliced by multiple stage video segments has high temporal smoothness and visual consistency, the embodiment introduces precise post-processing and continuous generation constraints:
[0094] First, in terms of segment splicing, the embodiment adopts multiple techniques to maintain the naturalness of the transition between adjacent frames. Specifically, the optical flow alignment technique is used to analyze the pixel motion trend at the boundary of adjacent stage segments, and combined with the color matching algorithm, the color tone mutation caused by the randomness of the generative model is eliminated. Further, the system uses a flow-guided crossfade method at the segment boundary, which uses flow information to guide the direction and intensity of the crossfade, thereby achieving a highly smooth and natural transition to avoid picture jumps or flickering.
[0095] Second, in terms of consistency control during generation, the embodiment uses the negative prompt used in step S102 throughout the generation process of the color video. The negative prompt continuously acts on the third large model, aiming to suppress the drift of geometric structure and style. This continuous constraint ensures that the objects, backgrounds and overall visual style in the scene can maintain high visual and semantic consistency with the initial image and the enhanced scene set throughout the duration of the video sequence, preventing common object deformation, relocation or background distortion in long-time sequence generation.
[0096] In the above embodiment, through the above mechanism, the embodiment ensures that the final output complete color video not only meets the semantic instructions in terms of content, but also meets the quality standards that can be used for high-precision trajectory inversion in terms of vision and timing.
[0097] In one embodiment, the step of combining the color video and the corresponding depth video to inversely deduce the robot control trajectory data corresponding to the video timing specifically includes:
[0098] According to the color video and the depth video, the pose vectors of the target object and the end effector in each frame of the video are obtained to form a six-dimensional pose sequence;
[0099] The inverse dynamics network is used to take the adjacent frame pose vectors in the six-dimensional pose sequence as input, calculate the pose difference component and perform feature extraction, and predict the control increment parameter describing the action instruction of the robot at this time step;
[0100] The control increment parameter is integrated with the corresponding timestamp and end pose data to form the robot control trajectory data.
[0101] Specifically, first, the color video and the corresponding depth video are subjected to multi-modal processing. Through visual feature recognition positioning, three-dimensional reconstruction combined with depth information, and other methods, the six-dimensional pose vectors of the target object and the end effector (e.g., a robot arm) involved in the task at each frame of the video are accurately obtained. These pose vectors are then arranged in chronological order to form a full-time continuous six-dimensional pose sequence s t =[x,y,z,roll,pitch,yaw]. Among them, x, y, z are spatial positions. These three parameters define the absolute three-dimensional coordinates of the object in the global coordinate system (e.g., the robot base or world coordinate system). The x-axis, y-axis, and z-axis represent the displacement of the object along these three orthogonal directions, accurately determining the position of the object in space. roll, pitch, yaw are spatial poses. These three parameters, usually expressed in the form of Euler angles, define the rotation angles of the object around its own coordinate axes, i.e., its orientation. roll (roll angle) is usually the rotation around the X-axis, pitch (pitch angle) is the rotation around the Y-axis, and yaw (yaw angle) is the rotation around the Z-axis. These three angles combine to completely describe all possible orientation states of the object in space.
[0102] After obtaining the complete pose time sequence, it is input into the inverse dynamics network (IDM-A) for control quantity reasoning. The IDM-A model takes the pose vectors of adjacent frames as input, first calculates the differential quantity of the pose change, and extracts motion features through an encoding network; then predicts the control increment parameter through a multi-layer perceptron, which is used to describe the action instructions of the robot at this time step, such as joint angle change Δq or end pose increment Δx.
[0103] Finally, the predicted control increment parameter, corresponding timestamp information, and the original extracted end pose data are integrated and encapsulated to form the robot control trajectory data. This trajectory data is the input basis for subsequent physical calibration and post-processing.
[0104] Further, in one embodiment, after obtaining the robot control trajectory data, the method further includes consistency optimization of the robot control trajectory data, which specifically includes:
[0105] inputting the control parameters predicted by the inverse dynamics into a differentiable physics simulation engine to generate a simulation video, determining a trajectory optimization error by calculating the difference between the simulation video and the real video, and propagating the optimization error back to the inverse dynamics network to correct the control parameters.
[0106] Specifically, the embodiment introduces a bidirectional consistency optimization mechanism based on the Brax differentiable physics engine and optical flow estimation to calibrate the robot control trajectory data predicted by the inverse dynamics network (IDM-A) in a closed loop.
[0107] In the optimization process, the single action instruction predicted by the IDM-A at each time step is referred to as a control increment parameter; and all the control increment parameters successively predicted by the IDM-A over the entire task timing collectively constitute a control signal sequence for inputting into the simulation engine. This mechanism aims to ensure that the final trajectory data meet strict physical feasibility and timing consistency.
[0108] The embodiment is forward consistency optimization, and the control signal sequence output by the IDM-A is input into the Brax simulation environment of the differentiable physics engine for execution. The simulation engine strictly calculates the motion of each joint of the robot, the end trajectory, and the corresponding physical feedback according to the control signal, and generates a virtual simulation video frame. The simulation generated video is compared with the real video (including color video and depth video) frame by frame, and the consistency degree of the simulation action and the real action is measured by calculating the pixel-level difference and structural similarity evaluation (SSIM). If the deviation exceeds the preset threshold, the error is back propagated to the IDM-A network by using the differentiability of Brax. This mechanism can automatically correct the network parameters to ensure that the control signal output by the IDM-A in the next reasoning can drive the simulation result to more accurately fit the real video action. For example, if the position of the robot arm grabbing the cup in the simulation deviates by 3 cm from the real video, the system will adjust the joint angular velocity component at this stage in the reverse direction until the grabbing position in the simulation picture coincides with the video picture, thereby ensuring the physical feasibility of the trajectory.
[0109] In the above embodiment, by introducing a bidirectional consistency optimization mechanism based on the differentiable physics engine and optical flow estimation, the problem of physical and timing inconsistency between video observation and control signal in the existing trajectory derivation method is fundamentally solved. The forward optimization uses the differentiable physics engine to verify and correct the control signal in a closed loop, ensuring the physical feasibility and dynamic constraints of the final trajectory; the reverse optimization converts the visual observation motion into the control signal through optical flow analysis, ensuring that the trajectory and the video are strictly aligned in amplitude and timing. The bidirectional calibration mechanism eliminates the cumulative error and uncertainty in the inverse dynamics reasoning process, so that the finally output robot control trajectory data has high physical reproducibility and actual executability, thereby greatly improving the training value of the data.
[0110] Further, in one embodiment, the pixel motion between adjacent video frames is analyzed based on optical flow estimation, and combined with depth information to convert into spatial motion, forming a visual backstepping control signal; by comparing the difference between the visual backstepping control signal and the inverse dynamics predictive control parameter, the inverse dynamics predictive control parameter is corrected.
[0111] Specifically, the present embodiment is reverse consistency optimization, which focuses on ensuring the alignment of timing and amplitude. Based on optical flow estimation algorithm, the pixel motion between adjacent video frames is analyzed, the displacement and direction change of the end effector in the pixel space are extracted, and the depth information is used to convert it into three-dimensional space motion, forming a visual backstepping control signal. Then, the visual backstepping control signal is compared with the control increment parameter output by IDM-A in numerical value and direction, the purpose is to make them highly consistent in amplitude, direction and timing. For example, when the optical flow estimation shows that the end of the robot arm moves forward 5mm in 0.1s, while the control increment parameter predicted by IDM-A corresponds to an end displacement of 7mm, the system will automatically adjust the network output or the proportion coefficient, so that they tend to be consistent in subsequent iterations, ensuring that the control trajectory and the motion rhythm observed by the video are completely aligned.
[0112] Further, in one embodiment, depth data is used to assist in outlier rejection and accurate recovery of spatial position.
[0113] Specifically, the depth video is used to assist the state estimation module in high-precision tracking of target objects and end effectors involved in the task. Specifically, the application of depth data is embodied in the following two key aspects:
[0114] Outlier rejection: When performing feature point matching and tracking based on color video, it is inevitable to produce false matching points. Depth information provides a reliable three-dimensional geometric constraint, which can be used to judge the difference between the depth value of the outlier and the surrounding points, so as to accurately identify and reject unreasonable two-dimensional feature points or false three-dimensional re-projection points. This significantly improves the robustness and accuracy of pose estimation.
[0115] Accurate recovery of spatial position: Depth information combined with camera intrinsic parameters is the key to accurate back-projection of two-dimensional image coordinates to three-dimensional space coordinates. By using real depth data that has been calibrated in scale, it can ensure that the recovered pose information has very high three-dimensional spatial accuracy, effectively solving the problem of cumulative error caused by relying only on vision or a single modality estimation.
[0116] Through the above application, the present embodiment ensures that the final six-dimensional pose sequence obtained is stable, continuous and high-precision.
[0117] Further, in one embodiment, the generated robot control trajectory data is subjected to velocity, acceleration and torque limiting processing to ensure that the signal is physically feasible within the execution range of the robot;
[0118] Further, in one embodiment, the trajectory is smoothed by a spline curve or quadratic programming algorithm to make the motion continuous and non-jittering and to comply with the dynamics constraints.
[0119] Specifically, the robot control trajectory data (including position, velocity, acceleration and torque parameters) needs to be subjected to velocity limiting, acceleration limiting and torque limiting processing. This processing ensures that the amplitude of the trajectory signal will not exceed the maximum execution range allowed by the actual robot system hardware and dynamics constraints, thereby ensuring the physical feasibility of the signal within the execution environment of the robot.
[0120] After the limiting processing, the robot control trajectory data also needs to be further smoothed by a spline curve or quadratic programming algorithm. The purpose of this smoothing operation is to eliminate high-frequency noise and jitter that may be generated in the inverse dynamics reasoning and optimization process, so that the final robot motion trajectory exhibits continuous and non-jittering motion in time and ensures that it complies with the dynamics constraints of the robot.
[0121] In this embodiment, a strict quality assurance mechanism is established at the end of the trajectory data generation process. By using depth information for outlier rejection and accurate recovery of spatial position in pose extraction, this method significantly improves the three-dimensional spatial accuracy and stability of the six-dimensional pose sequence, overcoming the cumulative error problem that can be easily introduced by relying solely on color vision, providing a reliable input basis for high-precision inverse dynamics derivation. At the same time, the generated control trajectory is subjected to velocity / acceleration / torque limiting and spline curve / quadratic programming smoothing processing, which fundamentally ensures that the trajectory data complies with the physical constraints of the robot in terms of signal amplitude and is continuous and non-jittering in terms of motion form. This combined mechanism ensures that the final output control trajectory data has extremely high physical reproducibility and actual executability, and can be directly used to drive real robots or high-performance simulators.
[0122] In one embodiment, the step of obtaining the pose vector of the target object and the end effector in each frame of the video to form a six-dimensional pose sequence according to the color video and the depth video specifically includes:
[0123] Determining the two-dimensional image coordinates of the target object and the end effector in the color video using a visual feature recognition model;
[0124] Converting the two-dimensional image coordinates into three-dimensional spatial position information in combination with corresponding depth video and camera model parameters;
[0125] calculating relative kinematic parameters of the object and the end effector by feature correlation between sequential frames;
[0126] performing temporal smoothing on the relative kinematic parameters to obtain the six-dimensional pose sequence.
[0127] Specifically, the color video is analyzed frame by frame using a visual feature recognition model (e.g., a target detection and key point recognition model) to determine the two-dimensional image coordinates (i.e., pixel positions) of the target object and the end effector in each frame. Then, by combining the corresponding depth video and the pre-calibrated camera model parameters (e.g., camera intrinsic parameters), the two-dimensional pixel positions are converted into three-dimensional spatial position information through back-projection operation. This process realizes the conversion from two-dimensional image observation to three-dimensional coordinates with real scale constraints. After obtaining continuous three-dimensional spatial position information, a feature correlation technique between sequential frames (e.g., continuous frame matching or feature point tracking) is used to analyze the motion of the object and the effector. By comparing and calculating the motion state between adjacent frames, the relative kinematic parameters of the object and the end effector are calculated, including all six degrees of freedom parameters such as displacement and rotation.
[0128] Finally, the calculated relative kinematic parameters are smoothed (e.g., fused with inertial constraints or using a filtering algorithm). This smoothing process aims to eliminate measurement noise and transient jitter in the tracking process.
[0129] In this embodiment, through the above steps, it is ensured that the final obtained six-dimensional pose sequence is continuous, stable and high-precision, and it is ensured that the pose data extracted from the video can be used as reliable and accurate kinematic input for subsequent inverse dynamics network reasoning.
[0130] In one embodiment, the step of performing frame-by-frame monocular depth estimation on the color video to generate a depth video aligned with the color video specifically includes:
[0131] performing frame-by-frame monocular depth estimation on the color video to obtain a predicted depth video;
[0132] extracting pixels corresponding to the active object region from the depth image corresponding to the initial image as a depth reference;
[0133] calculating a linear scale mapping relationship based on the statistical characteristics of the predicted depth video and the depth reference in the active object region;
[0134] applying the linear scale mapping relationship to the predicted depth video to generate a depth video aligned with the color video.
[0135] Specifically, the embodiment aims to generate depth information that is strictly aligned with the color video in both time and space, and has real-world scale. Frame-by-frame monocular depth estimation is performed on the color video (e.g., using the DepthAnything model) to obtain a predicted depth video with only relative depth information. To convert the relative depth to real scale, first, the pixel points corresponding to the task-related active object regions (such as the regions where the target object and the end effector are located) are extracted from the real depth image corresponding to the initial image, and the real depth values of these pixel points are taken as the depth reference. Based on the depth value distribution of the active object regions in the predicted depth video and the statistical characteristics (such as mean, variance, median, etc.) of the depth reference, a linear scale mapping relationship (i.e., a scale factor α and an offset β) is calculated. The mapping relationship reflects the fixed proportional deviation between the predicted depth and the real depth. The calculated linear scale mapping relationship is applied to all pixel values of the predicted depth video.
[0136] In the embodiment, through this global linear calibration operation, a depth video with real-world scale that is strictly aligned with the color video in both time and space is finally generated. The depth data with real-world scale ensures the accuracy of subsequent pose estimation and trajectory inversion.
[0137] Further, in an embodiment, the depth video is stored in a single-channel high-bit-depth format to ensure the accuracy and precision of the depth information.
[0138] Specifically, the depth video is packaged and output in MKV format, and is encoded in single-channel 16-bit or float16 precision. This storage format avoids the quantization error caused by the insufficient bit depth of the common 8-bit image format (such as RGB image), so that the real-world depth values after scale calibration can be accurately recorded. By using high-bit-depth format, the accuracy and precision of the depth data are ensured not to be lost in the storage link, providing a reliable data foundation for subsequent high-precision pose estimation and inverse dynamics analysis.
[0139] Through the above embodiments, the present application utilizes a large model to perform structured semantic disassembly of tasks, providing precise action guidance and semantic constraints for the entire data generation process. At the same time, by using constrained image enhancement technology, the appearance and environmental diversity of the data are greatly increased without changing the scene geometry, effectively enhancing the environmental generalization ability of embodied intelligent models trained based on this data. By using the seamless connection strategy of the final frame in the previous stage as the starting frame of the subsequent stage, the visual coherence of long-time action video is ensured. More importantly, by calibrating and aligning the monocular depth estimation results to the real scale, reliable three-dimensional spatial constraints are provided for subsequent precise pose extraction, overcoming the lack of accurate depth information. By introducing a bidirectional consistency optimization mechanism of forward physical calibration using a differentiable physics engine and reverse time sequence correction using optical flow estimation, the problem of physical mismatch and time sequence deviation between video observation and control trajectory is solved, ensuring that the final output trajectory data has high physical reproducibility and time sequence strict alignment. Finally, by automatically aligning the task semantics and video timeline, the method realizes the integrated and automated output of color video, depth video, control trajectory, and structured semantic annotation file. It greatly reduces the cost of manual collection and fine annotation, and provides key technical support for training robust embodied intelligent models.
[0140] In addition, some embodiments of the present application also provide an electronic device. The electronic device can be various forms of digital computers, such as laptops, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and the like. The electronic device can also be various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices.
[0141] The electronic device includes one or more processors and a memory having stored computer program instructions that, when executed, cause the processor to perform the steps of the method provided by any one or more embodiments described above. Figure 2An exemplary structural diagram of the electronic device is disclosed. The electronic device includes one or more processors 1101, memory 1102, and an interface for connecting the components, including a high-speed interface and a low-speed interface. The components are connected to each other through different buses, and can be mounted on a common main board or otherwise installed as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display a GUI on an external input / output device such as a display device coupled to the interface. In some other embodiments, a plurality of processors and / or buses can be used with a plurality of memories and a plurality of memory, if necessary. Also, a plurality of electronic devices can be connected, each device providing part of the necessary operations. Among them, the components shown herein, their connections and relationships, and their functions are merely examples, and are not intended to limit the implementation of the present application described and / or claimed herein.
[0142] The electronic device can further include input means 1103 and output means 1104. The processor 1101, the memory 1102, the input means 1103, and the output means 1104 can be connected by a bus or otherwise, and are connected by a bus in the figure.
[0143] The input means 1103 can receive input digital or character information, and generate key signal input related to user settings and function control of the electronic device, such as touch screen, keypad, mouse, trackpad, touchpad, pointing stick, one or more mouse buttons, trackball, joystick, etc. The output means 1104 can include a display device, an auxiliary lighting device (e.g., LED), and a tactile feedback device (e.g., a vibration motor), etc. The display device can include, but is not limited to, a liquid crystal display, a light-emitting diode display, and a plasma display. In some embodiments, the display device can be a touch screen.
[0144] To provide interaction with the user, the electronic device can be a computer. The computer has a display device (e.g., a cathode ray tube or an LCD monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback); and input from the user can be received in any form (e.g., voice input or tactile input).
[0145] In the embodiments of the present application, the computer program / instruction stored on the computer readable medium is executed by the processor to implement the steps of the method provided by any one or more of the embodiments described above. The computer readable medium can be included in the electronic device described in the embodiments above, or can exist separately and not be assembled into the device. The computer readable medium carries one or more computer readable instructions.
[0146] The memory 1102 can be used as a non-transitory computer readable medium to store non-transitory software programs, non-transitory computer executable programs and modules. The processor 1101 executes various functions and data processing of the server by running the non-transitory software programs, instructions and modules stored in the memory 1102, so as to implement the program instructions / modules corresponding to the method provided by any one or more of the embodiments described above in the embodiments of the present application.
[0147] The memory 1102 can include a program storage area and a data storage area. The program storage area can store an operating system and at least one application required by a function. The data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 1102 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory 1102 can optionally include a memory disposed remotely with respect to the processor 1101, and these remote memories can be connected to the electronic device through a network. Examples of the network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0148] It should be noted that the computer readable medium described in the present application can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared or semiconductor system, device or component, or any combination of the above. More specific examples of computer readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component.
[0149] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technologies, read-only optical discs, digital versatile discs or other optical storage, magnetic cassette tapes, magnetic tape discs storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device.
[0150] Computer program code for carrying out operations of the present application can be written in one or more programming languages or combinations of languages including object-oriented, such as Java, Smalltalk, C++, conventional procedural programming languages, such as the C programming language or similar programming languages. Program code can be executed entirely on a user computer, partially on a user computer, as a standalone software package, partially on a user computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user computer through any kind of network, including a local area network or a wide area network, or can be connected to an external computer (for example, through an Internet service provider to connect through the Internet).
[0151] In the above embodiments, all or part can be implemented by software, hardware, firmware or any combination thereof. For example, a dedicated integrated circuit, a general-purpose computer or any other similar hardware device can be used. In some embodiments, the software program of the present application can be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drive or soft disk and similar devices. In addition, some steps or functions of the present application can be implemented by hardware, for example, as a circuit cooperating with the processor to perform each step or function.
[0152] The computer program product provided by the embodiments of the present application includes one or more computer programs / instructions, which, when executed by a processor, generate all or part of the processes or functions described in the embodiments of the present application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (such as floppy disk, hard disk, magnetic tape), optical media (such as DVD), or semiconductor media (such as solid state disk), etc.
[0153] The flowcharts or block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that shown in the figures. For example, two blocks noted in succession can actually be executed substantially concurrently, or they can sometimes be executed in reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, as well as combinations of blocks in the block diagrams and / or flowcharts, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or they can be implemented by a combination of dedicated hardware and computer instructions.
[0154] The scope of the present application is defined by the appended claims rather than the description preceding it, so all changes that come within the meaning and range of equivalency of the claims are to be embraced within the scope of the present application. No reference signs in the claims should be considered as limiting the scope of the claims to the features identified by the reference signs. Furthermore, the word "comprising" does not exclude other elements or steps, and the singular "a" or "an" does not exclude the plural. Multiple units or apparatuses stated in an apparatus claim can also be implemented by one unit or apparatus through software or hardware. The words "first", "second" and the like do not imply any particular order but are used to identify different components. They also do not indicate or imply any relative importance of the referenced elements.
[0155] The above merely provides specific examples of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims, and the above examples should be regarded as exemplary and non-limiting.
Claims
1. A method for embodied intelligent data synthesis and annotation, characterized in that, The method includes: The system receives an embodied intelligence task described in natural language, along with the corresponding initial and depth images. The first major model decomposes the embodied intelligence task into a semantic sequence of task stages with action type labels. The initial image is perturbed and transformed, and a set of scene images with appearance differences from the initial image is generated by the second large model; Based on the semantic sequence of the task stage and the set of scene images, a color video corresponding to the semantic sequence of the task stage is generated through the third major model; Perform frame-by-frame monocular depth estimation on the color video to generate a depth video aligned with the color video; By combining the color video and the corresponding depth video, the robot control trajectory data corresponding to the video time sequence is derived through inverse dynamics; Based on the duration of each segment in the color video, the semantic sequence of the task stage is mapped to the timeline of the color video to generate a structured semantic annotation file; Output the color video, the depth video, the robot control trajectory data, and the structured semantic annotation file.
2. The method according to claim 1, characterized in that, The process of decomposing the embodied intelligence task into a semantic sequence of task stages labeled with action types using the first major model specifically includes: The first major model, based on the natural language description, decomposes the task into at least two executable stages; The semantic sequence of task stages is composed of the decomposed executable stages arranged in chronological order; The semantic sequence of the task phase is a strongly constrained structured format, which includes action type labels, native language instructions, and target language instructions.
3. The method according to claim 1, characterized in that, The step of performing a perturbation transformation on the initial image and generating a set of scene images with appearance differences from the initial image through the second large model specifically includes: The second large model is prompted to perform perturbation transformation on the lighting and color distribution of the initial image by positive prompts; And / or through negative prompts, make the set of scene images generated by the second large model consistent with the geometric structure and semantic information of the initial image.
4. The method according to claim 1, characterized in that, The step of generating a color video corresponding to the semantic sequence of the task stage using a third model based on the semantic sequence of the task stage and the scene image set specifically includes: The color video is composed of sequentially spliced stage video segments, and the number of stage video segments is consistent with the number of stages in the semantic sequence of the task stages. The third major model uses the semantic sequence of the task stage as the generation constraint and the image in the scene image set as the starting frame of the first stage video segment; subsequent stage video segments use the last frame of the previous stage video segment as their own starting frame to generate the color video.
5. The method according to claim 1, characterized in that, The step of combining the color video and the corresponding depth video to deduce the robot control trajectory data corresponding to the video time sequence through inverse dynamics specifically includes: Based on the color video and depth video, the pose vectors of the target object and the end effector in each frame of the video are obtained to form a six-dimensional pose sequence. Using an inverse dynamics network with the attitude vectors of adjacent frames in the six-dimensional attitude sequence as input, the attitude difference components are calculated and features are extracted to predict the control increment parameters that describe the robot's action command at that time step. The control increment parameters are integrated with the corresponding timestamps and end-effector pose data to form the robot control trajectory data.
6. The method according to claim 5, characterized in that, The step of obtaining the pose vectors of the target object and the end effector in each frame of the video based on the color video and the depth video, and forming a six-dimensional pose sequence, specifically includes: The two-dimensional image coordinates of the target object and the end effector are determined in the color video using a visual feature recognition model. By combining the corresponding depth video and camera model parameters, the two-dimensional image coordinates are converted into three-dimensional spatial position information; The relative kinematic parameters of the object and the end effector are calculated by feature correlation between time frames. The relative kinematic parameters are temporally smoothed to obtain the six-dimensional attitude sequence.
7. The method according to claim 1, characterized in that, The step of performing frame-by-frame monocular depth estimation on the color video to generate a depth video aligned with the color video specifically includes: Perform frame-by-frame monocular depth estimation on the color video to obtain the predicted depth video; From the depth image corresponding to the initial image, extract the pixels corresponding to the active object region as the depth reference; Based on the statistical characteristics of the predicted depth video and the depth benchmark in the region of the moving object, a linear scale mapping relationship is calculated. The linear scale mapping relationship is applied to the predicted depth video to generate a depth video aligned with the color video.
8. An electronic device, characterized in that, The electronic device includes: One or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method as described in any one of claims 1 to 7.
9. A computer-readable medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 7.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Body-equipped agent training system and method
CN118194966A
Mechanical arm motion control method based on multi-agent cooperation
CN120620234A
Automatic action trajectory labeling system and method based on visual language model
CN120747818A
Robot operation track generation method based on structure perception and knowledge enhancement reasoning
CN120765961A
Robotic grasping using efficient vision transformer
WO2025049074A1