Method and device for predicting operation track of motion-free labeling robot based on body flow representation

By using a motionless annotation method based on embodied flow representation, combined with multimodal information and robot kinematic constraints, accurate robot operation trajectories are generated, solving the problems of data dependence and joint constraint differences in existing technologies, and enabling efficient robot operation in complex environments.

CN120839772APending Publication Date: 2025-10-28INST OF AUTOMATION CHINESE ACAD OF SCI +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510862711.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing methods for predicting robot maneuvers rely on large-scale datasets with action annotations, which makes data collection difficult and introduces noise. They are also inadequate for handling flexible objects, occlusion interference, and non-object displacement manipulation. Furthermore, the kinematic constraints of robot joints vary significantly and are difficult to handle uniformly.

Method used

A motion-free annotation method based on embodied flow representation is adopted. By acquiring the robot's embodied flow, language task instructions, and visual features as multimodal conditional features, and combining them with the robot's standard structural description file and camera parameters, positive kinematic mapping is performed to decompose the motion components of each joint and generate accurate robot operation trajectories.

Benefits of technology

It achieves accuracy and robustness in robot operation without motion annotation, can handle flexible objects and occlusion, and improves the robot's generalization ability and operation success rate in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120839772A_ABST
    Figure CN120839772A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for predicting an operation track of an action-free labeling robot based on body flow representation, and the method comprises the steps: taking an initial point position in the body flow of the robot, a text feature in a language task instruction and a visual feature in an initial image as multi-modal condition features; the multi-modal condition features are input into the robot trajectory prediction model to obtain body flow information, and the robot trajectory prediction model can comprehensively consider motion of the robot, object related information contained in a language instruction and visual features in an initial image, so that more accurate body flow information is generated; the robot can accurately interact with a specific object; according to the method, the motion constraint information of each joint of the robot is obtained, the body flow information is decomposed into the motion component of each joint through forward kinematics mapping based on the motion constraint information of each joint, and kinematics constraints of different joints of the robot are fully considered, so that the accurate control of the robot motion is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot perception and operation strategy learning technology, and in particular to a method and apparatus for predicting the operation trajectory of a robot without motion annotation based on embodied flow representation. Background Technology

[0002] Robotic control policy learning has received widespread attention in recent years. Current cutting-edge methods primarily employ imitation learning frameworks, directly mapping visual observations and verbal commands to low-level robot actions through end-to-end vision-language-action models. However, these methods suffer from a drawback: they heavily rely on large-scale datasets with action annotations. Such datasets are not only difficult to collect at scale but also prone to introducing noise, leading to a decrease in the robustness of the robot's policies.

[0003] Besides the limited datasets with action annotations, a large number of unannotated manipulation videos contain rich prior motion information, which can alleviate the problem of data scarcity. Currently, some technologies have proposed to infer sub-targets through image or video generation models to improve sample efficiency and generalization in the policy learning process. However, these methods are essentially still in the category of imitation learning and still require data with underlying action annotations for training. At the same time, there is also a method that completely avoids action annotation, namely, predicting the object-centered optical flow and inferring robot actions from changes in objects between video frames. Although such methods have certain potential, they show obvious limitations in complex real-world application scenarios: (1) Rigid assumption limitation: such methods assume that the motion changes of all parts of the object are uniform, making it difficult to effectively handle flexible and deformable objects; (2) Susceptible to occlusion interference: action inference depends on changes in the object's state, and when part of the object is occluded, it is difficult to obtain accurate motion information; (3) Difficulty in handling non-object displacement manipulation: when the manipulation action is not accompanied by obvious object displacement (e.g., rotating a switch or pressing a button), existing methods are difficult to infer the correct robot action. The above limitations severely restrict the generalization ability of such methods in a wider range of real-world manipulation scenarios.

[0004] However, the implementation of the robot-centric embodied flow approach still faces two major technical challenges: (1) Although embodied flow prediction emphasizes the robot's own motion, most manipulation tasks require precise interaction between the robot and specific objects. If the robot's motion is predicted directly while ignoring the state of the specific object, the robot's actions may be inaccurate. Especially when the language instructions contain explicit object-related information, it is impossible to complete the interaction task required by the specific instructions. (2) Different joints of a robot usually have their own unique kinematic constraints, such as significant differences in range of motion, degrees of freedom, and dynamic capabilities, making it difficult to uniformly process and calculate actions. Summary of the Invention

[0005] This invention provides a method and apparatus for predicting the trajectory of a robot without motion annotation based on embodied flow representation, in order to overcome the deficiencies of existing robot trajectory prediction methods.

[0006] This invention provides a method for predicting the trajectory of an action-unannotated robot based on embodied flow representation, comprising the following steps: The robot acquires the initial point position in its embodied flow, textual features in the language task instructions, and visual features in the initial image. The initial point position, the text features, and the visual features are used as multimodal conditional features, and the multimodal conditional features are input into the robot trajectory prediction model to obtain the robot's embodied flow information output by the robot trajectory prediction model. Based on the robot's standard structural description file, the motion constraint information of each joint of the robot is obtained. Based on the motion constraint information of each joint, as well as the intrinsic and extrinsic parameters of the camera, the body flow information is decomposed into motion components of each joint through forward kinematic mapping. Based on the motion components, the trajectory is sent.

[0007] According to the present invention, a method for predicting the trajectory of a robot without motion annotation based on embodied flow representation is provided. The training steps of the robot trajectory prediction model include: Obtain multimodal conditional features of the samples and an initial robot trajectory prediction model; the initial robot trajectory prediction model includes an embodied flow prediction model and a target image prediction model; The sample multimodal conditional features are input into the embodied flow prediction model, which determines the first predicted noise distribution of the pixel displacement of the embodied point in the image sequence, and the embodied flow prediction result is determined based on the first predicted noise distribution. The embodied flow prediction result and the sample multimodal conditional features are fused to obtain conditional prior features, and the conditional prior features are input into the target image prediction model, which then determines the second prediction noise distribution of the target image. Based on the difference between the first predicted noise distribution and the first true noise distribution, the body flow loss is determined, and based on the difference between the second predicted noise distribution and the second true noise distribution, the image prediction loss is determined. Based on the embodied flow loss and the image prediction loss, a target loss is determined, and the parameters of the initial robot trajectory prediction model are iterated based on the target loss to obtain the robot trajectory prediction model.

[0008] According to the present invention, a method for predicting the operation trajectory of a robot without motion annotation based on embodied flow representation is provided, wherein the first prediction noise distribution is a positive noise distribution.

[0009] According to the present invention, a method for predicting the trajectory of a robot without motion annotation based on embodied flow representation is provided. The method involves obtaining motion constraint information of each joint of the robot based on a standard structural description file of the robot, and decomposing the embodied flow information into motion components of each joint through forward kinematic mapping based on the motion constraint information of each joint and the intrinsic and extrinsic parameters of the camera. The method includes: Based on the robot's standard structural description file, the motion constraint information of each joint of the robot is obtained, and the motion constraint information of each joint, the intrinsic and extrinsic parameters of the camera, and the body flow information are input into the joint motion prediction model to obtain the motion components of each joint output by the joint motion prediction model. The training steps of the joint motion prediction model include: Obtain the sample body flow information and the initial joint motion prediction model; The motion constraint information of each joint, the intrinsic and extrinsic parameters of the camera, and the sample embodied flow information are input into the initial joint motion prediction model. The initial joint motion prediction model determines the three-dimensional trajectory of each point in each joint at the current moment, as well as the two-dimensional projection position of the next moment. Based on the three-dimensional trajectory and the two-dimensional projection position, the reprojection error is determined. Based on the reprojection error, the parameters of the initial joint motion prediction model are iterated to obtain the joint motion prediction model.

[0010] According to the present invention, a method for predicting the trajectory of a robot without motion annotation based on embodied flow representation is provided, wherein determining the reprojection error based on the three-dimensional trajectory and the two-dimensional projection position includes: The reprojection error is determined based on the pose of the target actuator, the transformation relationship between each joint and the target actuator, the three-dimensional trajectory, and the two-dimensional projection position.

[0011] According to the present invention, a method for predicting the trajectory of a robot without motion annotation based on embodied flow representation includes the following steps for determining the initial point position in the robot's embodied flow: Obtain the initial image frame; Based on a multimodal segmentation model, the physical region where the robot's robotic arm is located in the initial image frame is identified; A predetermined number of pixels are randomly selected within the mask area of ​​the embodied region as embodied key points; The position of the initial point is obtained by tracking the position of the embodied keypoint in the initial image frame using a pre-trained point tracking model.

[0012] The present invention also provides a device for predicting the trajectory of a robot without motion annotation based on embodied flow representation, comprising the following units: The acquisition unit is used to acquire the initial point position in the robot's embodied flow, the text features in the language task instructions, and the visual features in the initial image; The input unit is used to take the initial point position, the text features and the visual features as multimodal conditional features, and input the multimodal conditional features into the robot trajectory prediction model to obtain the robot's embodied flow information output by the robot trajectory prediction model; The trajectory sending unit is used to obtain the motion constraint information of each joint of the robot based on the robot's standard structural description file, and based on the motion constraint information of each joint, as well as the intrinsic and extrinsic parameters of the camera, decompose the body flow information into motion components of each joint through forward kinematic mapping, and send the trajectory based on the motion components.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for predicting the motion-annotated robot operation trajectory based on embodied flow representation as described above.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for predicting the motion-annotated robot operation trajectory based on embodied flow representation as described above.

[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the method for predicting the motion-annotated robot operation trajectory based on embodied flow representation as described above.

[0016] The present invention provides a method and apparatus for predicting robot operation trajectories without motion annotation based on embodied flow representation. On the one hand, it uses the initial point position in the robot's embodied flow, text features in the language task instructions, and visual features in the initial image as multimodal conditional features. These multimodal conditional features are input into the robot trajectory prediction model to obtain the robot's embodied flow information output by the robot trajectory prediction model. In this way, the robot trajectory prediction model can comprehensively consider the robot's own motion, object-related information contained in the language instructions, and visual features in the initial image, thereby generating more accurate embodied flow information. This enables the robot to accurately interact with specific objects and complete the interactive tasks required by specific instructions. On the other hand, based on the robot's standard structural description file, it obtains the motion constraint information of each joint of the robot. Based on the motion constraint information of each joint, as well as the intrinsic and extrinsic parameters of the camera, it decomposes the embodied flow information into motion components of each joint through forward kinematic mapping. Based on the motion components, trajectory is distributed. This fully considers the kinematic constraints of different joints of the robot and can generate motion components that conform to their motion constraints for each joint, thereby achieving precise control of the robot's movements and solving the problem that the significant differences in kinematic constraints of different joints are difficult to handle uniformly. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is one of the flowcharts illustrating the method for predicting the trajectory of a robot without motion annotation based on embodied flow representation provided by the present invention.

[0019] Figure 2 This is the second flowchart of the method for predicting the operation trajectory of a robot without motion annotation based on embodied flow representation provided by the present invention.

[0020] Figure 3 This is a schematic diagram of the structure of the robot operation trajectory prediction device based on embodied flow representation without motion annotation provided by the present invention.

[0021] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0023] The terms "first," "second," etc., used in this invention are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and that the objects distinguished by "first," "second," etc., are generally of the same class.

[0024] Figure 1 This is one of the flowcharts illustrating the method for predicting the trajectory of a robot without motion annotation based on embodied flow representation provided by the present invention, such as... Figure 1 As shown, the method includes steps 110, 120 and 130.

[0025] Step 110: Obtain the initial point position in the robot's embodied flow, the text features in the language task instructions, and the visual features in the initial image.

[0026] Specifically, embodiments of the present invention fundamentally overcome the limitations of existing technologies by transforming the prediction task from object-centered optical flow to robot-centered embodied flow. The embodied flow prediction method no longer relies on object properties or states, enabling generalization to both rigid and flexible objects, and more effectively handling occlusion and non-object displacement manipulation. Furthermore, the robot's own kinematic structure is usually visible in actual manipulation; even under partial occlusion, sufficient motion information can still be extracted from visible joint movements.

[0027] Specifically, first, the initial point positions in the robot's embodied flow, textual features from the language task instructions, and visual features from the initial images are acquired. Embodied flow refers to the motion trajectory of key points on the robot's body (such as a robotic arm) within an image sequence. These key points represent reference positions of the robot's structure, and their movement reflects the robot's operational behavior. In this embodiment, these key points are located on the robotic arm, and the robot's actions are represented by tracking their pixel displacements within the image sequence.

[0028] Here, the initial point location can be encoded using an MLP (Multilayer Perceptron), the text features in the language task instructions can be extracted using a CLIP (Contrastive Language–Image Pre-training) model, and the visual features in the initial image can be extracted using a ResNet model.

[0029] The initial point location refers to the pixel coordinates of the embodied flow keypoints in the first frame of the image sequence. These coordinates serve as the starting position for subsequent embodied flow prediction. In practice, various methods can be used to determine the initial point location, such as extracting the robotic arm region based on a Segment Anything Model (SAM) and randomly sampling within that region.

[0030] Language task instructions refer to natural language instructions that describe the tasks the robot needs to perform, such as "move the red object to the table." Language task instructions contain the task's objectives and constraints, and are crucial information for the robot to understand the task.

[0031] Text features refer to the vector representations obtained after encoding language task instructions. Text features can be extracted using various natural language processing techniques, such as encoding instructions using a pre-trained language model. Here, the pre-trained language model can be a BERT (Bidirectional Encoder Representations from Transformers) model, a Word2Vec model, or a cascaded multi-layer convolutional neural network (CNN), etc., and this embodiment of the invention does not specifically limit its use.

[0032] An image sequence refers to a series of consecutive image frames captured by a camera. Image sequences contain visual information about the robot's operation process and are an important source of information for the robot to perceive its environment and its own state.

[0033] Visual features refer to feature vectors extracted from the first frame (initial image) of an image sequence, used to describe the content of the image. Visual features can be extracted using various computer vision techniques, such as using convolutional neural networks (CNNs) to extract global or local features of an image, or using deep neural networks (DNNs), or a combination of CNNs and DNNs, etc. This embodiment of the invention does not specifically limit these methods.

[0034] Step 120: The initial point position, the text features, and the visual features are used as multimodal conditional features, and the multimodal conditional features are input into the robot trajectory prediction model to obtain the robot's embodied flow information output by the robot trajectory prediction model.

[0035] Specifically, after obtaining the initial point location, text features, and visual features, the initial point location, text features, and visual features can be used as multimodal conditional features.

[0036] A robot trajectory prediction model is used to predict the embodied flow of a robot based on multimodal conditional features. The model employs a diffusion probability model to model the pixel displacements of embodied points in an image sequence. The diffusion process is guided by multimodal conditional features, which include: the initial image... Visual features, language task instructions Textual features, and the location of the embodied initial point. The structural encoding, when combined, forms a condition vector. It is used to provide contextual guidance to the diffusion model at each time step to achieve semantic constraints and task alignment on the embodied flow generation process.

[0037] Embodied flow information refers to the predicted motion trajectory of keypoints on a robot's body within an image sequence. Embodied flow information can be represented as the pixel coordinates of keypoints in each frame of the image, or the pixel displacements of keypoints between adjacent frames. Keypoints can be robot joints, end effectors, or other components. By analyzing the positional changes of these keypoints in consecutive image frames, their motion trajectories can be obtained.

[0038] Step 130: Based on the robot's standard structural description file, obtain the motion constraint information of each joint of the robot, and based on the motion constraint information of each joint, as well as the intrinsic and extrinsic parameters of the camera, decompose the body flow information into motion components of each joint through forward kinematic mapping, and send the trajectory based on the motion components.

[0039] Specifically, after obtaining the embodied flow information, the motion constraint information of each joint of the robot can be obtained based on the robot's standard structural description file. Based on the motion constraint information of each joint, as well as the intrinsic and extrinsic parameters of the camera, the embodied flow information is decomposed into motion components of each joint through forward kinematic mapping. Based on the motion components, the trajectory is sent out.

[0040] Here, the Unified Robot Description Format (URDF) is an XML file used to describe the robot's structure, joints, links, and other information. The URDF file contains the robot's geometric, kinematic, and dynamic parameters, forming the basis for robot control and simulation.

[0041] Motion constraint information refers to the range of motion and speed limits of each joint of the robot. This constraint information is provided by URDF files and is used to ensure the safety and feasibility of robot motion.

[0042] The camera's intrinsic parameters describe its internal parameters, such as focal length and principal point coordinates; the camera's extrinsic parameters describe the camera's position and orientation in the world coordinate system. These parameters are used to convert image coordinates to world coordinates.

[0043] Forward kinematic mapping refers to the process of calculating the pose of the robot's end effector based on the robot's joint angles. Motion components refer to the motion components of each joint of the robot, such as the changes in joint angles. Trajectory transmission refers to sending the calculated motion components of each joint to the target end effector, which then performs the corresponding action.

[0044] For example, the geometric properties and relative positions of each joint are obtained using a Robot Structure Description File (URDF). Combined with intrinsic and extrinsic parameters from the camera, the embodied flow sampling points in the image are mapped to 3D space and assigned to the corresponding joints based on their positional distribution. Sampling points must meet the following conditions: visible in consecutive frames, significant displacement, and having a valid depth value. Each sampling point must belong to only one joint region within the image plane to ensure the uniqueness and accuracy of the assignment.

[0045] Based on the predicted motion components of each joint, the motion is sent to the physical or simulated robot for execution through the robot control system (such as the ROS control interface) to complete the target manipulation task under the language conditions.

[0046] Finally, the obtained trajectory results are encapsulated into standard control commands and sent to the robot arm control interface via the Franka ROS control package. The ROS system is responsible for executing the trajectory tracking task, which includes time synchronization, speed adjustment, and dynamic constraints. The control system has robust response capabilities to minor disturbances, ensuring that the robot end effector accurately executes the task trajectory in three-dimensional space, and improving the overall system's adaptability and stability in complex environments.

[0047] Furthermore, the methods of this invention can be deployed in the standard simulation platform Meta-World environment and real robot hardware systems to comprehensively test their execution capabilities under different perception conditions and task requirements. The test scenarios cover a variety of tasks such as conventional object handling and placement, fine manipulation, and flexible object manipulation.

[0048] Typical tasks, including flexible object manipulation, occlusion interference handling, and non-object displacement interactions (such as rotating switches and pressing buttons), were selected to compare and evaluate the performance differences between the proposed method and existing object-centered optical flow strategies, and to verify the learning ability and policy generalization of the proposed embodied flow method in complex scenarios.

[0049] Quantitative evaluation was conducted using key indicators such as success rate. Experimental results show that the proposed method achieves significant performance improvements in the aforementioned tasks, possesses stronger visual understanding capabilities and action generation stability, and fully verifies the feasibility and superiority of the proposed method under conditions without action annotation.

[0050] Understandably, this method, by introducing embodied flow representation, enables robot trajectory prediction without action annotation, reducing data annotation costs. Simultaneously, by utilizing multimodal information fusion and robot kinematic constraints, it improves the accuracy and feasibility of trajectory prediction.

[0051] The method provided in this invention, on the one hand, uses the initial point position in the robot's embodied flow, text features in the language task instructions, and visual features in the initial image as multimodal conditional features. These multimodal conditional features are then input into the robot trajectory prediction model to obtain the robot's embodied flow information output by the robot trajectory prediction model. In this way, the robot trajectory prediction model can comprehensively consider the robot's own motion, object-related information contained in the language instructions, and visual features in the initial image, thereby generating more accurate embodied flow information. This enables the robot to accurately interact with specific objects and complete the interaction tasks required by specific instructions. On the other hand, based on the robot's standard structural description file, the motion constraint information of each joint of the robot is obtained. Based on the motion constraint information of each joint, as well as the intrinsic and extrinsic parameters of the camera, the embodied flow information is decomposed into motion components of each joint through forward kinematic mapping. Based on the motion components, trajectory is distributed. This fully considers the kinematic constraints of different joints of the robot and can generate motion components that conform to their motion constraints for each joint, thereby achieving precise control of the robot's movements and solving the problem that the significant differences in kinematic constraints of different joints are difficult to handle uniformly.

[0052] Based on the above embodiments, the training steps of the robot trajectory prediction model include: Step 210: Obtain the multimodal conditional features of the samples and the initial robot trajectory prediction model; the initial robot trajectory prediction model includes a body flow prediction model and a target image prediction model. Step 220: Input the sample multimodal conditional features into the embodied flow prediction model, and use the embodied flow prediction model to determine the first prediction noise distribution of the pixel displacement of the embodied point in the image sequence, and determine the embodied flow prediction result based on the first prediction noise distribution; Step 230: The embodied flow prediction result and the sample multimodal conditional features are fused to obtain conditional prior features, and the conditional prior features are input into the target image prediction model, and the target image prediction model determines the second prediction noise distribution of the target image. Step 240: Determine the embodied flow loss based on the difference between the first predicted noise distribution and the first true noise distribution, and determine the image prediction loss based on the difference between the second predicted noise distribution and the second true noise distribution; Step 250: Based on the embodied flow loss and the image prediction loss, determine the target loss, and perform parameter iteration on the initial robot trajectory prediction model based on the target loss to obtain the robot trajectory prediction model.

[0053] Specifically, to better train the robot trajectory prediction model so that it can more accurately predict the robot's embodied flow and thus improve the success rate of robot operations, the robot trajectory prediction model can be trained based on the following steps: First, sample multimodal conditional features and an initial robot trajectory prediction model are obtained. The initial robot trajectory prediction model includes a body flow prediction model and a target image prediction model. The sample multimodal conditional features are the input data used to train the robot trajectory prediction model. These features include the initial point position of the samples, sample text features, and sample visual features. This sample data can be obtained by collecting and preprocessing a video training dataset of robot operations.

[0054] Among them, a video training dataset for robot manipulation strategy learning is constructed. The dataset covers a variety of typical manipulation tasks. Each sample includes a continuous RGB image frame sequence captured by a fixed-view camera, as well as the corresponding natural language task instruction. The sample does not contain action annotations, control signals or joint parameter information during robot execution. Only visual and language information is retained as supervision signals to support the subsequent end-to-end learning framework based on images and instructions.

[0055] The initial robot trajectory prediction model refers to the robot trajectory prediction model that has not been trained. The initial robot trajectory prediction model can be initialized with random parameters, or a pre-trained model can be used as the initial robot trajectory prediction model. This embodiment of the invention does not specifically limit this.

[0056] Here, the embodied flow prediction model is used to predict the motion trajectory of the robot's body keypoints in the image sequence. The target image prediction model is used to predict the target image after the robot's operation is completed. The target image prediction model can assist the embodied flow prediction model to improve the accuracy of trajectory prediction.

[0057] After obtaining the multimodal conditional features, the sample multimodal conditional features can be input into the embodied flow prediction model. The embodied flow prediction model determines the first predicted noise distribution of the pixel displacement of the embodied point in the image sequence, and the embodied flow prediction result is determined based on the first predicted noise distribution.

[0058] Among them, the embodied flow prediction model uses a diffusion probability model to model the pixel displacement of the embodied point in the image sequence, constructing a positive noise process: in, Represents the actual motion trajectory of the embodied point at the initial moment (including pixel coordinates and visibility). express The actual movement trajectory of the embodied point at any given moment. For time steps The noise intensity below, This represents the standard matrix. During the reverse generation process, the embodied flow prediction model learns from the noisy state. To reconstruct the original trajectory, the reverse distribution is defined as: in, Indicates control signal, This represents the mean of the first predicted noise distribution. This represents the variance of the first predicted noise distribution.

[0059] Then, the embodied flow prediction results and sample multimodal conditional features are fused to obtain conditional prior features. These conditional prior features are then input into the target image prediction model, which determines the second prediction noise distribution of the target image. This involves constructing an auxiliary image prediction task corresponding to the target final state frame of the embodied trajectory, reconstructing it using a diffusion probability model isomorphic to the embodied flow. The conditional prior features of the target image at each time step include... ,in It provides the current time step embodied flow prediction results, realizing the two-way fusion of action information and visual results.

[0060] After obtaining the first predicted noise distribution and the second predicted noise distribution, the embodied flow loss can be determined based on the difference between the first predicted noise distribution and the first true noise distribution. The formula for the embodied flow loss is as follows: in, Indicates the first predicted noise distribution. This represents the first true noise distribution. This represents the embodied flow loss. The first true noise distribution is the noise distribution obtained by adding noise to the embodied flow data forward, meaning the first predicted noise distribution is a forward noise distribution.

[0061] It is understandable that the greater the difference between the first predicted noise distribution and the first true noise distribution, the greater the embodied flow loss; the smaller the difference between the first predicted noise distribution and the first true noise distribution, the smaller the embodied flow loss.

[0062] Furthermore, the image prediction loss can be determined based on the difference between the second predicted noise distribution and the second true noise distribution, which is the noise distribution obtained by the forward noise addition process of the target image data.

[0063] It is understandable that the greater the difference between the second predicted noise distribution and the second true noise distribution, the greater the image prediction loss; the smaller the difference between the second predicted noise distribution and the second true noise distribution, the smaller the image prediction loss.

[0064] Finally, the target loss can be determined based on the embodied flow loss and the image prediction loss, and the initial robot trajectory prediction model can be iterated based on the target loss. The initial robot trajectory prediction model after parameter iteration is used as the robot trajectory prediction model.

[0065] The formula for determining the target loss, based on the embodied flow loss and the image prediction loss, is as follows: in, Indicates target loss. Indicates the loss of embodied flow, This represents the image prediction loss. This represents the weighting coefficient.

[0066] It should be noted that the target loss is a joint loss, which is jointly optimized by the embodied flow loss and the image prediction loss to ensure the consistency between the embodied trajectory and the final image of the task in terms of geometric path and semantic target.

[0067] It should be noted that the target image prediction model and the target image alignment module guide the embodied flow to not only restore the robot's actual motion trajectory, but also to meet the task objectives specified in the language description, thereby enhancing the semantic consistency of task-related interactions.

[0068] It is understood that, by introducing a body-centric optical flow modeling framework in this embodiment of the invention, the method can learn effective robot operation strategies from video data without underlying action annotations, significantly improving generalization ability and task execution efficiency in complex environments. The core idea of ​​this method is to replace the traditional object-centric action estimation path with an optical flow modeling approach centered on the robot's own joint motion, thereby effectively avoiding typical problems such as difficulty in modeling flexible objects, lack of information on occluded areas, and lack of obvious object displacement. Simultaneously, by combining a target image alignment mechanism with a physical modeling action generation framework, end-to-end conversion from visual language input to physically executable actions is achieved.

[0069] The method provided in this invention trains a robot trajectory prediction model, enabling it to more accurately predict the robot's embodied flow, thereby improving the success rate of robot operations. By combining embodied flow loss and image prediction loss, the model's generalization ability and robustness can be improved.

[0070] Based on the above embodiments, step 130, which involves obtaining motion constraint information for each joint of the robot based on the robot's standard structural description file, and decomposing the body flow information into motion components for each joint through forward kinematic mapping based on the motion constraint information of each joint and the intrinsic and extrinsic parameters of the camera, includes: Step 131: Based on the robot standard structural description file, obtain the motion constraint information of each joint of the robot, and input the motion constraint information of each joint, the intrinsic and extrinsic parameters of the camera, and the body flow information into the joint motion prediction model to obtain the motion components of each joint output by the joint motion prediction model. The training steps of the joint motion prediction model include: Step 310: Obtain the sample body flow information and the initial joint motion prediction model; Step 320: Input the motion constraint information of each joint, the intrinsic and extrinsic parameters of the camera, and the sample embodied flow information into the initial joint motion prediction model. The initial joint motion prediction model determines the three-dimensional trajectory of each point in each joint at the current moment and the two-dimensional projection position of the next moment at the current moment. Based on the three-dimensional trajectory and the two-dimensional projection position, the reprojection error is determined. Step 330: Based on the reprojection error, perform parameter iteration on the initial joint motion prediction model to obtain the joint motion prediction model.

[0071] Specifically, motion constraint information of each joint of the robot can be obtained based on the robot's standard structural description file. The motion constraint information of each joint, the intrinsic and extrinsic parameters of the camera, and the embodied flow information are then input into the joint motion prediction model to obtain the motion components of each joint output by the joint motion prediction model.

[0072] The training steps for the joint motion prediction model are as follows: First, the sample embodied flow information and the initial joint motion prediction model are obtained. Then, the motion constraint information of each joint, the intrinsic and extrinsic parameters of the camera, and the sample embodied flow information are input into the initial joint motion prediction model. The initial joint motion prediction model determines the three-dimensional trajectory of each point in each joint at the current moment, as well as the two-dimensional projection position of the next moment. Based on the three-dimensional trajectory and the two-dimensional projection position, the reprojection error is determined.

[0073] The parameters of the initial joint motion prediction model can be preset or randomly generated; this embodiment of the invention does not impose specific limitations on this.

[0074] Sample embodied flow information is the input data used to train a joint motion prediction model. Sample embodied flow information can be obtained by collecting and processing a dataset of robot operation videos. For example, a pre-trained point tracking model can be used to predict the position of an initial embodied point in consecutive image frames, obtaining its time-varying pixel coordinates and visibility information to form a complete time-series trajectory, which can be used as a supervisory label for embodied flow prediction.

[0075] A 3D trajectory refers to the movement trajectory of a robot in three-dimensional space. A 3D trajectory can be obtained by converting image coordinates to world coordinates.

[0076] The two-dimensional projection position refers to the position of the three-dimensional trajectory projected onto the image. The two-dimensional projection position can be calculated using camera parameters and a projection model.

[0077] Furthermore, the reprojection error can be determined based on the pose of the target actuator, the transformation relationship between each joint and the target actuator, the 3D trajectory, and the 2D projection position. For example, after obtaining the correspondence between the embodied points and the joints, a depth-sensing reconstruction method is first used to restore the 2D trajectory of the sampling points to 3D coordinates; then, based on the pose of the target actuator... The transformation relationships between each joint and the target end actuator are calculated using the robot's kinematic model. And construct the reprojection error function: in, This indicates the reprojection error. Indicates camera projection operation. and The first The first joint The three-dimensional trajectory of each point and its two-dimensional projection position at the next moment. Indicates the pose of the target actuator. This indicates the transformation relationship between each joint and the target end actuator.

[0078] Understandably, by minimizing the error To obtain the optimal end effector pose that meets the kinematic constraints. It is used to generate action instructions that can be directly transmitted to the robot control system, ensuring that the predicted actions are feasible and consistent in accuracy at the physical level.

[0079] Furthermore, the optimal end effector pose is solved by forward kinematics modeling and inverse kinematics solution.

[0080] 1) Forward kinematics modeling Used to represent the joint angles of a robotic arm Mapped to the pose of the end effector ,Right now: in, This represents the kinematic chain transformation established based on the Denavit-Hartenberg (DH) parameters of the robotic arm, with the output being a homogeneous transformation matrix consisting of displacement and translation matrices. , This represents the Lie group space.

[0081] 2) Solving inverse kinematics When the key points of the target trajectory are located and orientation Given a given condition, the inverse kinematics module solves for the joint angles that satisfy the following conditions. : To improve solution efficiency and accuracy, this invention employs numerical optimization methods (such as gradient descent and Newton's method) combined with redundancy analysis to address the problems of multiple solutions and singular poses. When multiple feasible solutions exist, the solution that satisfies both motion continuity and energy minimization is preferentially selected. , Represent the Lie algebra space.

[0082] Based on the above embodiments, the step of determining the initial point position in the robot's embodied flow includes: Step 410: Obtain the initial image frame; Step 420: Based on the multimodal segmentation model, identify the physical region where the robot's robotic arm is located in the initial image frame; Step 430: Randomly select a preset number of pixels within the mask range of the embodied region as embodied key points; Step 440: Using a pre-trained point tracking model, the position of the embodied keypoint is tracked in the initial image frame to obtain the initial point position.

[0083] Specifically, firstly, video data containing multiple manipulation tasks is collected. Each data point consists of a continuous RGB image sequence recorded by a fixed-position camera and corresponding language task instructions, without containing any underlying motion or joint annotation information. This step provides the input basis for subsequent embodied flow prediction.

[0084] Then, the initial image frame is obtained from the video data. The initial image frame refers to the first frame of the video data, which is used to determine the starting position of the embodied stream.

[0085] Furthermore, based on a multimodal segmentation model, the embodied region of the robot's arm in the initial image frame is identified. The multimodal segmentation model can be a SAM model. The embodied region refers to the area of ​​the robot's body (such as the robotic arm) in the image. That is, the embodied region of the robotic arm in the initial image frame is identified using the multimodal segmentation model, and a bounding box of this region is generated by fusing language prompts and image content. This bounding box is then further refined to obtain a pixel-level segmentation mask, providing spatial constraints for subsequent keypoint sampling.

[0086] Several pixels are randomly selected within the masked area of ​​the embodied region as embodied key points. The key points are evenly distributed in the image space to represent the reference point positions on the robot's body structure, which are used to track the pixel displacements that evolve over time.

[0087] By using a pre-trained point tracking model, the initial embodied point is tracked in consecutive image frames to obtain its pixel coordinates and visibility information that change over time, forming a complete time series trajectory, which is used as a supervision label for embodied flow prediction.

[0088] Finally, the trajectory coordinates of all key points are normalized to eliminate scale differences between different image resolutions, ensuring that the input data in the model learning process has a uniform spatial scale and contrast stability.

[0089] Based on any of the above embodiments Figure 2 This is the second flowchart illustrating the method for predicting the trajectory of a robot without motion annotation based on embodied flow representation provided by this invention. Figure 2 As shown, Figure 2The main structure of the method is illustrated. The left side shows the process of embodied flow prediction using a diffusion probability model, while the right side shows the process of using target image prediction to assist embodied flow prediction, enabling the embodied flow to conform to language instructions and object-related interactions. The middle section represents the module that uses visual features from the initial image, language task instructions, and initial point positions as control conditions for the diffusion model. Specifically, embodied masking restricts the sampling area, random sampling determines the initial position points of key points, and a multimodal conditional feature composed of the initial position points, textual features of the language task instructions, and visual features of the initial image is input into the embodied flow prediction model. The embodied flow prediction model generates the trajectory through multiple diffusion steps; simultaneously, the target image prediction model predicts the visual target corresponding to the action. Both are optimized together to improve the spatial accuracy and semantic consistency of trajectory prediction, ultimately driving the robot to execute the corresponding trajectory through joint motion decomposition.

[0090] The initial point location is encoded using an MLP model, the CLIP model is used to extract textual features from the language task instructions, and the ResNet model is used to extract visual features from the initial image.

[0091] This invention discloses a method for predicting robot operation trajectories without motion annotation based on embodied flow representation, aiming to solve the problems of low learning efficiency and weak generalization ability of existing methods for complex operation tasks under the lack of underlying motion supervision. The method mainly includes the following steps: acquiring video data containing only visual image sequences and language task instructions, without relying on underlying robot motion annotation; extracting the robot embodied region from the image, randomly sampling embodied key points, and constructing pixel-level motion trajectories as supervision signals; using a diffusion probability model to predict the future motion trajectory of embodied points, combining language task instructions and target image guidance to achieve semantic consistency; introducing a target image prediction task, co-training it with the embodied flow modeling process to ensure that the generated actions meet object interaction and task semantic constraints; using a URDF model for kinematic modeling, decomposing the predicted embodied flow into specific joint actions; solving for the optimal end pose through inverse kinematics, and encapsulating the results into control commands, driving the robot to execute the operation task through the ROS interface. Experiments on simulation platforms and real robot systems verified the effectiveness of the method in flexible objects, occluded environments, and non-object displacement tasks, significantly improving learning efficiency and execution success rate under unannotated conditions. This method features low data cost, high policy generalization, and good platform adaptability, making it suitable for various practical application scenarios such as service robots and industrial automation.

[0092] The following describes the robot trajectory prediction device based on embodied flow representation without motion annotation provided by the present invention. The robot trajectory prediction device based on embodied flow representation without motion annotation described below can be referred to in correspondence with the robot trajectory prediction method based on embodied flow representation without motion annotation described above.

[0093] Based on any of the above embodiments Figure 3 This is a schematic diagram of the structure of the robot operation trajectory prediction device based on embodied flow representation without motion annotation provided by the present invention, as shown below. Figure 3 As shown, the device includes: The acquisition unit 310 is used to acquire the initial point position in the robot's embodied flow, the text features in the language task instructions, and the visual features in the initial image; The input unit 320 is used to take the initial point position, the text features and the visual features as multimodal conditional features, and input the multimodal conditional features into the robot trajectory prediction model to obtain the robot's embodied flow information output by the robot trajectory prediction model; The trajectory sending unit 330 is used to obtain the motion constraint information of each joint of the robot based on the robot standard structural description file, and based on the motion constraint information of each joint, as well as the intrinsic and extrinsic parameters of the camera, decompose the body flow information into motion components of each joint through forward kinematic mapping, and send the trajectory based on the motion components.

[0094] The apparatus provided in this invention, on the one hand, uses the initial point position in the robot's embodied flow, text features in the language task instructions, and visual features in the initial image as multimodal conditional features, and inputs these multimodal conditional features into the robot trajectory prediction model to obtain the robot's embodied flow information output by the robot trajectory prediction model. In this way, the robot trajectory prediction model can comprehensively consider the robot's own motion, object-related information contained in the language instructions, and visual features in the initial image, thereby generating more accurate embodied flow information, enabling the robot to accurately interact with specific objects and complete the interaction tasks required by specific instructions. On the other hand, based on the robot's standard structural description file, it obtains the motion constraint information of each joint of the robot, and based on the motion constraint information of each joint, as well as the intrinsic and extrinsic parameters of the camera, it decomposes the embodied flow information into motion components of each joint through forward kinematic mapping. Based on the motion components, it performs trajectory distribution, fully considering the kinematic constraints of different joints of the robot, and can generate motion components that conform to their motion constraints for each joint, thereby achieving precise control of the robot's movements and solving the problem that the significant differences in kinematic constraints of different joints are difficult to handle uniformly.

[0095] Based on any of the above embodiments, a first training unit is further included, wherein the first training unit is specifically used for: Obtain multimodal conditional features of the samples and an initial robot trajectory prediction model; the initial robot trajectory prediction model includes an embodied flow prediction model and a target image prediction model; The sample multimodal conditional features are input into the embodied flow prediction model, which determines the first predicted noise distribution of the pixel displacement of the embodied point in the image sequence, and the embodied flow prediction result is determined based on the first predicted noise distribution. The embodied flow prediction result and the sample multimodal conditional features are fused to obtain conditional prior features, and the conditional prior features are input into the target image prediction model, which then determines the second prediction noise distribution of the target image. Based on the difference between the first predicted noise distribution and the first true noise distribution, the body flow loss is determined, and based on the difference between the second predicted noise distribution and the second true noise distribution, the image prediction loss is determined. Based on the embodied flow loss and the image prediction loss, a target loss is determined, and the parameters of the initial robot trajectory prediction model are iterated based on the target loss to obtain the robot trajectory prediction model.

[0096] Based on any of the above embodiments, the first predicted noise distribution is a positive noise distribution.

[0097] Based on any of the above embodiments, the trajectory sending unit 330 is specifically used for: Based on the robot's standard structural description file, the motion constraint information of each joint of the robot is obtained, and the motion constraint information of each joint, the intrinsic and extrinsic parameters of the camera, and the body flow information are input into the joint motion prediction model to obtain the motion components of each joint output by the joint motion prediction model. It also includes a second training unit, which specifically includes: The sample acquisition unit is used to acquire sample-specific flow information and the initial joint motion prediction model; An error determination unit is used to input the motion constraint information of each joint, the intrinsic and extrinsic parameters of the camera, and the sample body flow information into the initial joint motion prediction model. The initial joint motion prediction model determines the three-dimensional trajectory of each point in each joint at the current moment, as well as the two-dimensional projection position of the next moment. Based on the three-dimensional trajectory and the two-dimensional projection position, the reprojection error is determined. The parameter iteration unit is used to perform parameter iteration on the initial joint motion prediction model based on the reprojection error to obtain the joint motion prediction model.

[0098] Based on any of the above embodiments, the error determination unit is specifically used for: The reprojection error is determined based on the pose of the target actuator, the transformation relationship between each joint and the target actuator, the three-dimensional trajectory, and the two-dimensional projection position.

[0099] Based on any of the above embodiments, a unit for determining the initial point position is further included, wherein the unit for determining the initial point position is specifically used for: Obtain the initial image frame; Based on a multimodal segmentation model, the physical region where the robot's robotic arm is located in the initial image frame is identified; A predetermined number of pixels are randomly selected within the mask area of ​​the embodied region as embodied key points; The position of the initial point is obtained by tracking the position of the embodied keypoint in the initial image frame using a pre-trained point tracking model.

[0100] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communications bus 440. The processor 410 can call logic instructions in the memory 430 to execute a method for predicting the trajectory of a robot without motion annotation based on embodied flow representation. This method includes: acquiring the initial point position in the robot's embodied flow, text features in the language task instructions, and visual features in the initial image; using the initial point position, text features, and visual features as multimodal conditional features, and inputting the multimodal conditional features into a robot trajectory prediction model to obtain the robot's embodied flow information output by the robot trajectory prediction model; acquiring the motion constraint information of each joint of the robot based on the robot's standard structural description file, and decomposing the embodied flow information into motion components of each joint through forward kinematic mapping based on the motion constraint information of each joint and the intrinsic and extrinsic parameters of the camera; and issuing a trajectory based on the motion components.

[0101] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0102] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the above-described method for predicting the trajectory of a robot without motion annotation based on embodied flow representation. The method includes: acquiring the initial point position in the robot's embodied flow, text features in the language task instructions, and visual features in the initial image; using the initial point position, the text features, and the visual features as multimodal conditional features, and inputting the multimodal conditional features into a robot trajectory prediction model to obtain the robot's embodied flow information output by the robot trajectory prediction model; acquiring the motion constraint information of each joint of the robot based on the robot's standard structural description file, and decomposing the embodied flow information into motion components of each joint through forward kinematic mapping based on the motion constraint information of each joint and the intrinsic and extrinsic parameters of the camera; and issuing a trajectory based on the motion components.

[0103] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the above-described method for predicting the trajectory of a robot without motion annotation based on embodied flow representation. The method includes: acquiring the initial point position in the robot's embodied flow, text features in the language task instructions, and visual features in the initial image; using the initial point position, the text features, and the visual features as multimodal conditional features, and inputting the multimodal conditional features into a robot trajectory prediction model to obtain the robot's embodied flow information output by the robot trajectory prediction model; acquiring the motion constraint information of each joint of the robot based on the robot's standard structural description file, and decomposing the embodied flow information into motion components of each joint through forward kinematic mapping based on the motion constraint information of each joint and the intrinsic and extrinsic parameters of the camera; and issuing a trajectory based on the motion components.

[0104] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0105] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting the trajectory of a robot without motion annotation based on embodied flow representation, characterized in that, include: The robot acquires the initial point position in its embodied flow, textual features in the language task instructions, and visual features in the initial image. The initial point position, the text features, and the visual features are used as multimodal conditional features, and the multimodal conditional features are input into the robot trajectory prediction model to obtain the robot's embodied flow information output by the robot trajectory prediction model. Based on the robot's standard structural description file, the motion constraint information of each joint of the robot is obtained. Based on the motion constraint information of each joint, as well as the intrinsic and extrinsic parameters of the camera, the body flow information is decomposed into motion components of each joint through forward kinematic mapping. Based on the motion components, the trajectory is sent.

2. The method for predicting the trajectory of an unannotated robot based on embodied flow representation according to claim 1, characterized in that, The training steps of the robot trajectory prediction model include: Obtain multimodal conditional features of the samples and an initial robot trajectory prediction model; the initial robot trajectory prediction model includes an embodied flow prediction model and a target image prediction model; The sample multimodal conditional features are input into the embodied flow prediction model, which determines the first predicted noise distribution of the pixel displacement of the embodied point in the image sequence, and the embodied flow prediction result is determined based on the first predicted noise distribution. The embodied flow prediction result and the sample multimodal conditional features are fused to obtain conditional prior features, and the conditional prior features are input into the target image prediction model, which then determines the second prediction noise distribution of the target image. Based on the difference between the first predicted noise distribution and the first true noise distribution, the body flow loss is determined, and based on the difference between the second predicted noise distribution and the second true noise distribution, the image prediction loss is determined. Based on the embodied flow loss and the image prediction loss, a target loss is determined, and the parameters of the initial robot trajectory prediction model are iterated based on the target loss to obtain the robot trajectory prediction model.

3. The method for predicting the trajectory of a robot without motion annotation based on embodied flow representation according to claim 2, characterized in that, The first predicted noise distribution is a positive noise distribution.

4. The method for predicting the trajectory of an unannotated robot based on embodied flow representation according to any one of claims 1 to 3, characterized in that, The robot obtains motion constraint information for each joint based on the standard structural description file of the robot, and decomposes the body flow information into motion components of each joint through forward kinematic mapping based on the motion constraint information of each joint and the intrinsic and extrinsic parameters of the camera, including: Based on the robot's standard structural description file, the motion constraint information of each joint of the robot is obtained, and the motion constraint information of each joint, the intrinsic and extrinsic parameters of the camera, and the body flow information are input into the joint motion prediction model to obtain the motion components of each joint output by the joint motion prediction model. The training steps of the joint motion prediction model include: Obtain the sample body flow information and the initial joint motion prediction model; The motion constraint information of each joint, the intrinsic and extrinsic parameters of the camera, and the sample embodied flow information are input into the initial joint motion prediction model. The initial joint motion prediction model determines the three-dimensional trajectory of each point in each joint at the current moment, as well as the two-dimensional projection position of the next moment. Based on the three-dimensional trajectory and the two-dimensional projection position, the reprojection error is determined. Based on the reprojection error, the parameters of the initial joint motion prediction model are iterated to obtain the joint motion prediction model.

5. The method for predicting the trajectory of a robot without motion annotation based on embodied flow representation according to claim 4, characterized in that, The determination of reprojection error based on the three-dimensional trajectory and the two-dimensional projection position includes: The reprojection error is determined based on the pose of the target actuator, the transformation relationship between each joint and the target actuator, the three-dimensional trajectory, and the two-dimensional projection position.

6. The method for predicting the trajectory of an unannotated robot based on embodied flow representation according to any one of claims 1 to 3, characterized in that, The steps for determining the initial point position in the robot's embodied flow include: Obtain the initial image frame; Based on a multimodal segmentation model, the physical region where the robot's robotic arm is located in the initial image frame is identified; A predetermined number of pixels are randomly selected within the mask area of ​​the embodied region as embodied key points; The position of the initial point is obtained by tracking the position of the embodied keypoint in the initial image frame using a pre-trained point tracking model.

7. A device for predicting the trajectory of a robot without motion annotation based on embodied flow representation, characterized in that, include: The acquisition unit is used to acquire the initial point position in the robot's embodied flow, the text features in the language task instructions, and the visual features in the initial image; The input unit is used to take the initial point position, the text features and the visual features as multimodal conditional features, and input the multimodal conditional features into the robot trajectory prediction model to obtain the robot's embodied flow information output by the robot trajectory prediction model; The trajectory sending unit is used to obtain the motion constraint information of each joint of the robot based on the robot's standard structural description file, and based on the motion constraint information of each joint, as well as the intrinsic and extrinsic parameters of the camera, decompose the body flow information into motion components of each joint through forward kinematic mapping, and send the trajectory based on the motion components.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method for predicting the motion-unannotated robot operation trajectory based on embodied flow representation as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for predicting the motion-unannotated robot operation trajectory based on embodied flow representation as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for predicting the motion-unannotated robot operation trajectory based on embodied flow representation as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Three-dimensional continuous shape estimation method and system for flexible robot

    CN121018669A

  • A method and system for three-dimensional continuous shape estimation of a flexible robot

    CN121018669B