Robot action prediction method utilizing displacement vector traction and fusing historical information

By fusing historical information and displacement vectors in the robot motion prediction method, using the Transformer architecture and multimodal fusion encoder, the stuck problem caused by the lack of historical information memory function of robots in the prior art is solved, and the task success rate and flexibility of action prediction are improved.

CN119952703AActive Publication Date: 2025-05-09NORTHEASTERN UNIV FOSHAN GRADUATE SCHOOL OF INNOVATION

Patent Information

Application Number
CN202510169241.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-09
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

The existing robot action prediction method based on artificial intelligence large models lacks historical information memory function during task execution, resulting in the problem that the robot may be stuck when facing occlusion or environmental changes.

Method used

A robot action prediction method that uses displacement vectors to pull and fuse historical information is adopted. By collecting training data in a simulation environment, using a multimodal fusion encoder and a history information encoder to generate encoded information, and inputting this information into the Transformer architecture, combining action blocking and time integration technology to predict the end position of the robot's future multiple moments.

Benefits of technology

It improves the task success rate, solves the problem of robots being stuck when facing occlusion or environmental changes, and achieves more flexible and effective action prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119952703A_ABST
    Figure CN119952703A_ABST
Patent Text Reader

Abstract

The invention discloses a robot action prediction method utilizing displacement vector traction and fusing historical information, which comprises the following steps: collecting training data in a simulation environment, including text description of a task executed by a mechanical arm and gray level images of the tail end and the head of the mechanical arm; a multi-modal fusion encoder is used for carrying out multi-modal fusion encoding on the text description and the grayscale image, and multi-modal fusion encoding information is output; a historical information encoder is adopted to encode the multi-modal fusion encoding information and the historical joint posture information to obtain historical encoding information; calculating a joint displacement vector and an end effector displacement vector; the multi-modal fusion coding information, the displacement vector, the robot joint posture information, the end effector posture and historical coding information are input into a transformer architecture, the action partitioning and time integration technology in ACT is adopted, the end poses of the robot at multiple moments in the future are predicted, and the robot joint angles at the multiple moments in the future are obtained through inverse kinematics conversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of robot motion prediction, and relates to a robot motion prediction method that utilizes displacement vector traction and fuses historical information. Background Art

[0002] In recent years, artificial intelligence big models have developed rapidly and have been applied to many fields such as robotics. However, the existing robot motion prediction methods based on artificial intelligence big models are not yet mature and there are many problems.

[0003] One of the most advanced technologies currently available is Stanford University's Action Block Generative Model (ACT). This end-to-end robot task learning large model has a low overall success rate and has no historical information memory function during task execution. It only relies on the current frame to judge and predict the next action sequence. If the target object is blocked during the operation, the task cannot be completed. The existing BeT algorithm that uses historical information only uses image information, which is insufficient for describing the current environmental state.

[0004] In order to solve this problem, which involves the hindrance of imitation problems in the end-to-end autonomous driving model, if historical information is directly added to the Transformer backbone model, it is easy to cause a significant decline in model performance. The end-to-end robot task execution model will habitually predict the state of the next frame based on the state of the previous frame. If there is a certain pause in the process of task execution, the robot will be stuck in the process of model reasoning and prediction because it has learned this strategy lazily and does not perform the real operation, and does not deal with sudden or changing situations in the environment, causing the robotic arm to be in a fixed position and not change its posture. Summary of the invention

[0005] The present invention provides a robot motion prediction method that utilizes displacement vector traction and integrates historical information, so as to solve the problem that the existing model does not use historical information or only uses historical image information, so that when the model faces certain input situations, the robot's predicted motion may change within a very small range, causing the robot to get stuck and not move.

[0006] The present invention provides a robot motion prediction method using displacement vector traction and fusing historical information, comprising:

[0007] Step 1: Collect training data in a simulation environment, including text descriptions of the tasks performed by the robot arm and grayscale images of the end and head of the robot arm;

[0008] Step 2: Use a multimodal fusion encoder to perform multimodal fusion encoding of text description and grayscale image, and output multimodal fusion encoding information;

[0009] Step 3: Use the historical information encoder to encode the multimodal fusion encoding information and the historical joint posture information to obtain the historical encoding information;

[0010] Step 4: Calculate the joint displacement vector and the end effector displacement vector;

[0011] Step 5: Input the multimodal fusion coding information, joint displacement vector, end effector displacement vector, robot joint posture information, end effector posture and historical coding information into the transformer architecture, and use the action block and time integration technology in ACT to predict the end posture of the robot at multiple moments in the future, and obtain the robot joint angles at multiple moments in the future through inverse kinematics conversion.

[0012] The robot motion prediction method of the present invention, which utilizes displacement vector traction and integrates historical information, has the following beneficial effects:

[0013] The present invention collects training data in a simulation environment and generates multimodal fusion coding information using a multimodal fusion encoder. At the same time, a historical information encoder is used to encode to obtain historical coding information. The historical coding information, historical multimodal fusion coding information, joint displacement vectors, and end effector displacement vectors are input into the transformer architecture, and the action block and time integration technology in ACT are used to predict the end position of the robot at multiple moments in the future, and the robot joint angles at multiple moments in the future are obtained through inverse kinematics conversion, thereby improving the task success rate and solving the problem of the robot being stuck. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a flow chart of the robot motion prediction method of the present invention which utilizes displacement vector traction and integrates historical information. DETAILED DESCRIPTION

[0015] like Figure 1 As shown, a robot motion prediction method of the present invention using displacement vector traction and fusing historical information includes:

[0016] Step 1: Collect training data in a simulation environment, including text descriptions of the tasks performed by the robot arm and grayscale images of the end of the robot arm and head.

[0017] In the specific implementation, the training data set was obtained through the Coppeliasim simulation environment, a grasping task was designed, and the text instructions, camera images, and robot posture information during the simulation environment demonstration were collected. By executing multiple tasks, a certain scale of data set was collected. The image size is (160, 120), and the robot posture information includes: joint angles and end postures.

[0018] Step 2: Use the multimodal fusion encoder to perform multimodal fusion encoding of text description and grayscale image, and output multimodal fusion encoding information, specifically:

[0019] Step 2.1: Establish a multimodal fusion encoder, including a BERT model after knowledge distillation, a multi-layer perceptron and a visual encoder, wherein the visual encoder is composed of multiple residual blocks.

[0020] Step 2.2: Use the knowledge distilled BERT model to encode the text description of the task and obtain text encoding information.

[0021] Step 2.3: Input the text encoding information into the multilayer perceptron and obtain two coefficient tensors through two linear scaling learning.

[0022] Step 2.4: Input the two grayscale images into the visual encoder and extract the visual encoding information through the residual block.

[0023] Step 2.5: Multiply the text encoding information with the two coefficient tensors and then add it to the visual encoding information and input it to the next residual block; each residual block in the visual encoder performs the same operation and finally outputs a 3-dimensional tensor (512, 4, 5) as the multimodal fusion encoding information.

[0024] Step 3: Use the historical information encoder to encode the multimodal fusion encoding information and the historical joint posture information to obtain the historical encoding information, specifically:

[0025] Step 3.1: Build a history information encoder, including a compression module, a linear layer, and a transformer encoder.

[0026] Step 3.2: Input the multimodal fusion coding information into the compression module, linearly stretch the second and third dimensions to obtain two-dimensional multimodal fusion coding information (512, 20), and compress the second dimension by taking the average value to obtain the compressed multimodal fusion coding information (512, 1).

[0027] Step 3.3: Transform all joint posture information before the current moment into 512 dimensions through a linear layer.

[0028] Step 3.4: Concatenate the joint posture information of k moments, the compressed multimodal fusion coding information, and an all-zero tensor (1, 512) and input them into the transformer encoder. Take the encoded output of the all-zero tensor and convert it into a tensor (1, 1024) through a linear layer. Predict its variance and mean, and then reparameterize the output into a tensor (1, 512) as the historical coding information.

[0029] Step 4: Calculate the joint displacement vector and the end effector displacement vector, specifically:

[0030] The difference between the joint posture information at the previous moment and the joint posture information of the robot at the initial position is taken as the joint displacement vector, and the difference between the posture of the end effector at the previous moment and the posture of the initial position is taken as the end effector displacement vector.

[0031] Step 5: Input the multimodal fusion coding information, joint displacement vector, end effector displacement vector, robot joint posture information, end effector position and historical coding information into the transformer architecture, use the action block and time integration technology in ACT to predict the robot's end position at multiple moments in the future, and obtain the robot's joint angles at multiple moments in the future through inverse kinematics conversion, specifically:

[0032] Step 5.1: Input the multimodal fusion coding information, displacement vector, robot joint posture information, end effector posture and historical coding information into the transformer architecture, and use the action segmentation and time integration technology in ACT to predict the end posture of the robot at multiple moments in the future.

[0033] Step 5.2: Perform a weighted sum of the terminal postures at a certain moment in the future predicted by multiple historical moments to calculate the final terminal posture at a certain moment in the future, specifically:

[0034] The action block and time integration technology in ACT is used, that is, the terminal posture of a total of x moments in the future is predicted each time, and the number of prediction time steps is set to be the same as the number of historical encoding time steps, that is, x = k, which effectively avoids the robot from freezing due to long model calculation time; the exponential weighting scheme is used to perform weighted averaging on the predicted action sequence:

[0035]

[0036] in, It is the weight of the action at the tth moment solved based on the information at the i-th moment. m=0.01 is an adjustable parameter. The larger the m, the lower the weight of the action at a later time.

[0037] The final terminal posture at a certain moment in the future is calculated by weighted summing the terminal postures at a certain moment in the future predicted by multiple historical moments:

[0038]

[0039] in, is the action at the tth moment predicted at the jth moment.

[0040] Step 5.3: The final end position at multiple moments in the future is converted by inverse kinematics to obtain the robot joint angles at multiple moments in the future.

[0041] In the specific implementation, a composite loss function is used to train the model. One part is to calculate the KL divergence of the mean and variance solved in the action history information encoding and the Gaussian distribution commonly found in life, so that the distribution solved in the history information encoding is close to the Gaussian distribution; the other part is to calculate the MSE loss of the predicted action sequence and the actual action sequence.

[0042]

[0043] in is the end effector pose predicted by the model at k moments in the future after t, a t,t+k is the actual end effector pose in the training data at k moments after t, β is an adjustable parameter, It is the historical information encoder based on the currently observed environmental state information and historical action information a t-k,t The z distribution formed by the mean and variance obtained after encoding, N(0,1) is a Gaussian distribution.

[0044] The above description is only a preferred embodiment of the present invention and is not intended to limit the concept of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A robot motion prediction method using displacement vector traction and fusing historical information, characterized in that: include: Step 1: Collect training data in a simulation environment, including text descriptions of the tasks performed by the robot arm and grayscale images of the end and head of the robot arm; Step 2: Use a multimodal fusion encoder to perform multimodal fusion encoding of text description and grayscale image, and output multimodal fusion encoding information; Step 3: Use the historical information encoder to encode the multimodal fusion encoding information and the historical joint posture information to obtain the historical encoding information; Step 4: Calculate the joint displacement vector and the end effector displacement vector; Step 5: Input the multimodal fusion coding information, joint displacement vector, end effector displacement vector, robot joint posture information, end effector posture and historical coding information into the transformer architecture, and use the action block and time integration technology in ACT to predict the end posture of the robot at multiple moments in the future, and obtain the robot joint angles at multiple moments in the future through inverse kinematics conversion.

2. The robot motion prediction method using displacement vector traction and fusing historical information as claimed in claim 1, characterized in that: The step 2 is specifically as follows: Step 2.1: Establish a multimodal fusion encoder, including a BERT model after knowledge distillation, a multi-layer perceptron and a visual encoder, wherein the visual encoder is composed of multiple residual blocks; Step 2.2: Use the knowledge distilled BERT model to encode the text description of the task to obtain text encoding information; Step 2.3: Input the text encoding information into the multilayer perceptron and obtain two coefficient tensors through two linear scaling learning; Step 2.4: Input the two grayscale images into the visual encoder and obtain the visual encoding information through residual block extraction; Step 2.5: Multiply the text encoding information with the two coefficient tensors and then add it to the visual encoding information and input it to the next residual block; each residual block in the visual encoder performs the same operation and finally outputs a 3-dimensional tensor (512, 4, 5) as the multimodal fusion encoding information.

3. The robot motion prediction method using displacement vector traction and fusing historical information as claimed in claim 1, characterized in that: The step 3 is specifically as follows: Step 3.1: Establish a history information encoder, including a compression module, a linear layer, and a transformer encoder; Step 3.2: Input the multimodal fusion coding information into the compression module, linearly stretch the second and third dimensions to obtain the multimodal fusion coding information of two dimensions (512, 20), and compress the second dimension by taking the average value to obtain the compressed multimodal fusion coding information (512, 1); Step 3.3: Transform all joint posture information before the current moment into 512 dimensions through a linear layer; Step 3.4: Concatenate the joint posture information of k moments, the compressed multimodal fusion coding information, and an all-zero tensor (1, 512) and input them into the transformer encoder. Take the encoded output of the all-zero tensor and convert it into a tensor (1, 1024) through a linear layer. Predict its variance and mean, and then reparameterize the output into a tensor (1, 512) as the historical coding information.

4. The robot motion prediction method using displacement vector traction and fusing historical information as claimed in claim 1, characterized in that: The step 4 is specifically as follows: The difference between the joint posture information at the previous moment and the joint posture information of the robot at the initial position is taken as the joint displacement vector, and the difference between the posture of the end effector at the previous moment and the posture of the initial position is taken as the end effector displacement vector.

5. The robot motion prediction method using displacement vector traction and fusing historical information as claimed in claim 1, characterized in that: The step 5 is specifically as follows: Step 5.1: Input the multimodal fusion coding information, displacement vector, robot joint posture information, end effector posture and historical coding information into the transformer architecture, and use the action block and time integration technology in ACT to predict the end posture of the robot at multiple moments in the future; Step 5.2: Perform weighted summation of the terminal postures at a certain future moment predicted by multiple historical moments to calculate the final terminal posture at a certain future moment; Step 5.3: The final end position at multiple moments in the future is converted by inverse kinematics to obtain the robot joint angles at multiple moments in the future.

6. The robot motion prediction method using displacement vector traction and fusing historical information as claimed in claim 5, characterized in that: The step 5.2 is specifically as follows: The action block and time integration technology in ACT is used, that is, the terminal posture of a total of x moments in the future is predicted each time, and the number of prediction time steps is set to be the same as the number of historical encoding time steps, that is, x = k, which effectively avoids the robot from freezing due to long model calculation time; the exponential weighting scheme is used to perform weighted averaging on the predicted action sequence: in, is the weight of the action at time t obtained based on the information at time i. m = 0.01 is an adjustable parameter. The larger the m, the lower the weight of the action at a later time. The final terminal posture at a certain moment in the future is calculated by weighted summing the terminal postures at a certain moment in the future predicted by multiple historical moments: in, is the action at the tth moment predicted at the jth moment.

Citation Information

Patent Citations

  • Industrial robot pose precision online compensation method based on deep reinforcement learning

    CN113510709A

  • Control method and device of mobile operation robot and mobile operation robot

    CN117359633A

  • Mechanical arm grabbing prediction method, electronic equipment and storage medium

    CN119168835A

Cited By

  • Multi-mode tool body intelligent control method and device and electronic equipment

    CN121121698A

  • A multimodal embodied intelligent control method and device and electronic equipment

    CN121121698B