Robot imitation learning method based on intention understanding and subconsciousness execution

By introducing methods of intention understanding and subconscious execution in robot imitation learning, using deep learning and dynamic downsampling technology, the problem of insufficient robot execution efficiency and accuracy in the existing technology is solved, and more efficient and precise task execution is achieved.

CN119952726AActive Publication Date: 2025-05-09NORTHEASTERN UNIV FOSHAN GRADUATE SCHOOL OF INNOVATION

Patent Information

Application Number
CN202510386966.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-05-09
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

Existing robot imitation learning methods have challenges in execution efficiency, real-time and accuracy, especially in high-dimensional states and action spaces. The consumption of computing resources and multi-step prediction errors seriously affect the execution efficiency and accuracy of robot tasks.

Method used

A robot imitation learning method based on intention understanding and subconscious execution is adopted. By collecting robot demonstration trajectory data in the actual environment, calculating the intent impact measurement, and using smoothing filters and dynamic downsampling to generate the conscious trajectory data set, a deep learning-based action prediction model is established, and combined with the subconscious imitation learning mechanism, action prediction and execution are performed.

Benefits of technology

It improves the efficiency and accuracy of robot task execution, reduces the cumulative error of redundant calculations and multi-step prediction, improves the real-time and execution efficiency of the system, and solves the problems of robot being stuck and overloaded in computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119952726A_ABST
    Figure CN119952726A_ABST
Patent Text Reader

Abstract

The invention provides a robot imitation learning method based on intention understanding and subconsciousness execution, and the method comprises the steps: collecting a plurality of robot demonstration trajectory data to form a data set, and calculating the intention influence measurement according to the robot demonstration trajectory data; performing smoothing processing on the intention influence measurement by using a smoothing filter, calculating a dynamic down-sampling frequency, performing down-sampling on the robot demonstration trajectory data, and generating a consciousness trajectory data set; establishing an action prediction model based on deep learning, and training the action prediction model based on deep learning by using the consciousness trajectory data set; inputting current observation data, sensor data and action data into the trained action prediction model, outputting a predicted future action sequence, and predicting an execution action at the next moment through a cognitive unloading mechanism and a weighted accumulation prediction strategy; and the predicted execution action at the next moment is fed back to a robot actuator, new observation data are obtained, and the steps are repeated until the target task is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of robot imitation learning, and relates to a robot imitation learning method based on intention understanding and subconscious execution. Background Art

[0002] In the field of robot control, robot imitation learning is a key technology that automatically acquires task control strategies by learning from expert demonstrations. It is widely used in industrial production, service robots and other fields. In complex operation tasks, robots can complete tasks such as grasping, assembly, and precision manufacturing by imitating human operation demonstrations. However, existing imitation learning methods still face many challenges in practical applications, especially in terms of execution efficiency, real-time performance and accuracy.

[0003] Traditional robot imitation learning methods usually rely on a lot of calculations and multi-step predictions, especially in high-dimensional state and action spaces. The consumption of computing resources and multi-step prediction errors seriously affect the efficiency and accuracy of robot task execution. In addition, these methods fail to effectively solve the problem of redundant data, resulting in the need to process a large amount of irrelevant data during task execution, further increasing the computational burden and affecting the robot's real-time response capability.

[0004] At present, although there are action prediction models based on deep learning and transformer architecture that can improve prediction accuracy to a certain extent, these methods still face the problems of redundant data processing and real-time computing efficiency. In a dynamically changing environment, how to balance task accuracy and execution efficiency is still a difficult problem that needs to be solved.

[0005] As a learning mechanism that simulates human subconscious operations, subconscious imitation learning has been proposed as a solution to improve execution efficiency by reducing unnecessary calculations through cognitive offloading. However, existing subconscious imitation learning methods have not fully solved the problems of how to effectively reduce redundant calculations and avoid cumulative errors in multi-step predictions while ensuring task accuracy. Summary of the invention

[0006] In order to solve the above technical problems, the purpose of the present invention is to provide a robot imitation learning method based on intention understanding and subconscious execution, aiming to improve the efficiency and accuracy of robot task execution.

[0007] A robot imitation learning method based on intention understanding and subconscious execution of the present invention comprises:

[0008] Step 1: Collect multiple robot demonstration trajectory data in the actual environment to form a data set, and calculate the intention impact measure based on the robot demonstration trajectory data;

[0009] Step 2: Use a smoothing filter to smooth the intention impact measure, calculate the dynamic downsampling frequency, downsample the robot demonstration trajectory data based on the dynamic downsampling frequency, and generate a consciousness trajectory dataset;

[0010] Step 3: Establish a deep learning-based action prediction model and use the consciousness trajectory dataset to train the deep learning action prediction model;

[0011] Step 4: In the robot action execution phase, the current observation data, sensor data and action data are processed through steps 1 and 2 and input into the trained action prediction model, and the predicted future action sequence is output. The next execution action is predicted through the cognitive offloading mechanism and weighted cumulative prediction strategy;

[0012] Step 5: Feedback the predicted execution action at the next moment to the robot actuator, obtain new observation data, and repeat steps 4 and 5 until the target task is completed.

[0013] Furthermore, the step 1 is specifically as follows:

[0014] Step 1.1: Collect multiple robot demonstration trajectory data in the actual environment to form a dataset D = {τ0, τ1, …, τ N}, each trajectory τ n ,n=0,1,…,N consists of time steps t, including the robot’s visual observation data O (n,t) , sensor data S (n,t) and action data A (n,t) ;

[0015] The visual observation data is the image data of the robot working environment; the sensor data includes the joint position qpos (n,t) , joint speed qvel (n,t) and end effector torque eeft (n,t) ; Action data includes the target joint position α pos(n,t) and the target joint velocity α qvel(n,t) ;

[0016] Step 1.2: Calculate the intention impact measure I according to the following formula t :

[0017]

[0018] in, is a normalization function and ω is a weight vector that determines the relative importance of joints.

[0019] Furthermore, the step 2 is specifically as follows:

[0020] Step 2.1: Use the Butterworth filter to smooth the intention impact measure, the formula is as follows:

[0021]

[0022] Among them, s is a complex frequency variable, which is used to describe the frequency characteristics of the signal and the dynamic behavior of the system; m is the order of the Butterworth filter, which determines the steepness of the filter, ω c is the cut-off frequency;

[0023] Step 2.2: Calculate the dynamic downsampling frequency based on the smoothed intention impact measure and the sampling frequency of the robot demonstration trajectory data:

[0024]

[0025] Among them, f ds (t) is the dynamic downsampling frequency, f d is the original sampling frequency of the robot demonstration trajectory data, is a round-down operation, M represents the time scaling factor, f m is the target minimum sampling frequency;

[0026] Step 2.3: Downsample the original robot demonstration trajectory data according to the dynamic downsampling frequency to generate the consciousness trajectory dataset D ds ={τ ds(0) ,τ ds(1) ,…,τ ds(N)} for subsequent action prediction model training.

[0027] Furthermore, the step 3 is specifically as follows:

[0028] Step 3.1: Subtract each downsampled consciousness trajectory τ ds(n) Visual observation data O (n,t) Input into the pre-trained ResNet34 convolutional neural network to extract visual features F I , and then the visual feature F I and trajectory τ ds(n) The sensor data S in (n,t) The spliced ​​features are then further processed by the multi-layer perceptron MLP to generate a fused feature vector:

[0029] X (n,k) =MLP(concat(F I ,S (n,t) ))

[0030] Among them, concat(F I ,S (n,t)) indicates that F I and S (n,t) For vertical splicing, MLP is a feed-forward neural network consisting of multiple fully connected layers, which can learn complex nonlinear mapping relationships. MLP maps multimodal features to a unified feature space through the nonlinear activation function ReLU;

[0031] Step 3.2: For feature sequence Add position code PosEmb(k) to retain the time order information of the sequence. The position code is generated by sine and cosine functions. The formula is:

[0032]

[0033] Among them, d model is the dimension of the feature vector, i is the dimension index, T ds is the length of the consciousness trajectory generated by downsampling;

[0034] Step 3.3: Pass the feature sequence with position encoding added to the Transformer encoder. The encoder consists of multiple self-attention layers and feedforward neural network layers. The self-attention mechanism captures global context information by calculating the relationship between each time step in the feature sequence. The feature sequence processed by the encoder is calculated as follows:

[0035] H (n) =Transformer Encoder (X (n,k) +PosEmb(k))

[0036] Step 3.4: Feature sequence H after encoder processing (n) and historical action data A (n,k-1) The decoder is composed of multiple self-attention layers, cross-attention layers, and feed-forward neural network layers. The self-attention layer is used to capture the temporal dependencies in the historical action data, and the cross-attention layer is used to transform H (n) With A (n,k-1) Align and generate target action A (n,k) ; Using autoregression, we gradually predict the future action sequence, and the action data A generated at each step (n,k) As the input for the next step, it is used to predict the next action A (n,k+1) , the predicted target action in the future time step can be expressed as:

[0037] A (n,k) =Transformer Eecoder (H (n) ,A (n,k-1) )

[0038] Step 3.5: Calculate the mean square error (MSE) between the predicted action and the true action. The loss function is expressed as:

[0039]

[0040] Update the model parameters through the back-propagation algorithm to minimize the loss function L action ,Using the optimizer, Adam adjusts the model weights to ensure that the model can converge quickly.

[0041] Furthermore, the step 4 is specifically as follows:

[0042] Step 4.1: Get the current observation data, sensor data and action data, and process them using the methods in steps 1 and 2, and then input them into the trained action prediction model to predict future action sequences:

[0043] a T ={a T [T+1],a T [T+2],…,a T [T+K]}

[0044] Where K is the length of the action block; the next action to be executed is A T+1 It has been predicted K times in the past, which is represented by the prediction set:

[0045] As T+1 ={a (T-1) (T+1),a (T-2) (T+1),…,a (T-K+1) (T+1)}

[0046] Step 4.2: Perform the next execution action A through cognitive offloading mechanism and weighted cumulative prediction strategy T+1 Prediction:

[0047]

[0048] Among them, e -mk Represents the time weighting factor, m is the weight; COR T+1 It is a cognitive unloading mechanism used to calculate the degree of joint position overlap;

[0049]

[0050] Among them, As (T+1,;j) represents the predicted sequence of the jth joint action, μ j and σ j As (T+1,;j) The mean and standard deviation of , n is the number of robot joints, std(·) means to find the standard deviation; COR is the set threshold.

[0051] A robot imitation learning method based on intention understanding and subconscious execution has the following beneficial effects:

[0052] 1. During the data collection and processing stage, the method of the present invention analyzes the trajectory data through the intention impact measurement, reduces redundant information, and optimizes the data processing process.

[0053] 2. During the trajectory downsampling stage, the method of the present invention dynamically calculates the downsampling frequency, prioritizes the retention of mission-critical data, and reduces unnecessary calculations and data storage.

[0054] 3. In the action prediction stage, the method of the present invention adopts an action prediction model based on a transformer architecture, combined with a subconscious imitation learning mechanism, which effectively improves the accuracy and speed of task execution.

[0055] 4. During the action execution phase, the method of the present invention reduces unnecessary reasoning calculations through a cognitive offloading mechanism, thereby improving the real-time performance and execution efficiency of the system.

[0056] The present invention significantly improves the execution efficiency of robots in complex tasks, solves the problems of robots getting stuck and having excessive computational burden, has broad application prospects, and is particularly suitable for the fields of industrial and service robots with high real-time requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a flow chart of a robot imitation learning method based on intention understanding and subconscious execution of the present invention. DETAILED DESCRIPTION

[0058] like Figure 1 As shown, a robot imitation learning method based on intention understanding and subconscious execution of the present invention includes:

[0059] Step 1: Collect multiple robot demonstration trajectory data in the actual environment to form a data set, and calculate the intention impact measure based on the robot demonstration trajectory data, specifically:

[0060] Step 1.1: Collect multiple robot demonstration trajectory data in the actual environment to form a dataset D = {τ0, τ1, …, τ N}, each trajectory τ n ,n=0,1,…,N consists of time steps t, including the robot’s visual observation data O (n,t) , sensor data S (n,t) and action data A (n,t) .

[0061] The visual observation data is the image data of the robot working environment; the sensor data includes the joint position qpos (n,t), joint speed qvel (n,t) and end effector torque eeft (n,t) ; Action data includes the target joint position α pos(n,t) and the target joint velocity α qvel(n,t) .

[0062] Step 1.2: Calculate the intention impact measure I according to the following formula t :

[0063]

[0064] in, is a normalization function and ω is a weight vector that determines the relative importance of joints.

[0065] For example, in practical applications, ω is set to [0.1, 0.1, 0.1, 0.1, 0.1, 0.3, 0.1,]. This metric can quantify and prioritize the pattern information that is critical to task execution, thereby ensuring the effectiveness of subsequent data processing.

[0066] Step 2: Use a smoothing filter to smooth the intention impact measure, calculate the dynamic downsampling frequency, downsample the robot demonstration trajectory data based on the dynamic downsampling frequency, and generate the consciousness trajectory dataset, specifically:

[0067] Step 2.1: Use the Butterworth filter to smooth the intention impact measure, the formula is as follows:

[0068]

[0069] Where s is a complex frequency variable used to describe the frequency characteristics of the signal and the dynamic behavior of the system; m is the order of the Butterworth filter, which determines the steepness of the filter. The higher the order, the steeper the transition band of the filter and the better the filtering effect, but the higher the computational complexity, usually set to 4 or 6. c is the cut-off frequency.

[0070] If the sampling frequency is 100Hz and the cutoff frequency is 10Hz, the normalized cutoff frequency is:

[0071]

[0072] Step 2.2: Calculate the dynamic downsampling frequency based on the smoothed intention impact measure and the sampling frequency of the robot demonstration trajectory data:

[0073]

[0074] Among them, f ds (t) is the dynamic downsampling frequency, f dis the original sampling frequency of the robot demonstration trajectory data, is a round-down operation, M represents the time scaling factor, f m is the target minimum sampling frequency.

[0075] Step 2.3: Downsample the original robot demonstration trajectory data according to the dynamic downsampling frequency to generate the consciousness trajectory dataset D ds ={τ ds(0) ,τ ds(1) ,…,τ ds(N)}, for subsequent action prediction model training.

[0076] Step 3: Establish a deep learning-based action prediction model and use the consciousness trajectory dataset to train the deep learning action prediction model. The action prediction model uses the encoder and decoder in the Transformer architecture to predict future actions. The model accurately predicts the target action in the future time step by pattern recognition of historical trajectory data. Specifically:

[0077] Step 3.1: Subtract each downsampled consciousness trajectory τ ds(n) Visual observation data O (n,t) Input into the pre-trained ResNet34 convolutional neural network to extract visual features F I , and then the visual feature F I and trajectory τ ds(n) The sensor data S in (n,t) The spliced ​​features are then further processed by the multi-layer perceptron MLP to generate a fused feature vector:

[0078] X (n,k) =MLP(concat(F I ,S (n,t) ))

[0079] Among them, concat(F I ,S (n,t) ) indicates that F I and S (n,t) For vertical splicing, MLP is a feedforward neural network composed of multiple fully connected layers. It can learn complex nonlinear mapping relationships. MLP maps multimodal features to a unified feature space through the nonlinear activation function ReLU.

[0080] Step 3.2: For feature sequence Add position code PosEmb(k) to retain the time order information of the sequence. The position code is generated by sine and cosine functions. The formula is:

[0081]

[0082] Among them, d model is the dimension of the feature vector, i is the dimension index, T ds is the length of the consciousness trajectory generated by downsampling.

[0083] Step 3.3: Pass the feature sequence with position encoding added to the Transformer encoder. The encoder consists of multiple self-attention layers and feedforward neural network layers. The self-attention mechanism captures global context information by calculating the relationship between each time step in the feature sequence. The feature sequence processed by the encoder is calculated as follows:

[0084] H (n) =Transformer Encoder (X (n,k) +PosEmb(k))

[0085] Step 3.4: Feature sequence H after encoder processing (n) and historical action data A (n,k-1) The decoder is composed of multiple self-attention layers, cross-attention layers, and feed-forward neural network layers. The self-attention layer is used to capture the temporal dependencies in the historical action data, and the cross-attention layer is used to transform H (n) With A (n,k-1) Align and generate target action A (n,k) ; Using autoregression, we gradually predict the future action sequence, and the action data A generated at each step (n,k) As the input for the next step, it is used to predict the next action A (n,k+1) , the predicted target action in the future time step is expressed as:

[0086] A (n,k) =Transformer Eecoder (H (n) ,A (n,k-1) )

[0087] Step 3.5: Calculate the mean square error (MSE) between the predicted action and the true action. The loss function is expressed as:

[0088]

[0089] Update the model parameters through the back-propagation algorithm to minimize the loss function L action ,Using the optimizer, Adam adjusts the model weights to ensure that the model can converge quickly.

[0090] Step 4: In the robot action execution phase, the current observation data, sensor data and action data are processed through steps 1 and 2 and input into the trained action prediction model, and the predicted future action sequence is output. The next execution action is predicted through the cognitive unloading mechanism and weighted cumulative prediction strategy. Specifically:

[0091] Step 4.1: Get the current observation data, sensor data and action data, and process them using the methods in steps 1 and 2, and then input them into the trained action prediction model to predict future action sequences:

[0092] a T ={a T [T+1],a T [T+2],…,a T [T+K]}

[0093] Where K is the length of the action block; the next action to be executed is A T+1 It has been predicted K times in the past, which is represented by the prediction set:

[0094] As T+1 ={a (T-1) (T+1),a (T-2) (T+1),…,a (T-K+1) (T+1)}

[0095] Step 4.2: Perform the next execution action A through cognitive offloading mechanism and weighted cumulative prediction strategy T+1 Prediction:

[0096]

[0097] Among them, e -mk Represents the time weighting factor, m is the weight; COR T+1 It is a cognitive unloading mechanism used to calculate the degree of joint position overlap;

[0098]

[0099] Among them, As (T+1,;j) represents the predicted sequence of the jth joint action, μ j and σ j As (T+1,;j) The mean and standard deviation of , n is the number of robot joints, std(·) means to find the standard deviation; COR is the set threshold.

[0100] In the above step 4.2, the overlapping joint position information of the historical trajectory is evaluated to determine whether the step can be skipped, and the current execution action prediction A is directly generated. T+1 =a (T-1)(T+1). If the historical predicted joint position exceeds the set threshold, COR T+1 >COR, the subsequent weighted cumulative prediction process can be skipped to reduce the computational burden.

[0101] Step 5: Feedback the predicted execution action at the next moment to the robot actuator, obtain new observation data, and repeat steps 4 and 5 until the target task is completed.

[0102] The above description is only a preferred embodiment of the present invention and is not intended to limit the concept of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A robot imitation learning method based on intention understanding and subconscious execution, characterized in that: include: Step 1: Collect multiple robot demonstration trajectory data in the actual environment to form a data set, and calculate the intention impact measure based on the robot demonstration trajectory data; Step 2: Use a smoothing filter to smooth the intention impact measure, calculate the dynamic downsampling frequency, downsample the robot demonstration trajectory data based on the dynamic downsampling frequency, and generate a consciousness trajectory dataset; Step 3: Establish a deep learning-based action prediction model and use the consciousness trajectory dataset to train the deep learning action prediction model; Step 4: In the robot action execution phase, the current observation data, sensor data and action data are processed through steps 1 and 2 and input into the trained action prediction model, and the predicted future action sequence is output. The next execution action is predicted through the cognitive offloading mechanism and weighted cumulative prediction strategy; Step 5: Feedback the predicted execution action at the next moment to the robot actuator, obtain new observation data, and repeat steps 4 and 5 until the target task is completed.

2. The robot imitation learning method based on intention understanding and subconscious execution as claimed in claim 1, characterized in that: The step 1 is specifically as follows: Step 1.1: Collect multiple robot demonstration trajectory data in the actual environment to form a dataset D = {τ0, τ1, …, τ N }, each trajectory τ n ,n=0,1,…,N consists of time steps t, including the robot’s visual observation data O (n,t) , sensor data S (n,t) and action data A (n,t) ; The visual observation data is the image data of the robot working environment; the sensor data includes the joint position qpos (n,t) , joint speed qvel (n,t) and end effector torque eeft (n,t) ; Action data includes the target joint position α pos(n,t) and the target joint velocity α qvel(n,t) ; Step 1.2: Calculate the intention impact measure I according to the following formula t : in, is a normalization function and ω is a weight vector that determines the relative importance of joints.

3. The robot imitation learning method based on intention understanding and subconscious execution as claimed in claim 1, characterized in that: The step 2 is specifically as follows: Step 2.1: Use the Butterworth filter to smooth the intention impact measure, the formula is as follows: Among them, s is a complex frequency variable, which is used to describe the frequency characteristics of the signal and the dynamic behavior of the system; m is the order of the Butterworth filter, which determines the steepness of the filter, ω c is the cut-off frequency; Step 2.2: Calculate the dynamic downsampling frequency based on the smoothed intention impact measure and the sampling frequency of the robot demonstration trajectory data: Among them, f ds (t) is the dynamic downsampling frequency, f d is the original sampling frequency of the robot demonstration trajectory data, is a round-down operation, M represents the time scaling factor, f m is the target minimum sampling frequency; Step 2.3: Downsample the original robot demonstration trajectory data according to the dynamic downsampling frequency to generate the consciousness trajectory dataset D ds ={τ ds(0) ,τ ds(1) ,…,τ ds(N) }, for subsequent action prediction model training.

4. The robot imitation learning method based on intention understanding and subconscious execution as claimed in claim 1, characterized in that: The step 3 is specifically as follows: Step 3.1: Subtract each downsampled consciousness trajectory τ ds(n) Visual observation data O (n,t) Input into the pre-trained ResNet34 convolutional neural network to extract visual features F I , and then the visual feature F I and trajectory τ ds(n) The sensor data S in (n,t) The spliced ​​features are then further processed by the multi-layer perceptron MLP to generate a fused feature vector: X (n,k) =MLP(concat(F I ,S (n,t) )) Among them, concat(F I ,S (n,t) ) indicates that F I and S (n,t) For vertical splicing, MLP is a feed-forward neural network consisting of multiple fully connected layers, which can learn complex nonlinear mapping relationships. MLP maps multimodal features to a unified feature space through the nonlinear activation function ReLU; Step 3.2: For feature sequence Add position code PosEmb(k) to retain the time order information of the sequence. The position code is generated by sine and cosine functions. The formula is: Among them, d model is the dimension of the feature vector, i is the dimension index, T ds is the length of the consciousness trajectory generated by downsampling; Step 3.3: Pass the feature sequence with position encoding added to the Transformer encoder. The encoder consists of multiple self-attention layers and feedforward neural network layers. The self-attention mechanism captures global context information by calculating the relationship between each time step in the feature sequence. The feature sequence processed by the encoder is calculated as follows: H (n) =Transformer Encoder (X (n,k) +PosEmb(k)) Step 3.4: Feature sequence H after encoder processing (n) and historical action data A (n,k-1) The decoder is composed of multiple self-attention layers, cross-attention layers, and feed-forward neural network layers. The self-attention layer is used to capture the temporal dependencies in the historical action data, and the cross-attention layer is used to transform H (n) With A (n,k-1) Align and generate target action A (n,k) ; Using autoregression, we gradually predict the future action sequence, and the action data A generated at each step (n,k) As the input for the next step, used to predict the next action A (n,k+1) , the predicted target action in the future time step can be expressed as: A (n,k) =Transformer Eecoder (H (n) ,A (n,k-1) ) Step 3.5: Calculate the mean square error (MSE) between the predicted action and the true action. The loss function is expressed as: Update the model parameters through the back-propagation algorithm to minimize the loss function L action ,Using the optimizer, Adam adjusts the model weights to ensure that the model can converge quickly.

5. The robot imitation learning method based on intention understanding and subconscious execution as claimed in claim 1, characterized in that: The step 4 is specifically as follows: Step 4.1: Get the current observation data, sensor data and action data, and process them using the methods in steps 1 and 2, and then input them into the trained action prediction model to predict future action sequences: a T ={a T [T+1],a T [T+2],…,a T [T+K]} in, K is the length of the action block; the next action to be performed is A T+1 It has been predicted K times in the past, which is represented by the prediction set: As T+1 ={a (T-1) (T+1),a (T-2) (T+1),…,a (T-K+1) (T+1)} Step 4.2: Perform the next execution action A through cognitive offloading mechanism and weighted cumulative prediction strategy T+1 Prediction: Among them, e -mk represents the time weighting factor, m is the weight; COR T+1 It is a cognitive unloading mechanism used to calculate the degree of joint position overlap; Among them, As (T+1,;j) represents the predicted sequence of the jth joint action, μ j and σ j As (T+1,;j) The mean and standard deviation of , n is the number of robot joints, std(·) means to find the standard deviation; COR is the set threshold.

Citation Information

Patent Citations

  • Simulation learning mechanical arm grabbing method and device based on multi-scale sequence model

    CN116901071A

  • Robot long sequence task learning and planning method based on partial visual observation

    CN117565032A

  • Mechanical arm force feedback control method based on deep reinforcement learning

    CN119550332A

  • System and methods for robotic teleoperation intention estimation

    US20250083325A1

  • Robot learning from demonstration via meta-imitation learning

    WO2022012265A1

Cited By

  • Robot imitation learning method and device based on hybrid perception nonlinear model predictive control

    CN120244988A

  • Robot imitation learning method and device based on hybrid perception nonlinear model predictive control

    CN120244988B

  • Robot mechanical arm cooperative control method based on imitation learning

    CN121132692A

  • A robot manipulator cooperative control method based on imitation learning

    CN121132692B

  • Robot action sequence generation method and device, equipment and medium

    CN121733579A