A robot imitation learning method, control method, device, equipment and medium

By constructing a combination of prior knowledge information and visual information in robot imitation learning, and using Gaussian hybrid model and multi-head attention mechanism to train the prediction model, the problem of insufficient task success rate and generalization in the existing technology is solved, and the robot's understanding and execution ability of complex tasks is improved.

CN120134328BActive Publication Date: 2025-07-22PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510619026.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-22
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing end-to-end robot imitation learning methods have shortcomings in task success rate and generalization, which affects the robot's understanding of complex tasks and execution accuracy.

Method used

By constructing prior knowledge information based on the action sequence of the preset teaching trajectory set, using the Gaussian mixed model to extract the mean vector, the covariance matrix and the Gaussian distribution weight sequence, and combining the multi-head attention mechanism to build key dynamic features, and training the prediction model to obtain the predictive model learned by imitation.

Benefits of technology

It improves the robot's imitation learning method's robustness to tasks and environmental noise, and improves the robot's understanding of complex tasks and execution accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120134328B_ABST
    Figure CN120134328B_ABST
Patent Text Reader

Abstract

The present application discloses a robot imitation learning method, control method, device, equipment and medium. The method includes constructing prior knowledge information based on the action sequence in the teaching trajectory; determining key dynamic features based on the prior knowledge information and each time step in the teaching trajectory through a prediction model to be trained, and determining a predicted action sequence based on the key dynamic features; training the prediction model to be trained based on the predicted action sequence to obtain a prediction model after imitation learning. The present application first determines prior knowledge information from all actions in the teaching trajectory, and then combines the prior knowledge information with visual information and robot state information to predict future action sequences, so that the utilization efficiency of multi-modal information can be optimized through contrastive learning, and the robustness of the robot imitation learning method to task diversity and environmental noise is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and particularly relates to a robot imitation learning method, a control method, a device, a device and a medium. Background Art

[0002] Imitation learning enables a robot to learn specific technologies by observing a human instructor and is widely used in fields such as service robots and industrial robots. In order to improve the success rate and generalization ability of robot autonomous learning, how to quickly extract key information from demonstration data and maintain efficient learning in an open environment has become a current research hotspot.

[0003] Currently, an end-to-end robot imitation learning method is generally adopted. The end-to-end robot imitation learning method generally implicitly learns the latent representation of the robot task action trajectory and directly conducts end-to-end robot task operation learning. This will result in a low success rate of robot tasks and poor generalization of task operations, affecting the robot's understanding ability and execution accuracy of complex tasks.

[0004] Therefore, the existing technology still needs to be improved. Summary of the Invention

[0005] The technical problem to be solved by the present application is to provide a robot imitation learning method, a control method, a device, a device and a medium in view of the deficiencies of the existing technology.

[0006] To solve the above technical problem, a first aspect of the present application provides a robot imitation learning method. Specifically, the robot imitation learning method includes:

[0007] Construct prior knowledge information based on the action sequences in the demonstration trajectories in a preset demonstration trajectory set. Each demonstration trajectory in the preset demonstration trajectory set includes an action sequence and a time step sequence, and each time step in the time step sequence includes a robot state, visual information, and a reward;

[0008] Based on the prior knowledge information and each time step in the demonstration trajectory, the key dynamic features corresponding to each time step are determined through a prediction model to be trained, and the predicted action sequence corresponding to each time step is determined based on the key dynamic features corresponding to each time step;

[0009] Train the prediction model to be trained based on the predicted action sequence corresponding to each time step to obtain a prediction model after imitation learning.

[0010] In the robot imitation learning method, the demonstration trajectory includes all steps to complete a preset task, and the preset tasks corresponding to the demonstration trajectories in the preset demonstration trajectory set are the same.

[0011] The described robot imitation learning method, wherein constructing prior knowledge information based on the action sequences in the teaching trajectories in the preset teaching trajectory set specifically includes:

[0012] Input the action sequences in the teaching trajectories into a Gaussian mixture model, and extract the corresponding mean vector sequence, covariance matrix sequence, and Gaussian distribution weight sequence of the action sequences through the Gaussian mixture model;

[0013] Take the mean vector sequence, the covariance matrix sequence, and the Gaussian distribution weight sequence as the prior knowledge information of the teaching trajectories.

[0014] The described robot imitation learning method, wherein inputting the action sequences in the teaching trajectories into a Gaussian mixture model and extracting the corresponding mean vector sequence, covariance matrix sequence, and Gaussian distribution weight sequence of the action sequences through the Gaussian mixture model specifically includes:

[0015] Input the action sequences in the teaching trajectories into a Gaussian mixture model, and model the action sequences as a weighted combination of multiple Gaussian distributions through the Gaussian mixture model to obtain the action distribution corresponding to the action sequences;

[0016] Read the Gaussian distribution weights, mean vectors, and variance matrices of each Gaussian distribution in the action distribution;

[0017] Based on all the read Gaussian distribution weights, all the mean vectors, and all the variance matrices, determine the mean vector sequence, covariance matrix sequence, and Gaussian distribution weight sequence.

[0018] The described robot imitation learning method, wherein the sum of the Gaussian distribution weights in the Gaussian distribution weight sequence is equal to 1.

[0019] The described robot imitation learning method, wherein determining the key dynamic features corresponding to each time step based on the prior knowledge information and each time step in the teaching trajectories specifically includes:

[0020] Based on the prior knowledge information and each time step, construct the key vector and value vector corresponding to each time step through a multi-head attention mechanism;

[0021] Based on the preset learnable action sequence, construct the query vector for each time step through a multi-head attention mechanism;

[0022] Based on the key vector, value vector, and query vector, determine the key dynamic features corresponding to each time step through an attention learning mechanism.

[0023] The second aspect of the present application provides a control method for a robot. Specifically, the control method for the robot includes:

[0024] Obtain the current visual information and current robot state information of the robot at the current moment;

[0025] Input the current visual information and current robot state information into a prediction model trained through imitation learning, and output the actions of the robot at the next time step through the prediction model trained through imitation learning;

[0026] Control the robot to execute the actions at the next time step.

[0027] The third aspect of the present application provides a robot imitation learning device. Specifically, the robot imitation learning device includes:

[0028] A prior knowledge construction module, configured to construct prior knowledge information based on the action sequences in the teaching trajectories in a preset teaching trajectory set. Each teaching trajectory in the preset teaching trajectory set includes an action sequence and a time step sequence, and each time step in the time step sequence includes a robot state, visual information, and a reward;

[0029] A control module, configured to determine the key dynamic features corresponding to each time step based on the prior knowledge information and each time step in the teaching trajectory; determine the predicted action sequence corresponding to each time step based on the key dynamic features corresponding to each time step, and train the to-be-trained prediction model based on the predicted action sequence corresponding to each time step to obtain a prediction model trained through imitation learning.

[0030] The fourth aspect of the present application provides a computer-readable storage medium storing one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in any one of the above-mentioned robot imitation learning methods.

[0031] The fifth aspect of the present application provides a terminal device, which includes: a processor and a memory;

[0032] The memory stores a computer-readable program executable by the processor;

[0033] When the processor executes the computer-readable program, it implements the steps in any one of the above-mentioned robot imitation learning methods.

[0034] Beneficial effects: Compared with the prior art, the present application provides a robot imitation learning method, control method, device, equipment and medium. The method includes constructing prior knowledge information based on the action sequences in the teaching trajectories in a preset teaching trajectory set; determining the key dynamic features corresponding to each time step based on the prior knowledge information and each time step in the teaching trajectory through a prediction model to be trained, and determining the predicted action sequence corresponding to each time step based on the key dynamic features corresponding to each time step; training the prediction model to be trained based on the predicted action sequence corresponding to each time step to obtain a prediction model after imitation learning. The present application first determines prior knowledge information from all the actions in the teaching trajectory, and then combines the prior knowledge information with visual information and robot state information to predict future action sequences, so that the utilization efficiency of multi-modal information can be optimized through contrast learning, and the robot imitation learning method has good robustness to the diversity of tasks and environmental noise, enabling the robot imitation learning method to adapt to task scenarios with different initial states, greatly improving the efficiency, accuracy and generalization performance of robot task operation learning, and further improving the robot's understanding ability and execution accuracy for complex tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained without creative efforts based on these drawings.

[0036] Figure 1 It is a flowchart of the robot imitation learning method provided by the embodiment of the present application.

[0037] Figure 2 It is a principle example diagram of the robot imitation learning method provided by the embodiment of the present application.

[0038] Figure 3 It is a principle block diagram of the robot imitation learning device provided by the embodiment of the present application.

[0039] Figure 4 It is a principle block diagram of the terminal device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The embodiments of the present application provide a robot imitation learning method, control method, device, equipment and medium. To make the purpose, technical solutions and effects of the present application clearer and more definite, the following further details the present application with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0041] Those skilled in the art can understand that, unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of this application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0042] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which this application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0043] It should be understood that the sequence numbers and magnitudes of the steps in this embodiment do not mean the order of execution. The order of execution of each process is determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0044] The following further illustrates the application content by describing the embodiments in conjunction with the accompanying drawings.

[0045] This embodiment provides a robot imitation learning method, as Figure 1 shown, the method includes:

[0046] S10. Construct prior knowledge information based on the action sequences in the teaching trajectories in a preset teaching trajectory set.

[0047] Specifically, several teaching trajectories are used to describe all the steps to complete a preset task, and the preset task is a task completed by a robot. That is to say, the robot can complete the preset task according to the teaching trajectories. Among them, the preset tasks corresponding to each teaching trajectory in several teaching trajectories are the same, that is, each teaching trajectory is used to describe all the steps to complete the preset task, and several teaching trajectories can be different trajectories that can complete the preset task.

[0048] Each of a number of teaching trajectories includes an action sequence and a time step sequence. The action sequence includes all actions to complete a preset task, and the time step sequence includes trajectory information before the robot executes each action in the action sequence. The trajectory information includes visual information, robot state, and reward, and the action in the next time step can be predicted through the trajectory information included in this time step. Among them, the robot state is the state of the robot at this time step, which can be read in real time by the robot itself. The visual information is the image information (such as RGB image, etc.) collected through an image acquisition device (such as a camera, etc.) at this time step. The reward is the reward obtained by the robot for executing the task at this time step, and the action is the next action predicted based on the robot state, visual information, and reward.

[0049] For example: Suppose there are N = 50 teaching trajectories , , and each teaching trajectory includes a time step sequence and an action sequence . represents the number of time steps of the teaching trajectory. The th time step in the th teaching trajectory can be expressed as:

[0050] .

[0051] Among them, represents the robot state at the th time step in the th teaching trajectory, represents the visual information at the th time step in the th teaching trajectory, represents the reward at the th time step in the th teaching trajectory.

[0052] Then, imitation learning is to train a prediction model through the teaching trajectories so that the prediction model can learn how to predict the action in the next time step based on the visual information, robot state, and reward at each time step. Specifically, through imitation learning, the prediction model can learn a mapping relationship from the teaching trajectories:

[0053] .

[0054] Among them, represents the prediction model, which is used to receive the trajectory information as input and output an action corresponding to the trajectory information .

[0055] Further, after obtaining a number of teaching trajectories, the prediction model to be trained can be trained based on the number of teaching trajectories to obtain a prediction model after imitation learning. Among them, when training the prediction model to be trained based on a number of teaching trajectories, the number of teaching trajectories can be divided into a number of training batches, each training batch includes some of the teaching trajectories in the number of teaching trajectories, and then each training batch is used to perform one iteration of training on the prediction model to be trained until the prediction end condition, for example, the number of iterations reaches a preset number threshold and / or the loss function meets a preset requirement (such as the loss function value is less than a preset threshold, etc.). For the sake of illustration, here, it is assumed that each training batch includes one teaching trajectory for subsequent description.

[0056] Specifically, when training the prediction model to be trained based on a number of teaching trajectories, the prior knowledge information can be constructed first based on the action sequence in the teaching trajectory, and then the constructed prior knowledge information and the information included in each time step in the teaching trajectory are used as the input information of the prediction model to be trained. Among them, the prior knowledge information can be obtained through a Gaussian mixture model (GMM). Accordingly, the construction of the prior knowledge information based on the action sequence in the teaching trajectory in the preset teaching trajectory set specifically includes:

[0057] Input the action sequence in the teaching trajectory into the Gaussian mixture model, and extract the mean vector sequence, covariance matrix sequence, and Gaussian distribution weight sequence corresponding to the action sequence through the Gaussian mixture model;

[0058] Use the mean vector sequence, the covariance matrix sequence, and the Gaussian distribution weight sequence as the prior knowledge information of the teaching trajectory.

[0059] Specifically, the Gaussian mixture model is used to extract features from the action sequence and use the extracted feature information as prior knowledge information to assist the prediction model to be trained in action prediction. Among them, the Gaussian mixture model can effectively capture the multi-modal features of the action space, which is particularly important for tasks with multiple possible actions or execution methods, and can effectively improve the accuracy of action prediction. For example, in some complex tasks, there may be multiple different actions to complete the prediction target, and the Gaussian mixture model can capture the diversity of complex tasks.

[0060] Furthermore, the Gaussian mixture model is a probability-based model that can transform data into a weighted combination of multiple Gaussian distributions. To this end, after inputting the action sequence into the Gaussian mixture model, the Gaussian mixture model can model the action space corresponding to the action sequence as a Gaussian mixture distribution, and determine the mean vector sequence, covariance matrix sequence, and Gaussian distribution weight sequence based on the Gaussian mixture distribution. Among them, the mean vector sequence includes the mean vectors of each Gaussian distribution in the Gaussian mixture distribution, the covariance matrix sequence includes the covariance matrices of each Gaussian distribution in the Gaussian mixture distribution, and the Gaussian distribution weight sequence includes the weighting coefficients corresponding to each Gaussian distribution in the Gaussian mixture distribution.

[0061] Exemplarily, the step of inputting the action sequence in the teaching trajectory into the Gaussian mixture model and extracting the mean vector sequence, covariance matrix sequence, and Gaussian distribution weight sequence corresponding to the action sequence through the Gaussian mixture model specifically includes:

[0062] Input the action sequence in the teaching trajectory into the Gaussian mixture model, and model the action sequence as a weighted combination of multiple Gaussian distributions through the Gaussian mixture model to obtain the action distribution corresponding to the action sequence;

[0063] Read the Gaussian distribution weights, mean vectors, and variance matrices of each Gaussian distribution in the action distribution;

[0064] Based on all the read Gaussian distribution weights, all mean vectors, and all covariance matrices, determine the mean vector sequence, covariance matrix sequence, and Gaussian distribution weight sequence.

[0065] Specifically, the Gaussian mixture model models the action distribution of the action sequence as a weighted combination of multiple Gaussian distributions. For example, the action sequence as described above , represents the -th step of the -th trajectory, represents the dimension of the action, and the action sequence can be expressed as , represents the number of time steps included in the teaching trajectory. Then, the Gaussian mixture model distributes the action sequence as a weighted combination of multiple Gaussian distributions, that is:

[0066] ,

[0067] Among them, represents the action space of the action sequence; represents the number of Gaussian distributions, usually determined by methods such as cross-validation; represents the Gaussian distribution weight of the -th Gaussian distribution, and satisfies , represents the mean vector of the -th Gaussian distribution, and represents the covariance matrix of the

[0068] Furthermore, after reading the Gaussian distribution weights, mean vectors, and variance matrices of each Gaussian distribution in the action distribution, a mean vector sequence, a covariance matrix sequence, and a Gaussian distribution weight sequence can be determined based on all the read Gaussian distribution weights, mean vectors, and variance matrices. Among them, the mean vector sequence, the covariance matrix sequence, and the Gaussian distribution weight sequence can be respectively expressed as:

[0069] ,

[0070] ,

[0071] ,

[0072] where represents the Gaussian distribution weight sequence, represents the mean vector sequence, represents the covariance matrix sequence.

[0073] By using a Gaussian mixture model to extract features from the action sequence in this application, prior knowledge information for describing the task can be effectively extracted without additionally increasing the number of model parameters of the prediction model, and the prior knowledge information is directly used to replace the latent representation learned separately by a single network in explicit learning, thereby reducing the network complexity and model parameters of the prediction model. At the same time, through the prior knowledge information, the prediction model can obtain multi-modal information of the action sequence, improving the robot's understanding ability and execution accuracy for complex tasks.

[0074] S20. Based on the prior knowledge information and each time step in the demonstration trajectory, the prediction model to be trained determines the key dynamic features corresponding to each time step, and determines the predicted action sequence corresponding to each time step based on the key dynamic features corresponding to each time step.

[0075] Specifically, after obtaining the prior knowledge information, the prior knowledge information and the trajectory information in each time step are used as input items of the prediction model to be trained, and the prediction model to be trained combines and learns the prior knowledge information, visual information, and robot state to output the predicted action sequence. For example, in the above example, the input data of the prediction model to be trained can be expressed as:

[0076] ,

[0077] where Represents the input data of the prediction model to be trained.

[0078] Furthermore, the prediction model to be trained is a deep learning model for outputting a predicted action sequence based on the input data In an embodiment of the present application, the prediction model to be trained adopts a Transformer model. Specifically, as Figure 2 shown, the prior knowledge information and the trajectory information included in each time step are input into the encoder in the Transformer model, and the preset learnable action sequence in the Transformer model is input into the decoder in the Transformer model. Then, the input embedding representation is output through the encoder, and the learnable embedding representation is output through the decoder. Then, the key dynamic features are determined based on the input embedding representation and the learnable embedding representation.

[0079] Exemplarily, the determining of the key dynamic feature corresponding to each time step based on the prior knowledge information and each time step in the demonstration trajectory specifically includes:

[0080] Based on the prior knowledge information and each time step, construct the key vector and value vector corresponding to each time step through the multi-head attention mechanism;

[0081] Based on the preset learnable action sequence, construct the query vector for each time step through the multi-head attention mechanism;

[0082] Based on the key vector, value vector, and query vector, determine the key dynamic feature corresponding to each time step through the attention learning mechanism.

[0083] Specifically, as Figure 2As shown in the figure, first, the prior knowledge information and the trajectory information in the time step are input into the encoder. The content embedding vector is generated by the encoder. Then, the content embedding vector is concatenated with the position embedding vectors of the prior knowledge information and the trajectory information in the time step to obtain the input embedding representation. Based on this input embedding representation, the key vector and the value vector are constructed through the multi-head attention mechanism. Specifically, the encoder can connect the multi-head attention module and the feed-forward network module. The input embedding representation is input into the multi-head attention module, and the intermediate feature embedding is input through the multi-head attention module. Then, the intermediate feature embedding is input into the feed-forward network module, and the feature representation is input through the feed-forward network module. The key vector and the value vector are determined based on this feature representation. For example, this feature representation is used as the key vector and the value vector. On the other hand, the preset learnable action sequence is encoded into the action content embedding vector by the decoder, and the action content embedding vector is concatenated with the action position embedding vector corresponding to the preset learnable action sequence to obtain the action learnable embedding representation. Then, the learnable embedding representation is input into the multi-head attention module, and the action feature representation is input through the multi-head attention module. The query vector is constructed based on the action feature representation. For example, the action feature representation is used as the query vector, etc. Finally, after obtaining the key vector, the value vector, and the query vector, the key vector, the value vector, and the query vector are input into the multi-head attention module, and the key dynamic features are output through the multi-head attention module.

[0084] Further, after obtaining the key dynamic features, the key dynamic features can be input into the prediction module, and the prediction action sequence is output by the prediction module based on the key dynamic features. Among them, the prediction module can adopt a linear layer or a fully connected layer, etc.

[0085] S30. Train the to-be-trained prediction model based on the prediction action sequence corresponding to each time step to obtain the prediction model after imitation learning.

[0086] Specifically, after obtaining the prediction action sequence, the to-be-trained prediction model can be trained based on the prediction action sequence. Among them, when training the to-be-trained prediction model based on the prediction action sequence, a contrastive learning loss function can be used to determine the loss term, or other loss functions can be used to determine the loss term. For example, cross learning, etc. There is no specific limitation here. In addition, the existing training process can also be adopted for the training process, and it will not be specifically described here.

[0087] In summary, this embodiment provides a robot imitation learning method, which includes constructing prior knowledge information based on the action sequences in the teaching trajectories in a preset teaching trajectory set; determining the key dynamic features corresponding to each time step based on the prior knowledge information and each time step in the teaching trajectory through a prediction model to be trained, and determining the predicted action sequence corresponding to each time step based on the key dynamic features corresponding to each time step; training the prediction model to be trained based on the predicted action sequence corresponding to each time step to obtain a prediction model after imitation learning. In this application, prior knowledge information is first determined from all the actions in the teaching trajectory, and then the prior knowledge information is combined with visual information and robot state information to predict the future action sequence. In this way, the utilization efficiency of multi-modal information can be optimized through contrast learning, improving the robustness of the robot imitation learning method to task diversity and environmental noise, enabling the robot imitation learning method to adapt to task scenarios with different initial states, greatly improving the efficiency, accuracy, and generalization performance of robot task operation learning, and further enhancing the robot's understanding ability and execution accuracy for complex tasks.

[0088] Based on the above robot imitation learning method, this embodiment provides a robot control method, which specifically includes:

[0089] Obtain the current visual information and current robot state information of the robot at the current moment;

[0090] Input the current visual information and current robot state information into the prediction model after imitation learning, and output the action of the next time step of the robot through the prediction model after imitation learning;

[0091] Control the robot to execute the action of the next time step.

[0092] Based on the above robot imitation learning method, this embodiment provides a robot imitation learning device, as Figure 3 shown, the robot imitation learning device specifically includes:

[0093] A prior knowledge construction module 100, configured to construct prior knowledge information based on the action sequences in the teaching trajectories in a preset teaching trajectory set, where each teaching trajectory in the preset teaching trajectory set includes an action sequence and a time step sequence, and each time step in the time step sequence includes a robot state, visual information, and a reward;

[0094] The control module 200 is configured to determine the key dynamic features corresponding to each time step based on the prior knowledge information and each time step in the teaching trajectory; determine the predicted action sequence corresponding to each time step through a to-be-trained prediction model based on the key dynamic features corresponding to each time step, and train the to-be-trained prediction model based on the predicted action sequence corresponding to each time step to obtain a prediction model after imitation learning.

[0095] Based on the above robot imitation learning method, this embodiment provides a computer-readable storage medium storing one or more programs, which can be executed by one or more processors to implement the steps in the robot imitation learning method as described in the above embodiment.

[0096] Based on the above robot imitation learning method, this application also provides a terminal device, as Figure 4 shown, which includes at least one processor 20; a display screen 21; and a memory 22, and may further include a communication interface 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22, and the communication interface 23 can complete mutual communication through the bus 24. The display screen 21 is set to display a user guidance interface preset in the initial setting mode. The communication interface 23 can transmit information. The processor 20 can call the logical instructions in the memory 22 to execute the method in the above embodiment.

[0097] In addition, when the logical instructions in the above memory 22 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium.

[0098] The memory 22, as a computer-readable storage medium, can be set to store software programs and computer-executable programs, such as the program instructions or modules corresponding to the method in the embodiment of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, that is, implements the method in the above embodiment.

[0099] The memory 22 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, may also be a transient storage medium.

[0100] In addition, the specific processes of loading and executing multiple instructions in the above-mentioned storage medium and the terminal device have been described in detail in the above method, and will not be repeated here one by one.

[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A robot imitation learning method, characterized in that The described robot imitation learning method specifically includes: Constructing prior knowledge information based on the action sequences in the teaching trajectories in a preset teaching trajectory set, where each teaching trajectory in the preset teaching trajectory set includes an action sequence and a time step sequence, and each time step in the time step sequence includes a robot state, visual information, and a reward; Determining the key dynamic features corresponding to each time step based on the prior knowledge information and each time step in the teaching trajectory, and determining the predicted action sequence corresponding to each time step based on the key dynamic features corresponding to each time step; Training a prediction model to be trained based on the predicted action sequence corresponding to each time step to obtain a prediction model after imitation learning.

2. The robot imitation learning method according to claim 1, wherein, The teaching trajectory includes all steps to complete a preset task, and the preset tasks corresponding to the teaching trajectories in the preset teaching trajectory set are the same.

3. The robot imitation learning method according to claim 1, wherein, The constructing prior knowledge information based on the action sequences in the teaching trajectories in a preset teaching trajectory set specifically includes: Inputting the action sequence in the teaching trajectory into a Gaussian mixture model, and extracting the mean vector sequence, covariance matrix sequence, and Gaussian distribution weight sequence corresponding to the action sequence through the Gaussian mixture model; Taking the mean vector sequence, the covariance matrix sequence, and the Gaussian distribution weight sequence as the prior knowledge information of the teaching trajectory.

4. The robot imitation learning method according to claim 3, wherein The inputting the action sequence in the teaching trajectory into a Gaussian mixture model, and extracting the mean vector sequence, covariance matrix sequence, and Gaussian distribution weight sequence corresponding to the action sequence through the Gaussian mixture model specifically includes: Inputting the action sequence in the teaching trajectory into a Gaussian mixture model, and modeling the action sequence as a weighted combination of multiple Gaussian distributions through the Gaussian mixture model to obtain the action distribution corresponding to the action sequence; Reading the Gaussian distribution weight, mean vector, and variance matrix of each Gaussian distribution in the action distribution; Determining the mean vector sequence, covariance matrix sequence, and Gaussian distribution weight sequence based on all the read Gaussian distribution weights, all the mean vectors, and all the variance matrices.

5. The robot imitation learning method according to claim 3 or 4, characterized in that The sum of the Gaussian distribution weights in the Gaussian distribution weight sequence is equal to 1.

6. The robot imitation learning method according to claim 1, wherein The determining the key dynamic features corresponding to each time step based on the prior knowledge information and each time step in the teaching trajectory specifically includes: Constructing the key vector and value vector corresponding to each time step through a multi-head attention mechanism based on the prior knowledge information and each time step; Constructing the query vector corresponding to each time step through a multi-head attention mechanism based on a preset learnable action sequence; Determining the key dynamic features corresponding to each time step through an attention learning mechanism based on the key vector, value vector, and query vector.

7. A control method for a robot, characterized in that, Applying the prediction model after imitation learning obtained by using the robot imitation learning method according to any one of claims 1-6, the control method of the described robot specifically includes: Obtaining the current visual information and current robot state information of the robot at the current moment; Input the current visual information and the current robot state information into the prediction model trained by imitation learning, and output the actions of the robot at the next time step through the prediction model trained by imitation learning; Control the robot to execute the actions at the next time step.

8. A robot imitation learning device, characterized in that, The robot imitation learning device specifically includes: A prior knowledge construction module, configured to construct prior knowledge information based on the action sequences in the teaching trajectories in a preset teaching trajectory set, where each teaching trajectory in the preset teaching trajectory set includes an action sequence and a time step sequence, and each time step in the time step sequence includes a robot state, visual information, and a reward; A control module, configured to determine the key dynamic features corresponding to each time step based on the prior knowledge information and each time step in the teaching trajectory; determine the predicted action sequence corresponding to each time step based on the key dynamic features corresponding to each time step, and train the prediction model to be trained based on the predicted action sequence corresponding to each time step to obtain a prediction model trained by imitation learning.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the robot imitation learning method according to any one of claims 1-6.

10. A terminal device, characterized in that, Including: A processor and a memory; A computer-readable program executable by the processor is stored on the memory; When the processor executes the computer-readable program, the steps in the robot imitation learning method according to any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Trajectory prediction method for cross-modal semantic information supervision

    CN117009787A

  • Man-machine cooperation blanket trajectory imitation learning method based on coring motion primitives

    CN117359626A