Robot Skill Continuous Learning Method and Device Based on Imitation Learning

By introducing imitation learning and skill codebooks into the robot policy network, the problems of inflexible skill selection and incoherent execution in the existing technology are solved, and the logic, smooth and coherent actions of the robot in long-sequence tasks are realized, and the task success rate and anti-forgetfulness are improved.

CN119589681BActive Publication Date: 2025-06-24INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411975691.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-06-24
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

Existing robot policy networks have challenges in issues such as inflexible skill selection, incoherent execution, and the difficulty in finding abstract hierarchies containing rich meaningful sub-behaviors.

Method used

The robot skills continuous learning method based on imitation learning is adopted. Through the imitation learning strategy network combined with the skill codebook, observe images, status information and text instructions are obtained, action information is predicted, and network parameters are updated based on loss to achieve continuous learning and reuse of skills.

Benefits of technology

The logic, smoothness and coherence of the robot's actions when executing long sequence tasks is realized, the forward migration ability and average task success rate are improved, and the complexity of catastrophic forgetting and skill selection is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119589681B_ABST
    Figure CN119589681B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and apparatus for continuous learning of robot skills based on imitation learning. The method includes: for each of a plurality of tasks in sequence, before a preset termination condition is satisfied, the following operations are repeatedly executed in a loop: inputting an observation image, state information, a text instruction of the current task obtained in advance, and a skill codebook of the current task into an imitation learning policy network to predict the action information of the robot at the current moment, where the skill codebook of the current task is a plurality of learnable sub-skill vectors initialized in advance and includes skill information of the skill codebook of the previous task in the case of having a previous task; calculating a loss based on the predicted action information of the robot and the true action information of the robot in the expert demonstration data; and updating the imitation learning policy network and the skill codebook of the current task based on the loss. The problems of inflexible skill selection and difficulty in abstracting the sub-skill hierarchy in manual skill design are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of information science and robotics, and particularly to a method and apparatus for continuous learning of robot skills based on imitation learning. Background Art

[0002] Robots can perceive their open and dynamic environment, complete corresponding tasks according to human instructions, and interact with humans, with characteristics such as safety, flexible deployment, and simple operation. In recent years, they have been widely used in fields such as industrial manufacturing, medical rehabilitation, and home services. In the real world, many tasks we hope robots to solve may be different from the tasks the robots have completed in the past, but the skills they require are similar. Humans continuously accumulate experience through extensive skill learning and then quickly adapt to the environment and solve new tasks. Therefore, continuous learning of robot skills is of great significance for the future development of robot technology and is an important basis for the wide application of future robots in various fields.

[0003] In the robot policy network in the related art, there are problems such as inflexible skill selection, inconsistent execution, and difficulty in finding an abstract hierarchical structure of sub-behaviors with rich meanings in artificially designed skills. Summary of the Invention

[0004] The method and apparatus for continuous learning of robot skills based on imitation learning provided by the exemplary embodiments of the present disclosure can at least solve the above technical problems and other technical problems not mentioned above.

[0005] According to one aspect of the present disclosure, there is provided a method for continuous learning of robot skills based on imitation learning, the method including: for each of a plurality of tasks in sequence, before a preset termination condition is satisfied, the following operations are repeatedly executed: obtaining an observation image and state information of the robot at the current moment; inputting the observation image, the state information, a text instruction of the current task obtained in advance, and a skill codebook of the current task into an imitation learning policy network to predict action information of the robot at the current moment, where the skill codebook of the current task is a plurality of learnable sub-skill vectors initialized in advance and includes skill information of the skill codebook of a previous task in the case of having a previous task; calculating a loss based on the predicted action information of the robot and the true action information of the robot in the expert demonstration data; and updating the imitation learning policy network and the skill codebook of the current task based on the loss to perform the continuous learning process of the robot skills.

[0006] Optionally, the imitation learning policy network includes a perception layer, a skill inference layer, and an action execution layer; wherein, inputting the observation image, the state information, the text instruction of the current task obtained in advance, and the skill codebook into the imitation learning policy network to predict the action information of the robot at the current moment includes: inputting the observation image, the state information, and the text instruction into the perception layer to obtain the multi-modal fusion information at the current moment; inputting the multi-modal fusion information and the skill codebook into the skill inference layer to predict the skill vector at the current moment; and inputting the skill vector into the action execution layer to obtain the action information of the robot at the current moment.

[0007] Optionally, the perception layer includes a pre-trained text encoder, a learnable multi-layer perceptron, a pre-trained learnable visual encoder, a learnable state encoder, and a modality fusion module; wherein, inputting the observation image, the state information, and the text instruction into the perception layer to obtain the multi-modal fusion information at the current moment includes: inputting the text instruction into the text encoder to obtain a text vector; inputting the text vector into the multi-layer perceptron to obtain a text representation; inputting the observation image into the visual encoder to obtain an image representation; inputting the state information into the state encoder to obtain a state representation; after fusing the text vector and the image representation through the modality fusion module, splicing the fused features with the text representation and the state representation to obtain the multi-modal fusion information.

[0008] Optionally, the skill inference layer includes an attention module and a high-level decoder, and the high-level decoder includes a multi-head self-attention layer, a multi-head cross-attention layer, and a feed-forward network layer; wherein, inputting the multi-modal fusion information and the skill codebook into the skill inference layer to predict the skill vector at the current moment includes: inputting the multi-modal fusion information into the attention module to obtain weights for multiple learnable sub-skill vectors in the skill codebook; weighting the multiple learnable sub-skill vectors based on the weights to obtain multiple learnable weighted sub-skill vectors; inputting the multi-modal fusion information into the multi-head self-attention layer to obtain observation information containing temporal information; inputting the feature obtained by splicing the weighted sub-skill vectors and the observation information, and the observation information into the multi-head cross-attention layer, and inputting the output of the multi-head cross-attention layer into the feed-forward network layer to predict the skill vector.

[0009] Optionally, inputting the feature obtained by concatenating the weighted sub-skill vector and the observation information, and the observation information into the multi-head cross-attention layer includes: using the feature obtained by concatenating the weighted sub-skill vector and the observation information as the key and value of the multi-head cross-attention layer, and using the observation information as the query of the multi-head cross-attention layer to perform multi-head cross-attention.

[0010] Optionally, the action execution layer includes a bottom decoder and a Gaussian mixture model distribution layer, and the bottom decoder includes a multi-head self-attention layer and a feed-forward network layer; wherein, inputting the skill vector into the action execution layer to obtain the action information of the robot at the current moment includes: processing the skill vector through the multi-head self-attention layer, the feed-forward network layer, and the Gaussian mixture model distribution layer to obtain the action information of the robot at the current moment.

[0011] Optionally, for the calculation of each head in the multi-head self-attention layer and the multi-head cross-attention layer in the skill inference layer, and the multi-head self-attention layer in the action execution layer, the following operations are used to approximate the pattern of the current head: mapping the input of the current head through three projection matrices into a query, a key, and a value; stacking the weight matrices of the current head to obtain a first tensor; using CP decomposition to decompose the second tensor into the sum of at least one rank-one component, where the second tensor is a learnable tensor for the current task pre-created using pattern approximation parameters, and the component is obtained based on multiple shared components among multiple tasks, specific components of the current task, and a learnable coefficient vector pre-randomly initialized for the current task; performing the calculation of the current head based on the query, the key, the value, the first tensor, and the second tensor after CP decomposition.

[0012] Optionally, the preset termination condition is to test the success rate of the imitation learning policy network on the current task at the preset training times checkpoint until the preset termination training times are reached, or after the success rate is higher than the first preset ratio, the success rates obtained by continuous preset times of checkpoint tests are lower than the first preset ratio; in the case of satisfying the preset termination condition, terminate the training of the current task and save the model parameters of the imitation learning policy network in the case of the highest success rate.

[0013] Optionally, the method further includes: in the case of satisfying the preset termination condition, freezing the parameters of the most important second preset ratio in the convolutional layer and the linear layer in the imitation learning policy network as the specific parameters of the current task.

[0014] According to another aspect of the present disclosure, there is also provided a robot skill continuous learning device based on imitation learning, characterized in that the device includes: a task switching module configured to: sequentially for each of a plurality of tasks, loop through the operations of an information acquisition module, an action prediction module, a loss calculation module, and a parameter update module before a preset termination condition is met; the information acquisition module is configured to: acquire the observation image and status information of the robot at the current moment; the action prediction module is configured to: input the observation image, the status information, the text instruction of the current task pre-acquired, and the skill codebook of the current task into an imitation learning policy network to predict the action information of the robot at the current moment, wherein the skill codebook of the current task is a plurality of learnable sub-skill vectors initialized in advance, and in the case of having a previous task, includes the skill information of the skill codebook of the previous task; the loss calculation module is configured to: calculate a loss based on the predicted action information of the robot and the true action information of the robot in the expert demonstration data; the parameter update module is configured to: update the imitation learning policy network and the skill codebook of the current task based on the loss to perform the robot skill continuous learning process.

[0015] Optionally, the imitation learning policy network includes a perception layer, a skill inference layer, and an action execution layer; wherein, the action prediction module is configured to: input the observation image, the status information, and the text instruction into the perception layer to obtain the multi-modal fusion information at the current moment; input the multi-modal fusion information and the skill codebook into the skill inference layer to predict the skill vector at the current moment; input the skill vector into the action execution layer to obtain the action information of the robot at the current moment.

[0016] Optionally, the perception layer includes a pre-trained text encoder, a learnable multi-layer perceptron, a pre-trained learnable visual encoder, a learnable status encoder, and a modality fusion module; wherein, the action prediction module is configured to: input the text instruction into the text encoder to obtain a text vector; input the text vector into the multi-layer perceptron to obtain a text representation; input the observation image into the visual encoder to obtain an image representation; input the status information into the status encoder to obtain a status representation; after fusing the text vector and the image representation through the modality fusion module, splice the fused feature with the text representation and the status representation to obtain the multi-modal fusion information.

[0017] Optionally, the skill inference layer includes an attention module and a high-level decoder, and the high-level decoder includes a multi-head self-attention layer, a multi-head cross-attention layer, and a feed-forward network layer; wherein, the action prediction module is configured to: input the multi-modal fusion information into the attention module to obtain weights for multiple learnable sub-skill vectors in the skill codebook; weight the multiple learnable sub-skill vectors based on the weights to obtain multiple learnable weighted sub-skill vectors; input the multi-modal fusion information into the multi-head self-attention layer to obtain observation information containing temporal information; input the features obtained by concatenating the weighted sub-skill vectors and the observation information, and the observation information into the multi-head cross-attention layer, and input the output of the multi-head cross-attention layer into the feed-forward network layer to predict the skill vector.

[0018] Optionally, the action prediction module is configured to: use the features obtained by concatenating the weighted sub-skill vectors and the observation information as the key and value of the multi-head cross-attention layer, and use the observation information as the query of the multi-head cross-attention layer to perform multi-head cross-attention.

[0019] Optionally, the action execution layer includes a low-level decoder and a Gaussian mixture model distribution layer, and the low-level decoder includes a multi-head self-attention layer and a feed-forward network layer; wherein, the action prediction module is configured to: process the skill vector through the multi-head self-attention layer, the feed-forward network layer, and the Gaussian mixture model distribution layer to obtain the action information of the robot at the current moment.

[0020] Optionally, for the calculation of each head in the multi-head self-attention layer and the multi-head cross-attention layer in the skill inference layer, and the multi-head self-attention layer in the action execution layer, the mode approximation of the current head is implemented through the operation of a mode approximation module; wherein, the mode approximation module is configured to: map the input of the current head into a query, a key, and a value through three projection matrices; stack the weight matrix of the current head to obtain a first tensor; use CP decomposition to decompose a second tensor into the sum of at least one rank-one component, where the second tensor is a learnable tensor for the current task pre-created using mode approximation parameters, and the component is obtained based on multiple shared components among multiple tasks, specific components of the current task, and a learnable coefficient vector pre-randomly initialized for the current task; perform the calculation of the current head based on the query, the key, the value, the first tensor, and the second tensor after CP decomposition.

[0021] Optionally, the preset termination condition is to test the success rate of the imitation learning policy network on the current task at a preset training times checkpoint until the preset termination training times are reached, or after the success rate is higher than the first preset ratio, the success rates obtained from consecutive preset times checkpoints in the subsequent tests are lower than the first preset ratio; when the preset termination condition is met, terminate the training of the current task and save the model parameters of the imitation learning policy network in the case of the highest success rate.

[0022] Optionally, the device further includes: a parameter isolation module, configured to: when the preset termination condition is met, freeze the parameters of the most important second preset ratio in the convolutional layer and the linear layer of the imitation learning policy network as the specific parameters of the current task.

[0023] According to another aspect of the embodiments of the present disclosure, there is also provided an electronic device, including: at least one processor; at least one memory storing computer-executable instructions, wherein when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the continuous learning method for robot skills based on imitation learning as described in any one of the above.

[0024] According to another aspect of the embodiments of the present disclosure, there is also provided a computer-readable storage medium storing instructions, which when run by at least one processor, cause the at least one processor to execute the continuous learning method for robot skills based on imitation learning as described in any one of the above.

[0025] According to another aspect of the embodiments of the present disclosure, there is also provided a system including at least one computing device and at least one storage device storing instructions, wherein when the instructions are run by the at least one computing device, the at least one computing device is caused to execute the continuous learning method for robot skills based on imitation learning as described in any one of the above.

[0026] According to another aspect of the embodiments of the present disclosure, there is also provided a computer program product, including computer programs / instructions, which when executed by a processor, implement the continuous learning method for robot skills based on imitation learning as described in any one of the above.

[0027] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0028] A method and device for continuous learning of robot skills based on imitation learning according to the present disclosure. By introducing a skill codebook, a sub-skill library is created separately for each task. Through network training, the sub-skills utilized in each task are captured. When executing the current task, the sub-skill vector can also be selected from the skill codebook of the previous task, realizing the reuse of sub-skills, and enabling the actions generated by the robot when executing long-sequence (more complex or containing multiple goals) tasks to be more logical, smoother, and coherent.

[0029] In addition, by introducing a hierarchical architecture, the robot can adjust decisions online according to real-time perception information when executing long-sequence (more complex or containing multiple goals) tasks, learn a reasonable skill abstraction level, and has better forward transfer ability and higher average task success rate compared with other policy architectures.

[0030] In addition, after learning each task, pruning the main network, fine-tuning to freeze the parameters of the previous task, and separately creating task-specific parameters (sub-skill library and task-specific decomposition vector) and task-shared parameters for each task endow the robot with the ability to discover the same atomic skills hidden in different tasks, and have better backward transfer ability and anti-forgetting compared with existing continuous learning algorithms.

[0031] In addition, on the long-view dataset, it can effectively capture the similarities between tasks (such as information about skills used or environmental layout, etc.), transfer the knowledge obtained from learning new tasks, and improve the success rate of completing previous tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation of the present disclosure.

[0033] Figure 1 Showing the environmental settings of each task in the target dataset and the initial state of the robot's robotic arm as presented in the exemplary embodiments of the present disclosure;

[0034] Figure 2 Showing the flowchart of the method for continuous learning of robot skills based on imitation learning in the exemplary embodiments of the present disclosure;

[0035] Figure 3 Showing the schematic flow diagram of data processing through the imitation learning policy network in the exemplary embodiments of the present disclosure;

[0036] Figure 4 Showing the success rate of each task of the three policy network architectures on the target dataset during the continuous learning process in the exemplary embodiments of the present disclosure;

[0037] Figure 5Shows the success rate of each task of three continuous learning algorithms in the continuous learning process on the target dataset in the exemplary embodiments of the present disclosure;

[0038] Figure 6 Shows a block diagram of a robot skill continuous learning device based on imitation learning in the exemplary embodiments of the present disclosure;

[0039] Figure 7 Shows a block diagram of an electronic device in the exemplary embodiments of the present disclosure. Detailed implementation manners

[0040] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0041] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0042] It should be noted here that "at least one of several items" in the present disclosure all represents three parallel situations including "any one of the several items", "any combination of several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example, "performing at least one of step one and step two" means the following three parallel situations: (1) performing step one; (2) performing step two; (3) performing step one and step two.

[0043] First, introduce the abbreviations, English, and definitions of key terms involved in the present disclosure:

[0044] Catastrophic forgetting: After learning new knowledge, almost completely forget the content mastered before. It makes the artificial intelligence agent lack the ability to continuously adapt to the environment and incrementally (continuously) learn like a biological being.

[0045] Imitation Learning (IL): Train the machine to be able to replicate the continuous actions of humans, and then endow it with the ability to imitate humans to complete task goals.

[0046] Multi-task Learning: It is a form of joint learning where multiple tasks are learned in parallel and the results influence each other.

[0047] Continual Learning: It has the ability to continuously learn like humans, using the experience and knowledge learned from historical tasks to assist in learning newly emerging tasks, and this experience and knowledge accumulate continuously without forgetting old knowledge due to new tasks.

[0048] Experience Replay: A continual learning method that maintains a replay buffer storing training data from previous tasks. During the training of a certain task, a portion of the replay data is selected to be trained together with the current task data to achieve the effect of anti-forgetting.

[0049] Transformer: It is a deep learning model architecture used for natural language processing (NLP) and other sequence-to-sequence tasks.

[0050] PackNet: An existing method for adding multiple tasks in a single deep neural network while avoiding catastrophic forgetting.

[0051] Forward Transfer (FWT) is used to evaluate the ability of an agent to utilize prior knowledge to facilitate the learning of new tasks.

[0052] Backward Transfer (BT) is used to evaluate the ability of an agent to use newly acquired knowledge to improve the performance of the model on previous tasks.

[0053] Mask: It is used to select network parameters, and its value represents the task subscript to which the parameter belongs.

[0054] FiLM: An existing means capable of fusing multi-modal information, which is used in the present invention to fuse text information and picture information.

[0055] CODA-Prompt: An existing method for large model prompt learning.

[0056] Next, introduce the related technologies and their existing problems:

[0057] Imitation learning has made great progress in efficiently teaching robots to perform common manipulation tasks, especially through supervised learning using human teleoperation demonstrations or expert demonstration trajectories. Despite this promise, imitation learning methods have been limited to single short sequence tasks, such as opening a door or picking up a specific object, and have not performed well on multiple long sequence tasks.

[0058] When faced with long sequence tasks, traditional policy architectures process perception planning and action execution separately, which not only increases the complexity and latency of the system, but also makes it difficult to adjust decisions online based on real-time perception information, and the connection between skills is not smooth enough. By introducing skills, the robot can operate in the upper skill space instead of the lower action space, thereby achieving accurate reasoning over a longer time frame. Skills can be learned from demonstrations, trial-and-error learning, or manual definition, and these works show that a single Transformer policy can learn multiple skills. However, when sorting skills, connection problems may occur because the skills cannot transition smoothly. Some works solve this problem by learning to match the start and end sets of skills or treating the connection itself as a new skill, but this makes the skill library composition more complex and difficult to maintain. Decision Diffuser learns a generative model to combine skills, but the skill combination order needs to be specified in advance, which is not flexible enough.

[0059] In addition, the robot's state and action space is complex, and the behavior is abstract. Therefore, some hierarchical architectures that use skill classification before executing the skill often have difficulty determining the appropriate hierarchical level. That is, the artificially defined skills may be too general, resulting in the same skill corresponding to many different robot arm trajectories, and over-reliance on the subsequent action network; or they may be too specific, resulting in too small gaps between skill categories, and over-reliance on the accuracy of the previous skill network. Both of these situations will affect the robot's ability to perform tasks. Artificially defined sub-behaviors, such as option theory used in related technologies, hinder the exploration capabilities required for autonomous operation. In addition, referring to the skills corresponding to the tasks in the dataset, artificially explicitly defining and constraining the skill space will greatly reduce the flexibility of skill selection and the robot's ability to continuously learn new skills to complete new tasks, because the number of skills defined in this way is fixed and the types of skills cannot be continuously expanded as the robot continues to learn.

[0060] As an agent interacting with the world, a robot should be able to learn and execute different tasks in various application scenarios. The traditional approach often adopts the multi-task learning paradigm, training the model by learning all the data at once. Facing incremental learning goals, to avoid catastrophic forgetting, it is necessary to continuously retrain, which will consume a large amount of time and resources. In addition, when such agents as robots are deployed in the real environment, they do not know all specific tasks, and the storage and computing resources they can carry are extremely limited, so careful planning is required to avoid overflow. Therefore, multi-task learning cannot meet the needs of robot skill learning. Continuous learning enables robots to quickly adapt and learn in different environments and tasks without the need for retraining or redesign, only requiring incremental updates to the existing model, which greatly improves the task execution efficiency and endows the robot with the ability to continuously adapt to the open dynamic environment and different task goals.

[0061] The continuous learning of traditional robot policy networks often adopts the method of experience replay, but in actual deployment, the robot agent may not have enough storage space to save the data of previous tasks. There are also some works that collect large datasets of robots completing different tasks, store different skills (actions) taken by the robots when completing tasks as policies, and use large language models or large models pre-trained on large-scale robot task datasets to decompose tasks into these skills and obtain the final control output. When a new task needs to be learned, the skill library is expanded through fine-tuning. Although the powerful performance of large language models endows robots with strong learning abilities, large language models also suffer from catastrophic forgetting and will produce "hallucinations", that is, they do not fully consider conditions such as the real environment where the robot is located, the target requirements, and the execution ability of the robot itself, and output unrealistic task execution plans. In addition, collecting a large number of robot skill datasets and pre-training them not only consumes a large amount of manpower and material resources, but also requires an unimaginable time cost.

[0062] Generally speaking, the existing robot policy networks cannot adjust decisions online according to real-time perception information, it is difficult to determine a reasonable skill abstraction level, and the connection between skills is not smooth; most of them adopt simple continuous learning methods, resulting in low anti-forgetting ability and deployment feasibility.

[0063] To solve the above problems, the present disclosure provides a method and device for continuous learning of robot skills based on imitation learning. By introducing a skill codebook, a sub-skill library is created separately for each task. Through network training, the sub-skills utilized by each task are captured. When executing the current task, the sub-skill vector can also be selected from the skill codebook of the previous task, realizing the reuse of sub-skills, making the actions generated by the robot when executing long-sequence (more complex or containing multiple goals) tasks more logical, smoother and more coherent.

[0064] First, introduce the training dataset used in the exemplary embodiments of the present disclosure.

[0065] According to the exemplary embodiments of the present disclosure, training can be performed using a dataset that is used for decision-making continuous learning in the field of robot operation and includes multiple tasks.

[0066] Specifically, the dataset can include a series of rich and diverse tasks that can reflect human daily activities, such as turning on the stove, moving a book, opening a drawer, etc. Each task can have a corresponding language instruction, such as "Open the drawer at the top of the cabinet and put the bowl in it."

[0067] The dataset can be differentiated to obtain a spatial dataset, an object dataset, a target dataset, and a 100-dataset. Each of the first three task sets can contain ten tasks for studying the transfer ability of the model's spatial information, item information, and task target information knowledge.

[0068] Specifically, all tasks in the spatial dataset can require the robot to place a bowl on a plate among the same group of items. Although the shapes, patterns, etc. of the two bowls are the same, their positions or spatial relationships with other objects are different. Therefore, in order to successfully complete the tasks in the spatial dataset, the robot needs to continuously learn and remember new spatial relationships.

[0069] All tasks in the item dataset can require the robot to pick up and place a unique item. Therefore, in order to complete the tasks in the item dataset, the robot needs to continuously learn and remember new item types.

[0070] All tasks in the target dataset can have the same operating object with a fixed spatial relationship and can only differ in task objectives. Therefore, in order to complete the tasks in the target dataset, the robot needs to continuously learn new knowledge about motion and behavior.

[0071] The 100-dataset can include a total of 100 tasks that require interaction with different objects and utilization of multifunctional motion skills. Among them, these 100 tasks can be further divided into 90 short-range tasks (only one task objective or only one item is regarded as the operating object) and 10 long-range tasks (including two task objectives or requiring operation of two items). The default task order can be used during the experiment.

[0072] Figure 1 Show the environmental settings of each task in the target dataset and the initial state of the robot's robotic arm shown in the exemplary embodiments of the present disclosure.

[0073] Refer to Figure 1, the environment is set to operate in a kitchen environment. Tasks that the robot needs to complete include, for example, the first task in the upper left corner is to open the middle drawer, and the second task is to open the top drawer and then put the bowl in it... 10 pictures correspond to 10 task objectives.

[0074] Next, reference will be made to Figures 2 to 7 specifically describe the method and device for continuous learning of robot skills based on imitation learning of the present disclosure.

[0075] Figure 2 The flowchart of the method for continuous learning of robot skills based on imitation learning in an exemplary embodiment of the present disclosure is shown.

[0076] Reference Figure 2 , for each of multiple tasks in sequence, the following operations 201-204 are repeatedly executed before the preset termination condition is satisfied.

[0077] According to an exemplary embodiment of the present disclosure, the robot skill learning problem can be regarded as a Markov decision process with a finite time

[0078] According to an exemplary embodiment of the present disclosure, for ending the learning of the current task T i , the early stopping method can be adopted: that is, the preset termination condition can be to test the success rate of the imitation learning policy network on the current task at the preset training times checkpoint until the preset termination training times are reached, or after the success rate is higher than the first preset ratio, the success rates obtained from the subsequent continuous preset times checkpoints are lower than the first preset ratio; in the case where the preset termination condition is satisfied, terminate the training of the current task and save the model parameters of the imitation learning policy network in the case of the highest success rate.

[0079] Specifically, the first preset ratio can be 90%, and the continuous preset times can be two or three consecutive times, etc. That is, test the success rate of the model on the current task at the preset training times checkpoint until the preset termination training times are reached or the success rate is higher than 90% and the success rates of the subsequent two or three tests are both lower than 90%, then stop training and save the model parameters of the imitation learning policy network at the highest success rate.

[0080] For the setting of checkpoints, all data of the current task can be used in one round of training, and it is tested once every 5 rounds of training, and a total of 50 rounds of training are performed. Setting intermediate test checkpoints can facilitate viewing the success rate of the model in completing tasks during training for early stopping.

[0081] In operation 201, obtain the observation image and status information of the robot at the current moment.

[0082] Specifically, at the initial stage of training for each task, the continual learning task dataset can be initialized first, including a text description of each task objective (i.e., the text instruction of the current task), the observed images at different moments of the robot completing the corresponding task, the third-person observed images (the observed images of the robot), and the state information of the robot body at each moment (joint pose information, torque, gripper state, etc.). And the types of task sets, the number of tasks, and the task order can be selected.

[0083] One task in the dataset corresponds to a text instruction description. One task has multiple demonstration data, which can be segmented into multiple time series. Each time series contains two observed images at each moment and the state information of the robot body.

[0084] In operation 202, the observed image, the state information, the text instruction of the current task obtained in advance, and the skill codebook of the current task are input into the imitation learning policy network to predict the action information of the robot at the current moment. Among them, the skill codebook of the current task is multiple learnable sub-skill vectors initialized in advance. In the case of having previous tasks, it includes the skill information of the skill codebook of the previous tasks.

[0085] Specifically, the skill codebook of the current task is part of the continual learning algorithm of the present disclosure and can be obtained through initialization.

[0086] According to the exemplary embodiments of the present disclosure, the imitation learning policy network may include, but is not limited to, a perception layer, a skill inference layer, and an action execution layer. The observed image, the state information, and the text instruction can be input into the perception layer to obtain the multi-modal fusion information at the current moment; the multi-modal fusion information and the skill codebook are input into the skill inference layer to predict the skill vector at the current moment; the skill vector is input into the action execution layer to obtain the action information of the robot at the current moment.

[0087] Specifically, the perception layer can help the robot perceive the changing real environment and extract key environmental information according to the task objective; the skill inference layer can infer the skills to be executed based on the environmental information, the task objective, and the state of the robot to obtain the skill vector; the action execution layer can output the continuous value of the end effector action according to the skill vector corresponding to each decision time step.

[0088] The modal fusion of the perception layer can use FiLM to fuse the text information and the visual information to obtain the encoded vector of the multi-modal fusion information or t ; the skill inference layer, that is, the high-level decoder, can be a Transformer decoder. The input is the encoded vector of the multi-modal fusion information or t , and the output is the predicted skill vector z at the current moment t; The action execution layer, i.e., the bottom decoder, is a Transformer decoder with the same structure as the skill inference layer, but they have their own parameters. Its input is the skill vector z predicted at the current moment t , and the output is action information, i.e., the action vector a t .

[0089] For the perception layer, its main function is to help the robot recognize the changing real environment and extract key environmental information according to the task goal, so as to help the subsequent policy network generate a logical action sequence.

[0090] According to an exemplary embodiment of the present disclosure, the perception layer may include, but is not limited to, a pre-trained text encoder, a learnable multi-layer perceptron, a pre-trained learnable visual encoder, a learnable state encoder, and a modality fusion module. The text instruction can be input into the text encoder to obtain a text vector; the text vector is input into the multi-layer perceptron to obtain a text representation; the observed image is input into the visual encoder to obtain an image representation; the state information is input into the state encoder to obtain a state representation; after the text vector and the image representation are fused by the modality fusion module, the fused features are concatenated with the text representation and the state representation to obtain multi-modal fusion information.

[0091] Specifically, the perception layer may include a pre-trained text encoder - CLIP and a learnable multi-layer perceptron (MLP), a pre-trained learnable visual encoder - ResNet-18, a learnable state encoder - multi-layer perceptron (MLP), and a multi-modal fusion module - FiLM. Its input may specifically be: the task goal described in English text (such as Open the top drawer of the cabinet and put the bowl in it, etc.), i.e., the text instruction; the camera picture at the robot's wrist and the third-person perspective picture at the current moment, i.e., the observed image; the robot joint pose information, torque, gripper state, etc., i.e., the state information; and its output is the multi-modal fusion information at the current moment or t .

[0092] To fuse the task instructions of the text with the visual information of the robot, the multi-modal deep learning technology FiLM can be used to extract key environmental information according to the task objectives, so as to adjust the visual information features of the neural network based on the task instruction information, and then achieve modal fusion. The modal fusion module implemented by FiLM can use the task objectives described in the text instructions to correct the robot's visual information by extracting key environmental information; for key environmental information, for example, for the text instruction "Open the top drawer of the cabinet and put the bowl in it", then the upper drawer and the bowl in the visual information are the key information because they are the objects that need to be directly operated on.

[0093] Specifically, for a time series of length seq l en in a task, the text encoder can map the text instructions of the task into a text vector ir, with the dimension of (task i ndex,clip o uput d im); the visual encoder can map the picture sequence information (i.e., the observed image / visual information) of the two perspectives of the robot (the third-person perspective and the robot perspective) into image representations with the dimension of (batch s ize,seq l en,embed s ize); the state encoder can map the state sequence information of the robot's joint poses, torques, and gripper states into state representations with the same dimension

[0094] The body features of the robot are a set of vectors, referring to the joint poses and torques of the robotic arm; the gripper states, which correspond to two representations and two modalities after encoding.

[0095] Finally, the FiLM technology can be used to fuse the two-modal representations of the text information and the image information and then concatenate it with the text representation and the state representation obtained from the text vector through the multi-layer perceptron to obtain the multi-modal fusion information output by the perception layer, with the dimension of (batch s ize,seq l en,num m odalities,embed s ize), where seq l en = 10, num m odalities = 5.

[0096] For the FiLM modal fusion module, a neural network f with an intermediate feature x and an external network g that can output modulation parameters γ and β can be considered. The adjusted feature x' is:

[0097] γ,β = g(z)

[0098] x' = γ ⊙ x + β

[0099] Where z is the input of the external network g, representing the text vector encoded only by CLIP (not further encoded by the multi-layer perceptron at this time), ⊙ represents element-wise multiplication, γ and β are vectors of the same size as x, and x represents the image representation. Each element in it is used to modulate the corresponding feature in x, that is, the text information is encoded by g, and the obtained γ and β are used to correct the image representation, achieving the fusion effect of the text modality and the visual modality, so that the image representation contains text information. This will ultimately enable FiLM to perform conditional calculations on features without explicitly changing the architecture, that is, using the text information vector (equivalent to the encoded representative information about the text) to modulate the visual representation (equivalent to the encoded representative information about the picture). Therefore, the task text instruction embedding (i.e., the text vector) is used as the input of the fully connected feed-forward network g in the modal fusion module. This network outputs the scaling and transformation parameters of the image embedding, that is, γ and β, which also belong to the modal fusion module. These parameters modulate the image embedding before passing it to the high-level decoder.

[0100] After that, the skills to be executed can be inferred based on the environmental information, task objectives, and the state of the robot, obtaining a skill vector.

[0101] According to the exemplary embodiments of the present disclosure, the skill inference layer may include, but is not limited to, an attention module and a high-level decoder. The high-level decoder may include, but is not limited to, a multi-head self-attention layer, a multi-head cross-attention layer, and a feed-forward network layer. The multi-modal fusion information can be input into the attention module to obtain the weights for multiple learnable sub-skill vectors in the skill codebook; based on the weights, multiple learnable sub-skill vectors are weighted to obtain multiple learnable weighted sub-skill vectors; the multi-modal fusion information is input into the multi-head self-attention layer to obtain observation information containing temporal information; the features obtained by concatenating the weighted sub-skill vectors and the observation information, and the observation information are input into the multi-head cross-attention layer, and the output of the multi-head cross-attention layer is input into the feed-forward network layer to predict the skill vector.

[0102] Specifically, the input of the attention module is the multi-modal fusion information, and the output is the weights of the skill vectors in the skill codebook (skill library). According to these weights, a weighted sum of all skill vectors in the skill library can be performed to obtain the selected skill, and then it is combined with the h in the Transformer in the skill inference layer K and hV Combine, and finally the Transformer outputs the skill predicted at the current moment.

[0103] The specific role of the attention module can be to calculate the cosine similarity γ between the result of the Hadamard product of the multimodal fusion information and A and K to obtain a weight vector. Then, this similarity can be used as a weight to select a skill vector from the skill library and splice it into the h of the Transformer in the skill inference layer. K and h V The attention module belongs to a part of the skill inference layer.

[0104] In the high-level decoder of the skill inference layer, the input of the multi-head self-attention layer is the multimodal fusion information, and the output is the multimodal fusion information that incorporates temporal information, that is, the observation information; the input of the multi-head cross-attention layer is the multimodal fusion information incorporating temporal information and the weighted sub-skill vector, and its output passes through a feed-forward network to obtain the skill vector at the current moment.

[0105] According to an exemplary embodiment of the present disclosure, the features obtained by splicing and weighting the sub-skill vector and the observation information can be used as the key and value of the multi-head cross-attention layer, and the observation information can be used as the query of the multi-head cross-attention layer to perform multi-head cross-attention.

[0106] Specifically, the observation information considering temporal information (i.e., the multimodal fusion information incorporating temporal information) can be obtained first through the multi-head self-attention layer, then the weighted sub-skill vector is spliced at each moment, and then "multi-head cross-attention" is performed. The query of "multi-head cross-attention" is the output of the multi-head self-attention, and the key and value are the output of the multi-head self-attention spliced with the sub-skill vector.

[0107] In the high-level decoder of the skill inference layer, the Transformer decoder can be used to infer the skills to be executed based on the environmental information, task objectives, and the state of the robot. In the traditional decoder, the key and value of the cross-attention mechanism are the outputs of the Transformer encoder, which requires explicitly encoding all the skills that the robot may need to perform different tasks. The labor cost of this process is huge; specifically, in the traditional decoder, the key and value of the cross-attention mechanism are the outputs of the Transformer encoder, that is, it is necessary to use the Transformer encoder to encode the skills that the robot may adopt (requiring an additional training process and requiring manual definition of which skills the robot has adopted when completing different tasks, and which data represents which skills, that is, explicitly encoding the skills), which requires a large amount of manpower and material resources.

[0108] The present disclosure can implicitly model skills using a skill codebook (a set of learnable feature vectors) as keys and values, because although the motion trajectories of human experts when performing tasks may seem complex, they can usually be decomposed into more specific and repetitive sub-skills; creating a task-specific skill codebook can capture the sub-skills utilized in each task through network training. Specifically, a learnable skill codebook can be created separately for the i-th sub-task in different continual learning task sets, and it is ensured that each task skill codebook is orthogonal to each other. When training the network using a loss function, the skill codebook can be updated accordingly to model the sub-skills adopted by the manipulator motion trajectory.

[0109] In this way, the model can more effectively learn the underlying structure of expert motion. Moreover, when performing the current task, sub-skill vectors can also be selected from the skill codebooks of previous tasks, realizing sub-skill reuse, which helps to generate diverse and realistic motion trajectories, and complete new tasks by recombining the learned sub-skills according to the task text instructions. Specifically, using the robot multi-modal fusion information encoded by the perception layer or t obtaining the weights of the sub-skill vectors through the attention module, and then calculating the weighted sub-skill vectors. After that, or t first passing through the multi-head self-attention layer, the output of which is then concatenated with the weighted sub-skill vectors to obtain the keys and values of the multi-head cross-attention. Its output itself is used as the query input to the multi-head cross-attention layer and the feed-forward neural network, and then the predicted skill z t is output, with the dimension of (batch s size, seq l len, D).

[0110] The present disclosure can separately create a learnable skill codebook for the i-th sub-task (i = 1, 2, 3,..., 10) in different continual learning task sets where N represents the number of skill primitives (sub-skills), L represents the total number of heads of the keys and values in the multi-head self-attention layer, and D represents the dimension of the hidden layer (in the present invention, N = 10, L = 12, D = embed s size). Subsequently, whenever a new task needs to be learned, the skill codebooks belonging to the old tasks will also participate in the weighted sum of weights in order to utilize the learned knowledge and improve the learning speed. For the skill codebook matrix of one of the self-attention heads each row corresponds to an atomic skill, describing its unique characteristics in the latent space. In addition, there is also a matrix that corresponds one-to-one with this task skill codebook matrix Used to calculate the similarity between the input and different skill vectors. Moreover, in the multi-head self-attention layer of the Transformer decoder (i.e., the high-level decoder) at this layer, the prefix fine-tuning method in prompt learning can be used to calculate the weight of each sub-skill based on the observation information at the current moment t, obtaining the weighted sub-skill, and the weighted sub-skill vector s t Are concatenated with the keys and values in the multi-head cross-attention layer respectively.

[0111] Since the robotic arm needs to perform various similar skill actions when completing different tasks, and each task also has its specific skill actions. For example, for the two tasks of pulling out a drawer and rotating a button, they both include skill actions such as moving the robotic arm forward, left or right, so that the gripper reaches the target position, and also include skill actions such as controlling the gripper to close and grasp the target object and then performing corresponding operations; however, for the task of pulling out the drawer, the robotic arm needs to move backward after grasping the drawer handle to complete the task goal, while for the task of rotating the button, the robotic arm needs to rotate the gripper after grasping the button to complete the task goal. In order to maintain the independence of different task skills while sharing skills, and then improve the learning rate and success rate of new tasks, the present disclosure can obtain a skill vector by performing a weighted sum on the skill codebook:

[0112]

[0113] where, α t,n is the weight vector of the nth sub-skill (n = 1, 2,..., N) matrix at time t,

[0114] including the nth sub-skill matrix in the skill codebook (sub-skill library) learned from the previous i - 1 tasks and initialized for the ith task. The proportion of each sub-skill vector in the sub-skill library in the finally executed skill vector is determined by the multi-modal fusion information or t output by the perception layer and the similarity with each sub-skill vector in the sub-skill library. In addition, each sub-skill matrix S n will initialize a corresponding attention matrix A n for it, which helps to filter out the features related to the current task skill selection from the high-dimensional multi-modal mixed information. The present disclosure can use a simple feature selection attention method to perform an element-wise product on the multi-modal fusion information or t output by the perception layer and the attention matrix A n to create an attention-bearing matrix, and finally perform a similarity matching with the K matrix to obtain the weighted weight of the final skill codebook.

[0115] α t = γ(or t ⊙ A, K)

[0116] = γ(or t ⊙A1,K1),…,γ(or t ⊙A N ,K N )

[0117] To reduce the interference between existing knowledge and newly learned knowledge, the skill codebook of the previous task can be frozen, and orthogonal losses are introduced for P, K, and A:

[0118]

[0119] where B represents an arbitrary matrix.

[0120] Taking one of the multi-head self-attention layers as an example, its input is or t , and the query, key, and value can be represented by h Q , h K and h V respectively. In the exemplary embodiment of the present disclosure, h Q = h K = h V = or′ t , o′ t is the value after dimensional transformation of or t , and the dimension is (batch size , seq len × num m odalities, embed s ize).

[0121] Then the output of this layer can be expressed as:

[0122]

[0123] where W O , is the mapping matrix.

[0124] The present disclosure can evenly divide the weighted sub-skill vector s t into two parts:

[0125]

[0126] Next, in the way of prefix fine-tuning, they can be concatenated with h K and h V :

[0127] f P-T (s,h) = MSA(h Q ,[s K ; h K ,[sV ; h V )

[0128] For the action execution layer, continuous values of the end effector actions can be output according to the skill vectors corresponding to each decision time step.

[0129] According to an exemplary embodiment of the present disclosure, the action execution layer may include, but is not limited to, a bottom decoder and a Gaussian mixture model distribution layer. The bottom decoder may include, but is not limited to, a multi-head self-attention layer and a feed-forward network layer. The skill vector can be processed through the multi-head self-attention layer, the feed-forward network layer, and the Gaussian mixture model distribution layer to obtain the action information of the robot at the current moment.

[0130] Specifically, the bottom decoder in the action execution layer is a Transformer decoder having the same structure as the skill inference, but they have their own parameters. In addition, the action execution layer may further include a Gaussian mixture model distribution module. The input of the action execution layer is the skill z predicted at the current moment t , and the output is the action vector a t . The input of the bottom decoder is the skill vector at the current moment. The vector obtained after passing through the Transformer decoder is then input into the Gaussian mixture model distribution to obtain the reference signal vector of the end effector, that is, the action vector.

[0131] Among them, the skill vector of one time step can correspond to 1 actuator action. For the dimension (batch_size, seq_len, D), it means that batch_size groups of training data are used during training. Each group of training data includes skill vectors with a time length of seq_len, and each skill vector contains D values.

[0132] According to the skill vectors corresponding to the decision time steps, a reference signal, that is, an action (including the three-dimensional coordinates of the position, the attitude angle information, and the gripper switch information), can be generated. The end effector is a structure of a physical robotic arm and is responsible for calculating the joint torque information required to reach the position and attitude in the reference signal based on the current position and attitude of the robotic arm.

[0133] The low-level decoder in the action execution layer can use a Transformer decoder, whose input is the skill vector selected by the skill inference layer at each decision time step, and the output is the corresponding action distribution. However, simple regression cannot fully cover the rich and multi-modal distribution of expert motion trajectories. Even for the same expert facing the same task, the strategies (action trajectories) adopted may not be exactly the same. To solve this problem, the present disclosure can adopt a Gaussian mixture model (GMM) based on a multi-layer perceptron to model the multi-modal trajectory distribution latent in the same skill z. Finally, the robot executes the policy by sampling continuous values a of the end effector action from the output distribution t (batch s ize,seq l en,ac d im), where ac d im = 7. During evaluation, the next action can be parameterized by the mean value of the Gaussian model with the highest probability. The Gaussian mixture model expression is:

[0134]

[0135] where τ is the demonstration trajectory when the expert completes the task, are the parameters of the GMM, and p(τ|θ, c m ) is the m-th sub-Gaussian distribution model c contains M sub-distributions. The weight η m , representing the probability proportion of the m-th sub-model, 0 ≤ η k ≤ 1, GMMs are more expressive than simple multi-layer perceptrons because they can better capture the multi-modal trajectories in expert action data. The ultimate learning objective of the GMM model is to minimize the negative log-likelihood of the expert demonstration action trajectory τ:

[0136]

[0137] In addition, model optimization can also be performed through mode approximation.

[0138] According to an exemplary embodiment of the present disclosure, for the calculation of each head in the multi-head self-attention layer and the multi-head cross-attention layer in the skill inference layer, and the multi-head self-attention layer in the action execution layer, the mode approximation of the current head can be achieved through the following operations: mapping the input of the current head through three projection matrices into queries, keys, and values; stacking the weight matrix of the current head to obtain a first tensor; using CP decomposition to decompose the second tensor into the sum of at least one rank-one component, where the second tensor is a learnable tensor for the current task pre-created using mode approximation parameters, and the components are obtained based on multiple shared components among multiple tasks, specific components of the current task, and a learnable coefficient vector pre-randomly initialized for the current task; based on the queries, keys, and values, the first tensor, and the second tensor after CP decomposition, perform the calculation of the current head.

[0139] Specifically, the CP decomposition vector is part of the continuous learning algorithm of the present disclosure. The CP decomposition vector can be initialized when initializing the continuous learning algorithm.

[0140] The present disclosure proposes a continuous learning algorithm for robot skills based on mode approximation, including but not limited to the following steps:

[0141] Initialize the skill codebook S, matrices A and K for calculating weights, and the CP decomposition vector (task-sharing vectors U, V, and task-specific vector P) for implementing mode approximation;

[0142] When learning task T i initialize S i 、A i 、K i of the current task, and jointly participate in training with the task-sharing decomposition vectors U i 、V i and the task-specific decomposition vector P i .

[0143] More specifically, the task skill codebook S and matrices A and K can be initialized using a uniform distribution, and then Schmidt orthogonalization is performed; U and P can be randomly initialized using a Gaussian distribution, and V can be initialized to 0.

[0144] When learning task T i initialize S i 、A i 、K i of the current task, and jointly participate in training with the task-sharing decomposition vectors U i 、V i and the task-specific decomposition vector P i .

[0145] Perform Schmidt orthogonalization on S, A, and K, and freeze the skill codebook S corresponding to the previous i - 1 tasks (1,i-1), only train the skill codebook S corresponding to the current task i , the task-sharing decomposition vector U i , V i and the task-specific decomposition vector P i .

[0146] The core operation of Transformer is multi-head attention. For the calculation of one head of a multi-head self-attention, i.e., h i , the input can be mapped into queries, keys, and values through three projection matrices , and the attention weight matrices of one head of the multi-head attention can be stacked to obtain the tensor where D is the embed s ize, i.e., the dimension of the embedding. According to the parameter-efficient fine-tuning learning technique, the present disclosure can use the same set of mode approximation (MA) parameters to create a new learnable tensor Δw for learning new tasks and completing task transfer. Specifically, borrowing the idea of CANDECOMP / PARAFAC (CP) decomposition, Δw can be decomposed into the sum of R rank-one components, and each component can be formalized as the outer product of three decomposition vectors, for example:

[0147]

[0148] where is the decomposition vector of the r-th component, and each vector belongs to the corresponding mode matrix. For example, u r is the column vector of U = [u1,…,u r ,…,u R , represents the outer product, λ r is the coefficient scalar of each component, and R is the rank of the CP decomposition. For better understanding, each component forms the corresponding value of the tensor in the form of the sum of scalar products:

[0149]

[0150] where i, j, and k represent the subscripts of the three modes. U and V are the global components shared by tasks and can achieve knowledge transfer between different tasks. P is the task-specific component, and each task has its own P. In addition, to further capture the characteristics of each task, the present disclosure can randomly initialize the learnable coefficient vector for all the attention weight matrices in the i-th task. Using these three mode components, when the input tensor is or t , the learnable tensor ΔW after CP decomposition can be incorporated into the calculation of multi-head attention during the forward propagation process:

[0151]

[0152] The present disclosure may apply the approximation of the attention-based weight matrix in the multi-head self-attention of the skill inference layer, the multi-head cross-attention, and the multi-head self-attention module of the action execution layer.

[0153] Referring to Figure 2 , in operation 203, a loss is calculated based on the predicted action information of the robot and the true action information of the robot in the expert demonstration data.

[0154] Specifically, the loss function of the entire robot skill learning deep network may be:

[0155]

[0156] where ξ is the weighted hyperparameter of the loss function, which can be set to 0.1 in the present disclosure.

[0157] In operation 204, the imitation learning policy network and the skill encoding book of the current task are updated based on the loss to perform the continuous learning process of the robot's skills.

[0158] According to an exemplary embodiment of the present disclosure, when a preset termination condition is satisfied, the parameters of the most important second preset ratio in the convolutional layer and the linear layer in the imitation learning policy network can be frozen as specific parameters of the current task.

[0159] Specifically, the present disclosure proposes a robot skill continuous learning algorithm based on parameter isolation and pattern approximation. When initializing the skill encoding book S, the PackNet pruning object and percentage can be set; for example, the PackNet pruning object can be set to the convolutional layer and the linear layer, and the pruning percentage can be 75%, that is, 25% of the parameters corresponding to the current entire neural network can be retained as the specific parameters of the current task and frozen.

[0160] Next, the specific process of processing the network structure according to the PackNet idea, that is, pruning, fine-tuning, and freezing some network parameters, is introduced.

[0161] The PackNet network parameter mask can be initialized when initializing the continuous learning algorithm. PackNet is a continuous learning algorithm based on dynamic structure, aiming to prevent changes to the parameters that are important for previous tasks in continuous learning. To achieve this, PackNet can iteratively train, prune, fine-tune, and freeze certain parts of the network.

[0162] The pruning process in PackNet consists of two stages. First, based on the neural network parameter mask of task i, it is determined which parameters can be used for learning the current task. Then, the network is trained on this task. At the end of the training, the top 25% of the most important parameters are selected from each convolutional layer and fully connected layer, and the remaining parameters are pruned for learning subsequent tasks. Next, in order to regain better performance on this task, only the remaining parameters are used, and then learning and training are performed again for the current task, that is, fine-tuning, and then freezing. In specific implementation, the present disclosure may not train all bias and normalization layers, and the fine-tuning process may execute the same number of epochs (50 epochs) as the training.

[0163] Next, the entire interaction process of the robot skill learning problem of the present disclosure will be specifically introduced.

[0164] Since the robot skill learning problem can be regarded as a Markov decision process with a finite time This Markov decision process can be modeled as where S and A are the state-action spaces of the robot. μ0 is the initial state distribution, is the reward function, is the transition function.

[0165] If one wants to learn a robot policy from visual and task instruction information that can solve the corresponding task, a sequential decision-making environment can be considered. At time step t = 0, the text instruction i, the initial picture observation x0, and the robot body information p0 are provided to the policy π. The policy π then generates an action distribution π(·∣i,x0,p0), and then samples an action a0 from the distribution and applies it to the robot. Finally, the transition function makes the robot and the environment enter the next state according to the current state and the action taken. This process is repeated, and the policy π continuously samples an action a from the learned distribution t and transmits it to the robot, and stops the loop iteration when the termination condition (the task is completed or the robot collides with the environment, etc.) is reached. The entire interaction process i from the starting step t = 0 to the termination step t = T is called an episode or a piece of expert demonstration data.

[0166] The present disclosure is based on the assumption of a sparse reward setting, and the decision value of whether the target completes the given task can be used to replace the reward function. The goal of the robot can be to maximize the expectation of the reward function on the distribution of the state transition function with the initial state μ0:

[0167]

[0168] Then, based on the robot skill learning deep network at time t-1 and the true action of the robot at time t in the robot image observation, body pose information, corresponding task text instructions, and expert demonstration data during the time period from 0 to t, the robot action instruction predicted by the network can be obtained, and the loss between the predicted action instruction and the true action instruction can be calculated. For example, the negative log-likelihood loss function can be used for calculation to update the robot skill learning deep network at time t.

[0169] Define as N examples of task T k where l k ≤H, where o0 refers to the robot observation, including the text instruction i, the initial image observation x0, and the robot body information p0 (joint and gripper information). In practical applications, the observation o t is not Markovian. Therefore, based on the work of the partially observable Markov process, the historical observations can be integrated to represent the current state s t , that is During the training process, the following objective function can be used for imitation learning:

[0170]

[0171] where is the supervised learning loss function, and the π is a Gaussian mixture model, assuming that the data of the first K-1 tasks cannot be obtained when learning task T K .

[0172] Let t = t + 1, and repeat the above network update steps until the current training round is completed.

[0173] In addition, in the problem of continuous learning of the robot, the robot uses a policy π to successively learn on K tasks T 1 , …, T K . Assume that π is task-conditioned, that is, π(2∣s; T). For each task it is defined as the combination of the initial state distribution the target decision value g k and the text instruction i k . Assume that each task has the same H. When it comes to the kth task T k , the goal of the robot is to optimize the following formula:

[0174]

[0175] When learning the k-th task, for the current task and the k - 1 tasks that have been learned, the action at time t for the p-th task can be obtained according to the policy network π(·∣s; T), that is, the "deep network for robot skill learning based on hierarchical thinking and implicit skill codebook". The state obtained after interacting with the environment In the reward function g of the p-th task p The expectation of the sum of the values at all times is maximized.

[0176] Simply put, it is to search for a policy network that can maximize the reward function value of each task.

[0177] According to an exemplary embodiment of the present disclosure, a method for continuous learning of robot skills based on imitation learning is disclosed, aiming to endow the robotic arm with high task completion ability and continuous learning ability for new tasks. The steps that this method can include are as follows:

[0178] Initialize the continuous learning task dataset, including a text instruction description of each task objective, the observed images at the wrist of the robot at different times when completing the corresponding task, the third-person observed images, and the state information of the robot body at each time (joint pose information, torque, gripper state, etc.); select the type of task set, the number of tasks, and the task order;

[0179] Construct a hierarchical imitation learning policy network, which can include a perception layer, a skill inference layer, and an action execution layer;

[0180] Initialize the continuous learning algorithm: initialize the current task skill codebook, the PackNet network parameter mask, and the CP decomposition vector;

[0181] Learn and train task T i : Use imitation learning to train the network on the current task dataset;

[0182] End the learning of task T i : Adopt the early stopping method to end the learning of the current task, and then process and fine-tune the network structure according to the PackNet idea;

[0183] Let i = i + 1 and continuously learn the next task.

[0184] Among them, constructing a hierarchical imitation learning policy network, specifically, it can be a deep network for robot skill learning based on hierarchical thinking and implicit skill codebook, denoted as Hierarchy-T-C

[0185] (Hierarchy-Transformer-CodeBook). This network can avoid the limitations of manually defining skills, making the robot skill selection more flexible and the connection smoother.

[0186] Initialize the continuous learning algorithm. Specifically, it can be to initialize the continuous learning algorithm for robot skills based on parameter isolation and mode approximation, denoted as PackNet-MA (PackNet-Mode Approximation). This algorithm can achieve almost no forgetting effect, and can also utilize the similarity between tasks to improve the execution ability of previous tasks, and has a certain deployment feasibility. It includes initializing the skill codebook S, A, K matrix for obtaining the weight vector by taking attention with the multi-modal fusion information, and the CP decomposition vectors (task-sharing vectors U, V, and task-specific vector P) for realizing mode approximation, and setting the PackNet pruning object and percentage.

[0187] Learning task T i When, the S of the current task can be initialized i , A i , K i , and jointly with the task-sharing decomposition vectors U i , V i and the task-specific decomposition vector P i participate in the training.

[0188] Figure 3 The flowchart shows the process of data processing by the imitation learning policy network in an exemplary embodiment of the present disclosure.

[0189] Referring to Figure 3 , it shows the process of the imitation learning policy network processing data at two moments. The input of each encoder is the original image observation (visual information), the robot body pose (state) information, and the task target text instruction corresponding to the task at the current moment. The rectangles above each encoder represent the output of the encoder. Modal fusion is used to fuse text information and picture information, that is, the text vector and the image representation. The fusion result is concatenated with the state representation and the text representation further encoded by the multi-layer perceptron to obtain the multi-modal fusion information, that is, the output of the perception layer.

[0190] The output of the perception layer is divided into two paths: one path obtains the weight of the sub-skill vector of the sub-skill codebook through the attention module, and the weighted sub-skill vector is used for the multi-head cross-attention of the skill inference layer; the other path is directly input into the high-level decoder, and after the multi-head self-attention, its output is divided into three paths: one path is used as the query of the multi-head cross-attention, and the remaining two paths are concatenated with the weighted sub-skill vector to obtain the key and value. The query, key, and value are used together for the multi-head cross-attention, and the output enters the feed-forward network. After N layers of such operations, the skill vector at the current moment t is output.

[0191] Specifically, the role of attention is to calculate the cosine similarity between the result of performing a Hadamard product on the multi-modal fusion information and A and K. Subsequently, this similarity can be used as a weight to select skill vectors from the sub-skill codebook and splice them into h of the Transformer decoder in the skill inference layer. K and h V The attention module is part of the skill inference layer.

[0192] The skill vectors are input into the bottom decoder. After passing through N layers of the same (multi-head self-attention + feed-forward network) structure, the output enters the Gaussian mixture model distribution, and the action at the current moment can be obtained.

[0193] Among them, mode approximation can be achieved through CP decomposition for model optimization. The above Δw can be decomposed into the sum of R rank-one components, and each component can be formalized as the outer product of three decomposed vectors P, V, and U, as Figure 3 shown in the lower right corner, where λ is the coefficient scalar of each component.

[0194] Finally, the loss can be calculated based on the predicted action information of the robot and the real action information of the robot in the expert demonstration data.

[0195] According to the exemplary embodiments of the present disclosure, the present disclosure selects skills from the skill codebook through multi-modal perception information, and then controls the action according to the skill output, achieving the goals of online adjusting decisions according to real-time perception information and smooth and coherent connection between skills, overcoming the problem that it is difficult to find an abstract hierarchical structure of sub-behaviors with rich meanings in artificially designed skills; making full use of the PackNet idea and the mode approximation idea to create shared and task-specific parameter and skill libraries, effectively overcoming the forgetting of learned tasks, maintaining good learning ability and having high deployment feasibility.

[0196] According to the exemplary embodiments of the present disclosure, in order to explore the knowledge transfer ability and anti-forgetting performance of the agent in the decision-making task, the following three evaluation metrics can be adopted in the present disclosure: Forward Transfer (FWT) is used to evaluate the ability of the agent to utilize prior knowledge to facilitate the learning of new tasks; Negative Backward Transfer (NBT) is used to evaluate the agent's use of newly acquired knowledge to improve the model's performance on previous tasks; Area Under the Curve (AUC) of the success rate is used to comprehensively evaluate the overall continuous learning ability of the agent. All metrics are calculated based on the success rate because previous literature has shown that for operating strategies, the success rate is a more reliable metric than the training loss. A lower NBT means that the policy has better performance in previously seen tasks and better anti-forgetting ability; a higher FWT means that the policy learns faster on new tasks; and a higher AUC means that the overall performance of NBT and FWT is better. Specifically, define c i,j,e to represent the success rate of the agent that has learned the first i - 1 tasks on the j-th task after training for e rounds (e ∈ {0, 5, …, 50}) on the i-th task. Define c i,i as the highest success rate of the agent over all test rounds e of the current task i, define e * as the earliest training round when the agent achieves the best performance on task i, and assume that for all training rounds For a different task j ≠ i, define In summary, the specific calculation methods of the three evaluation metrics are as follows:

[0197]

[0198] The existing methods used in the experiments are introduced below:

[0199] In the aspect of continuous learning, the present disclosure uses three representative algorithms as comparison objects, namely the representative method based on data replay - Experience Replay (ER), the representative method based on regularization - Elastic Weight Consolidation (EWC), and the representative method based on architecture - PackNet. Experience Replay (ER) stores up to 1000 trajectories to avoid catastrophic forgetting. In each training iteration, the present disclosure uniformly samples 32 replay data trajectories from the memory and combines them with each batch of training data from the new task to participate in the training. The present disclosure uses the online update version of EWC, where the loss function This version uses the exponential moving average along the continuous learning process to update the Fisher information matrix, where k is the number of tasks. γ is set to 0.9 and λ is set to 5·10 4。In the PackNet benchmark, this experiment sets the cropping ratio to 75% and uses the same fine-tuning period as the training period.

[0200] In terms of the robot skill learning policy network, this disclosure uses the RESNET-RNN architecture and the RESNET-T architecture provided by LIBERO as comparison objects. The former uses a pre-trained ResNet as the visual backbone network to encode visual observations at each time step, and lets the LSTM be the temporal backbone network to process a series of encoded visual information. The language instruction is merged into the ResNet features through the FiLM technology to achieve the fusion of the text modality and the visual modality, and then added to the LSTM input respectively; the latter uses a similar ResNet-based visual backbone network, a pre-trained BERT encoder, and a Transformer decoder as the temporal backbone network to process sequences of visual tokens and language tokens.

[0201] During the actual experiment process, the selection of relevant parameters is as follows:

[0202] Hierarchy-T-C, embed s ize = 64, N of the skill codebook A = 100, storing 100 skill vectors, both Transformer decoders are 4 layers, each layer has 6 heads, dropout = 0.1. When paired with the PackNet-MA continual learning algorithm, embed s ize is 384. For each task, the agent is trained for a total of 50 rounds on 50 expert demonstration trajectories, evaluated every 5 rounds, and the agent is tested 20 times on this task, with a maximum number of steps of 600 for each test (if the agent cannot complete the task within 600 steps, the test is judged as failed), and the average success rate is recorded. This disclosure can use the Adam optimizer, with a batch size of 32, and use a cosine scheduler to adjust the learning rate during the training process of each task, from 0.0001 at the beginning to 0.00001 at the end. After completing the training of a task, the model checkpoint with the highest success rate is selected as the final policy for the current task, and then the agent is evaluated on all learned tasks, tested 20 times for each task, and the average success rate is recorded.

[0203] Table 1 records the continual learning performance of three agents using different neural network architectures on four task sets. Experience Replay (ER) is a classic, simple, and efficient continual learning algorithm and is widely used as a benchmark algorithm. Therefore, in this part of the simulation experiment, experience replay is uniformly selected as the continual learning algorithm to compare the performance of the policy neural network.

[0204] Table 1

[0205]

[0206] As can be seen from Table 1, the AUC of RESNET-T and Hierarchy-T-C is higher than that of RESNET-RNN on almost all datasets, indicating that using Transformer may be better than using RNN models when dealing with "time" series. The AUC of Hierarchy-T-C is significantly higher than that of RESNET-T on three datasets except for the long-horizon dataset, and it also has a similar AUC on the long-horizon dataset, indicating that the hierarchical architecture has more advantages in comprehensive ability than the traditional flat strategy architecture. In terms of the NBT metric, Hierarchy-T-C has increased by approximately 17%, 16%, 34%, and 17% compared to RESNET-T on the four datasets (from top to bottom in the table), and by approximately 1%, 18%, 18%, and 18% compared to RESNET-RNN. Especially on the object dataset that only involves a single task objective of "pick up and put down" (e.g., LIBERO-OBJECT), the NBT has almost doubled. This shows that the skill codebook can effectively store the atomic skills that the robot may adopt during different tasks, and continuously corrects them as the training process progresses, which helps the implicit modeling of the robot's skills and their reuse on different tasks, improving the anti-forgetting ability. In addition, the NBT of RESNET-RNN is significantly higher than that of Hierarchy-T-C on all datasets except for the long-horizon dataset, indicating that the flat strategy architecture is not suitable for the scenario of robot continuous learning. However, in terms of the FWT metric, Hierarchy-T-C and RESNET-T each have their own advantages and are both better than RESNET-RNN. Hierarchy-T-C is better than RESNET-T on the target dataset, they perform similarly on the spatial dataset, but RESNET-T is better than Hierarchy-T-C on the object dataset and the long-horizon dataset, indicating that the forward transfer ability and the speed of learning new tasks of Hierarchy-T-C are not as good as those of RESNET-T. This may be because the skill codebook does not reasonably isolate and reuse the skills used in previous tasks when learning new tasks, resulting in interference between tasks and affecting the plasticity of Hierarchy-T-C.

[0207] Figure 4 Shows the success rate of each task of three policy network architectures during the continuous learning process on the target dataset in the exemplary embodiments of the present disclosure.

[0208] Refer to Figure 4, the data points corresponding to the black vertical lines represent the success rates on the k-th task after training on the expert demonstration data of the k-th task, and the right continuous learning data points represent the success rates of the agent on the k-th task after learning the k+1, k+2... 10th tasks. It can be seen from the figure that the success rates of RESNET-RNN at all continuous learning stages in all tasks are significantly lower than those of the other two policy architectures, which once again proves that the learning ability and anti-forgetting performance of RNN are not satisfactory, and its ability to process time series tasks is inferior to that of Transformer. It is worth noting that Hierarchy-T-C has the highest success rate at all continuous learning stages in all tasks, especially in Task 3, with an improvement of about 50% compared to the other two policy architectures. In addition, observing the success rate curve of Task 1 ("Open the drawer in the middle layer of the cabinet"), it is found that after learning the fourth task ("Open the top drawer and put the bowl in it"), the success rates of Hierarchy-T-C and RESNET-T on Task 1 have increased by about 30% and 18% respectively, because both of these tasks contain the goal of "opening the drawer", which indicates that the Transformer network can utilize the similarity between tasks and perform knowledge transfer to help improve the success rate of previous tasks, while RNN basically does not have this ability.

[0209] Table 2 shows the continuous learning performance of four continuous learning algorithms. Since Hierarchy-T-C performs the best among all policy architectures, this neural network architecture is used in all simulation experiments in this part.

[0210] Table 2

[0211]

[0212] As can be seen from Table 2, the PackNet-MA continual learning algorithm proposed in this disclosure outperforms existing classical algorithms in all evaluation metrics. In particular, compared with PackNet, this method has better forward transfer ability, with the FWT improved by approximately 16% and 9% respectively, which indicates that the additionally added task-specific decomposition parameters effectively improve the learning ability of the policy network, enabling it to master new skills more quickly. In addition, this algorithm has improved by approximately 4.4% and 0.7% respectively compared with PackNet in terms of the NBT metric, and both are close to zero, indicating that the additionally added task-specific decomposition parameters can store the information of each task separately, preventing subsequent tasks from affecting it through fine-tuning and freezing, and greatly alleviating catastrophic forgetting. On the long-horizon dataset, the NBT even shows negative numbers, which means that as the robot continuously learns new tasks, the success rate of completing previous tasks also increases, indicating that the task-sharing decomposition parameters can effectively capture the similarities between tasks, and the introduction of the skill encoding book can also store the atomic skills covered in a series of long-horizon tasks, reusing the similar parts and adding and saving the different parts, not only achieving the purpose of anti-forgetting, but also realizing the effect of using the newly learned knowledge to help complete previous tasks. In addition, in terms of the overall performance AUC, this method also has improvements of approximately 14% and 11% respectively compared with PackNet.

[0213] Figure 5 Shows the success rate of each task of three continual learning algorithms during the continual learning process on the target dataset in the exemplary embodiments of the present disclosure.

[0214] Referring Figure 5 , the data points corresponding to the black vertical lines represent the success rate on the k-th task after training on the expert demonstration data of the k-th task, and the data points in the subsequent continual learning process on the right represent the success rate of the agent on the k-th task after learning the k+1, k+2... 10th tasks. Since the success rate of the EWC continual learning algorithm is too low, it is not shown.

[0215] Observing Figure 5It can be seen that the success rates of PackNet-MA and ER on the corresponding tasks after training on Tasks 2, 4, 6, 9, and 10 are significantly higher than those of PackNet, indicating that the plasticity of these two continual learning algorithms is relatively strong, which once again verifies their good forward transfer ability throughout the continual learning stage. This is because in the later stage of learning, after multiple pruning operations, the number of parameters available for learning new tasks in PackNet is extremely small. Even though there is prior knowledge from multiple previous tasks, it does not have the ability to complete new tasks. Secondly, on almost all task sets, the success rates of PackNet-MA and PackNet during the continual learning process can be maintained near the success rates after testing immediately after learning the current task, with only minor fluctuations compared to ER, confirming that the dynamic architecture method has excellent anti-forgetting ability. It is worth noting that in the middle and later stages of continual learning (such as Tasks 6 and 7), the success rate of PackNet-MA can be maintained at a relatively high level, while the success rate of PackNet fluctuates greatly and decreases significantly on these two tasks, indicating that task-specific decomposition parameters help improve anti-forgetting performance. In addition, on Tasks 1, 5, 6, and 7, the success rate of PackNet-MA can continue to increase and be maintained at a relatively high level as the continual learning process progresses, which shows that the shared decomposition vectors between tasks can capture similar skills between different tasks, perform backward transfer of knowledge, and improve the success rate on previous tasks. However, it can also be seen that in Task 4, after learning the 9th task, the success rate on Task 4 drops significantly, while the success rates of the other two continual learning algorithms are not affected, indicating that this task sharing may also cause interference between similar tasks.

[0216] Figure 6 The block diagram of a robot skill continual learning device based on imitation learning in an exemplary embodiment of the present disclosure is shown.

[0217] Referring to Figure 6 FIG. , an exemplary embodiment of the present disclosure further provides a robot skill continual learning device 600 based on imitation learning, which may include, but is not limited to, a task switching module 601, an information acquisition module 602, an action prediction module 603, a loss calculation module 604, and a parameter update module 605.

[0218] The task switching module 601 can sequentially execute the operations in the information acquisition module 602, the action prediction module 603, the loss calculation module 604, and the parameter update module 605 in a loop for each task among multiple tasks before the preset termination condition is met.

[0219] The information acquisition module 602 can acquire the observation image and status information of the robot at the current moment.

[0220] The action prediction module 603 can input the observed image, status information, the text instruction of the current task obtained in advance, and the skill codebook of the current task into the imitation learning policy network to predict the action information of the robot at the current moment. The skill codebook of the current task is a plurality of learnable sub-skill vectors initialized in advance, and in the case of having a previous task, it includes the skill information of the skill codebook of the previous task.

[0221] The loss calculation module 604 can calculate the loss based on the predicted action information of the robot and the true action information of the robot in the expert demonstration data.

[0222] The parameter update module 605 can update the imitation learning policy network and the skill codebook of the current task based on the loss to perform the continuous learning process of the robot's skills.

[0223] According to an exemplary embodiment of the present disclosure, the imitation learning policy network includes a perception layer, a skill inference layer, and an action execution layer. Among them, the action prediction module 603 can input the observed image, status information, and text instruction into the perception layer to obtain the multi-modal fusion information at the current moment; input the multi-modal fusion information and the skill codebook into the skill inference layer to predict the skill vector at the current moment; input the skill vector into the action execution layer to obtain the action information of the robot at the current moment.

[0224] According to an exemplary embodiment of the present disclosure, the perception layer includes a pre-trained text encoder, a learnable multi-layer perceptron, a pre-trained learnable visual encoder, a learnable status encoder, and a modality fusion module. Among them, the action prediction module 603 can input the text instruction into the text encoder to obtain a text vector; input the text vector into the multi-layer perceptron to obtain a text representation; input the observed image into the visual encoder to obtain an image representation; input the status information into the status encoder to obtain a status representation; after fusing the text vector and the image representation through the modality fusion module, splice the fused features with the text representation and the status representation to obtain the multi-modal fusion information.

[0225] According to an exemplary embodiment of the present disclosure, the skill inference layer includes an attention module and a high-level decoder. The high-level decoder includes a multi-head self-attention layer, a multi-head cross-attention layer, and a feed-forward network layer. Among them, the action prediction module 603 can input the multi-modal fusion information into the attention module to obtain weights for multiple learnable sub-skill vectors in the skill codebook. Based on the weights, multiple learnable sub-skill vectors are weighted to obtain multiple learnable weighted sub-skill vectors. The multi-modal fusion information is input into the multi-head self-attention layer to obtain observation information including temporal information. The features obtained by concatenating the weighted sub-skill vectors and the observation information, and the observation information are input into the multi-head cross-attention layer, and the output of the multi-head cross-attention layer is input into the feed-forward network layer to predict the skill vector.

[0226] According to an exemplary embodiment of the present disclosure, the action prediction module 603 can use the features obtained by concatenating the weighted sub-skill vectors and the observation information as the key and value of the multi-head cross-attention layer, and use the observation information as the query of the multi-head cross-attention layer to perform multi-head cross-attention.

[0227] According to an exemplary embodiment of the present disclosure, the action execution layer includes a low-level decoder and a Gaussian mixture model distribution layer. The low-level decoder includes a multi-head self-attention layer and a feed-forward network layer. Among them, the action prediction module 603 can process the skill vector through the multi-head self-attention layer, the feed-forward network layer, and the Gaussian mixture model distribution layer to obtain the action information of the robot at the current moment.

[0228] According to an exemplary embodiment of the present disclosure, for each head in the multi-head self-attention layer and the multi-head cross-attention layer in the skill inference layer, and the multi-head self-attention layer in the action execution layer, the pattern approximation of the current head is achieved through the operation of a pattern approximation module (not shown in the figure). Among them, the pattern approximation module can map the input of the current head through three projection matrices into a query, a key, and a value. Stack the weight matrix of the current head to obtain a first tensor. Use CP decomposition to decompose the second tensor into the sum of at least one rank-one component, where the second tensor is a learnable tensor for the current task pre-created using pattern approximation parameters, and the components are obtained based on multiple shared components among multiple tasks, specific components of the current task, and a learnable coefficient vector pre-randomly initialized for the current task. Based on the query, the key, the value, the first tensor, and the second tensor after CP decomposition, perform the calculation of the current head.

[0229] According to an exemplary embodiment of the present disclosure, the preset termination condition is to check the success rate of the imitation learning policy network on the current task at a preset training times checkpoint until the preset termination training times is reached, or after the success rate is higher than the first preset ratio, the success rates obtained from continuous preset times checkpoints in the subsequent are lower than the first preset ratio; in the case where the preset termination condition is satisfied, terminate the training of the current task and save the model parameters of the imitation learning policy network in the case of the highest success rate.

[0230] According to an exemplary embodiment of the present disclosure, the robot skill continuous learning device 600 based on imitation learning may further include, but is not limited to, a parameter isolation module (not shown in the figure), which can freeze the parameters of the most important second preset ratio in the convolutional layer and the linear layer of the imitation learning policy network as the specific parameters of the current task when the preset termination condition is satisfied.

[0231] It can be understood that in the exemplary embodiment of the robot skill continuous learning device 600 based on imitation learning, the specific implementation process is substantially the same as that of the exemplary embodiment of the robot skill continuous learning method based on imitation learning, and will not be elaborated here. The robot skill continuous learning device 600 based on imitation learning can be respectively configured as software, hardware, firmware or any combination of the above items for performing specific functions. For example, these devices can correspond to dedicated integrated circuits, or can correspond to pure software codes, or can also correspond to modules combining software and hardware. In addition, one or more functions implemented by these devices can also be uniformly executed by components in a physical entity device (such as a processor, a client or a server, etc.).

[0232] Figure 7 The block diagram of an electronic device showing an exemplary embodiment of the present disclosure.

[0233] Referring to Figure 7 , the electronic device 700 includes at least one memory 701 and at least one processor 702. A set of computer-executable instructions is stored in the at least one memory 701. When the set of computer-executable instructions is executed by the at least one processor 702, the robot skill continuous learning method based on imitation learning according to an exemplary embodiment of the present disclosure is executed.

[0234] As an example, the electronic device 700 can be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instruction set. Here, the electronic device 700 does not have to be a single electronic device, and can also be any assembly of devices or circuits that can execute the above instructions (or instruction sets) alone or jointly. The electronic device 700 can also be a part of an integrated control system or a system manager, or can be configured as a portable electronic device interconnected with a local or remote (for example, via wireless transmission) interface.

[0235] In the electronic device 700, the processor 702 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, and the like.

[0236] The processor 702 may execute instructions or code stored in the memory 701, where the memory 701 may also store data. The instructions and data may also be sent and received over a network via a network interface device, where the network interface device may employ any known transmission protocol.

[0237] The memory 701 may be integrated with the processor 702. For example, RAM or flash memory may be disposed within an integrated circuit microprocessor or the like. Additionally, the memory 701 may include a separate device, such as an external disk drive, a storage array, or other storage devices that may be used by any database system. The memory 701 and the processor 702 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, etc., such that the processor 702 can read files stored in the memory.

[0238] Furthermore, the electronic device 700 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 700 may be connected to each other via a bus and / or a network.

[0239] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions may also be provided, where when the instructions are executed by at least one computing device, the at least one computing device is caused to perform the above-described method for continuous learning of robotic skills based on imitation learning.

[0240] Examples of the computer-readable storage medium herein include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device configured to store a computer program and any associated data, data files and data structures in a non-transitory manner and to provide the computer program and any associated data, data files and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium can run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system such that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers. It should be noted that the instructions can also be used to perform additional steps other than the above steps or more specific processing when performing the above steps. The content of these additional steps and further processing has been mentioned during the description of the relevant methods, so it will not be repeated here to avoid redundancy.

[0241] Another embodiment of the present disclosure relates to a system including at least one computing device and at least one storage device storing instructions, wherein, when the instructions are run by the at least one computing device, the at least one computing device is caused to execute the above-mentioned method for continuous learning of robot skills based on imitation learning.

[0242] It should be noted that the system according to the exemplary embodiment of the present disclosure can fully rely on the running of computer programs or instructions to achieve the corresponding functions, that is, each unit corresponds to each step in the functional architecture of the computer program, so that the entire system is called by a dedicated software package (such as, lib library) to achieve the corresponding functions.

[0243] On the other hand, when the above system is implemented in software, firmware, middleware or microcode, the program code or code segment for performing the corresponding operations can be stored in a computer-readable medium such as a storage medium, so that at least one processor or at least one computing device can perform the corresponding operations by reading and running the corresponding program code or code segment.

[0244] According to an exemplary embodiment of the present disclosure, the storage device can be integrated with the computing device. For example, RAM or flash memory can be arranged within an integrated circuit microprocessor, etc. In addition, the storage device can include an independent device, such as an external disk drive, a storage array, or other storage devices that can be used by any database system. The storage device and the computing device can be operatively coupled, or can communicate with each other, for example, through an I / O port, a network connection, etc., so that the computing device can read the instructions stored in the storage device.

[0245] Another embodiment of the present disclosure relates to a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the continuous learning method for robot skills based on imitation learning described in any one of the above is implemented.

[0246] According to the continuous learning method and device for robot skills based on imitation learning provided by the present disclosure, by introducing a skill codebook, a sub-skill library is created separately for each task, and through network training, the sub-skills utilized by each task are captured. When executing the current task, the sub-skill vector can also be selected from the skill codebook of the previous task, realizing the reuse of sub-skills, making the actions generated by the robot more logical, smoother and more coherent when executing long-sequence (more complex or containing multiple goals) tasks.

[0247] In addition, by introducing a hierarchical architecture, the robot can adjust decisions online according to real-time perception information when executing long-sequence (more complex or containing multiple goals) tasks, and learn a reasonable skill abstraction level, having better forward transfer ability and higher average task success rate compared with other policy architectures.

[0248] In addition, after learning each task, pruning the main network, fine-tuning and freezing the parameters of the previous task, and separately creating task-specific parameters (sub-skill library and task-specific decomposition vector) and task-shared parameters for each task endow the robot with the ability to discover the same atomic skills hidden in different tasks, having better backward transfer ability and anti-forgetting compared with existing continuous learning algorithms.

[0249] In addition, on a long-view dataset, it can effectively capture the similarities between tasks (such as information such as skills used or environment layout), transfer the knowledge obtained from learning new tasks, and improve the success rate of completing previous tasks.

[0250] The above describes various exemplary embodiments of the present disclosure. It should be understood that the above description is merely exemplary and not exhaustive, and the present disclosure is not limited to the disclosed exemplary embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the present disclosure. Therefore, the scope of protection of the present disclosure should be determined by the scope of the claims.

Claims

1. A robot skills continuous learning method based on imitation learning, characterized in that: The method comprises: For each of the multiple tasks in turn, loop and perform the following operations until a preset termination condition is met: Get the robot's observation image and status information at the current moment; Inputting the observed image, the state information, the pre-acquired text instructions of the current task, and the skill codebook of the current task into the imitation learning strategy network, predicting the action information of the robot at the current moment, wherein the skill codebook of the current task is a plurality of pre-initialized learnable sub-skill vectors, and in the case of a previous task, contains the skill information of the skill codebook of the previous task; Calculating the loss based on the predicted motion information of the robot and the real motion information of the robot in the expert demonstration data; The imitation learning strategy network and the skill codebook of the current task are updated based on the loss to execute the continuous skill learning process of the robot.

2. The method for continuous learning of robot skills based on imitation learning as claimed in claim 1, characterized in that: The imitation learning strategy network includes a perception layer, a skill inference layer and an action execution layer; The step of inputting the observed image, the state information, the pre-acquired text instructions of the current task, and the skill codebook into the imitation learning strategy network to predict the action information of the robot at the current moment includes: Inputting the observed image, the state information and the text instruction into the perception layer to obtain multimodal fusion information at the current moment; Inputting the multimodal fusion information and the skill codebook into the skill inference layer to predict the skill vector at the current moment; The skill vector is input into the action execution layer to obtain the action information of the robot at the current moment.

3. The method for continuous learning of robot skills based on imitation learning as claimed in claim 2, characterized in that: The perception layer includes a pre-trained text encoder, a learnable multi-layer perceptron, a pre-trained learnable visual encoder, a learnable state encoder, and a modality fusion module; The step of inputting the observed image, the state information and the text instruction into the perception layer to obtain the multimodal fusion information at the current moment includes: Inputting the text instruction into the text encoder to obtain a text vector; Inputting the text vector into the multi-layer perceptron to obtain text representation; Inputting the observed image into the visual encoder to obtain an image representation; Inputting the state information into the state encoder to obtain a state representation; After the text vector and the image representation are fused by the modal fusion module, the fused features are spliced ​​with the text representation and the state representation to obtain the multimodal fusion information.

4. The method for continuous learning of robot skills based on imitation learning as claimed in claim 2, characterized in that: The skill inference layer includes an attention module and a high-level decoder, and the high-level decoder includes a multi-head self-attention layer, a multi-head cross-attention layer, and a feedforward network layer; The step of inputting the multimodal fusion information and the skill codebook into the skill inference layer to predict the skill vector at the current moment includes: Inputting the multimodal fusion information into the attention module to obtain weights for a plurality of learnable sub-skill vectors in the skill codebook; Weighting the multiple learnable sub-skill vectors based on the weights to obtain multiple learnable weighted sub-skill vectors; Inputting the multimodal fusion information into the multi-head self-attention layer to obtain observation information containing time series information; The features obtained by concatenating the weighted sub-skill vector and the observation information, and the observation information are input into the multi-head cross-attention layer, and the output of the multi-head cross-attention layer is input into the feedforward network layer to predict the skill vector.

5. The method for continuous learning of robot skills based on imitation learning as claimed in claim 4, characterized in that: The step of concatenating the weighted sub-skill vector and the features obtained by the observation information, and inputting the observation information into the multi-head cross attention layer includes: The features obtained by concatenating the weighted sub-skill vector and the observation information are used as the key and value of the multi-head cross-attention layer, and the observation information is used as the query of the multi-head cross-attention layer to perform multi-head cross-attention.

6. The method for continuous learning of robot skills based on imitation learning as claimed in claim 4, characterized in that: The action execution layer includes a bottom layer decoder and a Gaussian mixture model distribution layer, and the bottom layer decoder includes a multi-head self-attention layer and a feedforward network layer; The step of inputting the skill vector into the action execution layer to obtain the action information of the robot at the current moment includes: The skill vector is processed by the multi-head self-attention layer, the feedforward network layer and the Gaussian mixture model distribution layer to obtain the action information of the robot at the current moment.

7. The method for continuous learning of robot skills based on imitation learning as claimed in claim 6, characterized in that: For the calculation of each head in the multi-head self-attention layer and the multi-head cross-attention layer in the skill inference layer, and the multi-head self-attention layer in the action execution layer, the mode approximation of the current head is achieved by the following operations: Mapping the input of the current head into query, key and value through three projection matrices; Stacking the weight matrices of the current head to obtain a first tensor; Decomposing the second tensor into a sum of at least one rank-one component using CP decomposition, wherein the second tensor is a learnable tensor for the current task that is pre-created using a pattern approximation parameter, and the component is obtained based on a plurality of shared components between the plurality of tasks, a specific component of the current task, and a learnable coefficient vector that is pre-randomly initialized for the current task; Based on the query, the key and value, the first tensor, and the CP-decomposed second tensor, a computation of the current head is performed.

8. The method for continuous learning of robot skills based on imitation learning as claimed in claim 1, characterized in that: The preset termination condition is to test the success rate of the imitation learning strategy network on the current task at a preset number of training checkpoints until a preset number of termination training times is reached, or after the success rate is higher than a first preset ratio, the success rate obtained by subsequent consecutive preset number of checkpoint tests is lower than the first preset ratio; if the preset termination condition is met, the training of the current task is terminated, and the model parameters of the imitation learning strategy network with the highest success rate are saved.

9. The method for continuous learning of robot skills based on imitation learning as claimed in claim 1, characterized in that: The method further comprises: When the preset termination condition is met, the most important parameters of the second preset ratio in the convolutional layer and the linear layer in the imitation learning strategy network are frozen as specific parameters of the current task.

10. A robot skill continuous learning device based on imitation learning, characterized in that: The device comprises: The task switching module is configured to: for each of the multiple tasks, cyclically execute the operations in the information acquisition module, the action prediction module, the loss calculation module, and the parameter updating module before a preset termination condition is met; The information acquisition module is configured to: acquire the observed image and state information of the robot at the current moment; The action prediction module is configured to: input the observed image, the state information, the pre-acquired text instructions of the current task, and the skill codebook of the current task into the imitation learning strategy network, and predict the action information of the robot at the current moment, wherein the skill codebook of the current task is a plurality of pre-initialized learnable sub-skill vectors, and contains the skill information of the skill codebook of the previous task in the case of a previous task; The loss calculation module is configured to: calculate the loss based on the predicted action information of the robot and the real action information of the robot in the expert demonstration data; The parameter updating module is configured to update the imitation learning strategy network and the skill codebook of the current task based on the loss to execute the continuous skill learning process of the robot.

Citation Information

Patent Citations

  • Systems and methods for data-driven movement skill training

    CN110998696A

  • Robot motion skill learning method fusing text instruction and motion information

    CN117428780A