A robot continual imitation learning method and system based on a double diffusion model

CN120395812BActive Publication Date: 2026-10-09HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510427504.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2026-10-09
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

[0004]1.图像生成质量低:现有持续模仿学习方法使用条件GAN或变分自编码器作为生成器,在生成高分辨率图像时往往不稳定,生成的图像质量较差,难以提供足够的视觉信息支持策略学习,尤其是在需要精细感知的任务中(如手术操作、精密装配等)

Benefits of technology

[0069]1. This invention utilizes a latent diffusion model (LDM) to generate high-resolution images in a low-dimensional latent space, significantly improving generation quality and ensuring that the robot can acquire sufficient visual information to support policy learning. It generates continuous task trajectories through a temporal conditional diffusion model, avoiding the trajectory fragmentation problem caused by traditional frame-by-frame generation methods, ensuring the coherence of action sequences. Diffusion generation in a low-dimensional latent space greatly reduces computational complexity, supports real-time generation of high-definition images, and meets practical deployment requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120395812B_ABST
    Figure CN120395812B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of robot learning, and particularly discloses a robot continuous imitation learning method and system based on a double diffusion model. The method comprises the following steps: generating old task data in a latent space by using a pre-trained latent diffusion model; generating image and state data related to the old task by guiding with a task identifier and combining a conditional diffusion generation model; modeling to construct a diffusion model, realizing multi-modal mapping from state to action, generating diversified action sequences; generating high-quality image and state data by using a pre-trained variational autoencoder and a conditional diffusion model; ensuring the consistency in time and the accuracy of global semantics of the generated data by using a time sequence continuity constraint and a semantic consistency constraint; and feedback optimizing training the latent diffusion model. The application can ensure that the model has good generalization ability and stability in complex tasks by dynamically adjusting the loss weight, combining the collaborative training of the imitation learning strategy and the generator, and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot learning technology, and more specifically, relates to a robot continuous imitation learning method and system based on a dual diffusion model. Background Technology

[0002] With the rapid development of robotics technology, imitation learning has become an important means for robots to master complex tasks. By allowing robots to mimic expert behavior (such as human operations or pre-programmed strategies), imitation learning can effectively reduce the complexity of task programming and improve the robot's adaptability in real-world scenarios. However, in practical applications, robots often need to learn multiple tasks consecutively (such as grasping, assembly, and navigation), which brings a key challenge: catastrophic forgetting. That is, when learning a new task, robots often forget previously learned tasks, leading to a significant decline in performance.

[0003] Currently, common methods for solving this problem can be mainly divided into three categories: regularization-based methods, architecture-based methods, and replay-based methods. Among them, replay-based methods have attracted much attention due to their ease of implementation and robust performance. Replay-based methods are mainly divided into two types: experience replay and generative replay. Experience replay stores a portion of the data in memory. When training a new task, ER samples data from memory and combines it with the training data of the current task. However, this method requires a large amount of storage space, making it impractical and lacking scalability in real-world robotic tasks. Generative replay, which generates pseudo-data of old tasks through generative models (such as Generative Adversarial Networks, GANs) and replays this data when learning new tasks to help the robot retain old task data, has been widely used to achieve continuous learning capabilities in different domains. However, existing generative replay-based continuous imitation learning methods for robots have the following main problems:

[0004] 1. Low image generation quality: Existing continuous imitation learning methods use conditional GANs or variational autoencoders as generators, which are often unstable when generating high-resolution images. The generated images are of poor quality and cannot provide enough visual information to support policy learning, especially in tasks that require fine perception (such as surgical operations, precision assembly, etc.).

[0005] 2. Poor temporal continuity: Existing methods typically generate images or states frame by frame, ignoring the temporal correlation of the task trajectory, resulting in fragmented generated trajectories that cannot form a coherent sequence of actions, affecting the smoothness and accuracy of the robot's task execution.

[0006] 3. High computational overhead: Existing methods typically generate images in high-dimensional pixel space, resulting in high computational complexity and difficulty in meeting real-time requirements, especially in scenarios requiring the processing of high-resolution images.

[0007] 4. Lack of support for task diversity: Existing methods have difficulty generating multimodal action distributions and cannot adapt to the multiple solutions that may exist in complex tasks (such as different grasping methods or path planning).

[0008] These problems severely limit the effectiveness of generative playback methods in continuous imitation learning of robots, and there is an urgent need for a solution that can efficiently generate high-quality, continuous trajectories and support multimodal actions. Summary of the Invention

[0009] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a robot continuous imitation learning method and system based on a dual diffusion model. It generates high-quality visual trajectories through a latent diffusion model (LDM), combines a diffusion model-based imitation learning strategy to generate multimodal action distributions, and introduces a bidirectional feedback mechanism to optimize the synergy between generated data and strategy learning, thereby achieving continuous imitation learning for robot visual tasks.

[0010] To achieve the above objectives, according to one aspect of the present invention, a continuous imitation learning method for robots based on a dual diffusion model is proposed, comprising the following steps:

[0011] Step 1: By utilizing a pre-trained latent diffusion model, old task data is generated in the latent space. Guided by task identifiers and combined with a conditional diffusion generation model, image and state data related to the old task are generated. The generated data is then mixed with real data from the new task to form a hybrid dataset.

[0012] Step 2: A diffusion model is constructed through feature encoding, multimodal input fusion, and context modeling of the Transformer encoder to achieve multimodal mapping from state to action, generate diverse action sequences, and maintain the smoothness and continuity of the actions.

[0013] Step 3: High-quality image and state data are generated through pre-trained variational autoencoders and conditional diffusion models. Temporal continuity constraints and semantic consistency constraints are used to ensure the temporal consistency and global semantic accuracy of the generated data.

[0014] Step four: Based on the hybrid dataset and the newly trained imitation learning strategy, the latent diffusion model is trained with feedback optimization. By dynamically adjusting the loss weights, it is ensured that the generator retains the knowledge of the old tasks while generating new task data, thereby achieving continuous optimization and adaptability of the system.

[0015] As a further preferred option, in step one, the task identifier c is utilized. task Generate old task data in the latent space, specifically:

[0016] (11) Generate latent code sequence {z} using conditional diffusion generation model.gen}, and then restored to a concatenated variable of image and state via a VAE decoder: p gen =VAE Decoder (z gen ), break down the splicing variables into images and their corresponding states: o gen ,s gen =split(p gen );

[0017] (12) The generated image o gen With the corresponding state data s gen The pre-trained diffusion model's imitation learning policy module is input together, and action output is obtained through action sampling:

[0018] a gen =Imitation learning strategy based on diffusion model (o gen ,s gen ),

[0019] (13) Generate the old task state-action pair (o gen ,s gen ,a gen ) and real data collected for new tasks Merge into a hybrid dataset

[0020] As a further preferred option, step two includes the following steps:

[0021] (21) Encode the image information in the real collected data or mixed dataset, and concatenate and linearly project it with the robot's body state information to generate a unified joint feature representation, thereby obtaining the joint feature h. t The joint feature h t It retains both visual semantic information and the robot's dynamic state, through the analysis of historical joint feature sequences {h}. t-k ,...,h t Use a Transformer encoder to perform context modeling and obtain the context vector c. t ;

[0022] (22) The conditional diffusion model is adopted. By gradually adding noise and reverse denoising, the multimodal mapping from state to action is learned. The advantage of the diffusion model is that it can generate diverse action sequences while maintaining the smoothness and continuity of the action.

[0023] As a further preferred option, step (21) includes the following steps:

[0024] Using a pre-trained ResNet18 on the input image o tFeature extraction is performed to obtain image features:

[0025]

[0026] Then this feature is compared with the robot's state. After concatenation, the data is mapped to the joint feature space via a linear transformation:

[0027]

[0028] The resulting joint features It retains both visual semantic information and the robot's dynamic state, providing a foundation for subsequent action and trajectory generation.

[0029] As a further preferred option, step (22) includes the following steps:

[0030] First, Gaussian noise is gradually added to the expert demonstration action a0 using a forward diffusion process to simulate the diffusion process of the action distribution:

[0031]

[0032] in, For Gaussian noise, a t For noise dispatching,

[0033] The reverse denoising stage uses a UNet-based network, which takes noise input as input and performs actions a. t Time step t and context vector c t To predict noise, and construct a loss function using the mean squared error:

[0034]

[0035] In the formula, This is the loss function.

[0036] As a further preferred embodiment, step three, which involves generating high-quality image and state data using a pre-trained variational autoencoder and conditional diffusion model, includes:

[0037] First, a pre-trained VAE encoder is used to process the real image. t With robot state s t The concatenated variable p t Compressing it into a low-dimensional latent space yields the latent code z. t :

[0038]

[0039] Then, conditional diffusion is performed in the latent space to generate a joint conditional vector by concatenating the task identifier with the time step t.

[0040] c = Concat(Embed(c) task ),TimeEmbed(t))

[0041] During the forward diffusion stage, noise is added to the latent code:

[0042]

[0043] Inverse denoising is then performed using conditional UNet, with the noise prediction loss being:

[0044]

[0045] In the formula, z t-1 Let be the latent space variables at time step t-1. For noise scheduling, ∈ represents Gaussian noise. The symbol for expected value indicates that the expected value is calculated with respect to the mean squared loss. φ (z t ,t,c) represents the noise predicted by the network.

[0046] As a further preferred embodiment, step three, the temporal continuity constraint, includes:

[0047] We introduce first-order and second-order smoothness constraints to construct the smoothness loss function:

[0048]

[0049] In the formula, λ1 and λ2 are the prime loss weight hyperparameters of the first-order smoothness constraint and the second-order smoothness constraint, respectively.

[0050] The first-order smoothness constraint includes:

[0051]

[0052] The second-order smoothness constraint includes:

[0053]

[0054] As a further preferred embodiment, in step three, the semantic consistency constraint includes:

[0055] For the generated image o gen and real images real High-level features were extracted using frozen ResNet18:

[0056] f gen =ResNet18(o gen ),f real =ResNet18(o real ),

[0057] And L2 distance is used as the feature matching loss:

[0058]

[0059] In the formula, f gen To generate image o gen The extracted features, f real To obtain from real images real High-level features extracted from [the data].

[0060] As a further optimization, in step three, the overall model loss includes:

[0061]

[0062] Where, λ smooth and λ feature These are hyperparameters that control the weights of various losses.

[0063] According to another aspect of the present invention, a rectangular electrical connector robot assembly module under point cloud and geometric constraints is also provided, comprising:

[0064] The data replay module is used to generate old task data in the latent space by utilizing a pre-trained latent diffusion model. Guided by task identifiers and combined with a conditional diffusion generation model, it generates image and state data related to the old task. The generated data is then mixed with real data from the new task to form a hybrid dataset.

[0065] The Imitation Learning Strategy module is used to build a diffusion model through feature encoding, multimodal input fusion, and context modeling of the Transformer encoder, so as to realize multimodal mapping from state to action, generate diverse action sequences, and maintain the smoothness and continuity of the action.

[0066] The image and state generator module is used to generate high-quality image and state data through pre-trained variational autoencoders and conditional diffusion models. Through temporal continuity constraints and semantic consistency constraints, the generated data is ensured to be consistent in time and accurate in global semantics.

[0067] The feedback optimization module is used to perform feedback optimization training on the latent diffusion model based on the hybrid dataset and the newly trained imitation learning strategy. By dynamically adjusting the loss weights, it ensures that the generator retains the knowledge of the old tasks while generating new task data, thereby achieving continuous optimization and adaptability of the system.

[0068] In summary, compared with the prior art, the above-described technical solutions conceived by this invention mainly possess the following technical advantages:

[0069] 1. This invention utilizes a latent diffusion model (LDM) to generate high-resolution images in a low-dimensional latent space, significantly improving generation quality and ensuring that the robot can acquire sufficient visual information to support policy learning. It generates continuous task trajectories through a temporal conditional diffusion model, avoiding the trajectory fragmentation problem caused by traditional frame-by-frame generation methods, ensuring the coherence of action sequences. Diffusion generation in a low-dimensional latent space greatly reduces computational complexity, supports real-time generation of high-definition images, and meets practical deployment requirements.

[0070] 2. The imitation learning strategy based on the diffusion model in this invention can generate diverse action sequences while maintaining the smoothness and continuity of the actions. This makes the robot's actions more natural and precise when performing tasks, effectively improving the quality of task execution. Furthermore, through the feature matching constraint module, the generated images exhibit excellent global semantic consistency, providing more accurate data support for the robot's visual perception.

[0071] 3. The overall system framework of this invention is rationally designed, with each module working collaboratively to maximize the utilization of computing resources. The data replay module efficiently generates old task data, the imitation learning strategy module quickly learns action mappings, and the generator module and feedback optimization module work closely together to achieve continuous model optimization. This efficient resource utilization method enables the system to complete training and optimization in a short time, improving overall work efficiency.

[0072] 4. In the model training and optimization process, this invention reduces the dependence on hardware resources through reasonable module partitioning and efficient algorithm design. For example, the latent diffusion generation module generates data in the latent space, reducing the computational complexity of direct operations in the high-dimensional image space, thereby saving energy and computing resources.

[0073] 5. This invention combines data playback, diffusion models, and feedback optimization techniques to enable the robot to effectively retain knowledge from old tasks while continuously learning new ones. This accumulation and reuse of knowledge not only improves the robot's adaptability to new tasks but also makes it possible for it to continuously learn and grow in complex and ever-changing environments. Attached Figure Description

[0074] Figure 1 This is a framework diagram of a robot continuous imitation learning method based on a dual diffusion model according to an embodiment of the present invention;

[0075] Figure 2 This is a flowchart illustrating the training process of the robot imitation learning strategy based on the diffusion model according to an embodiment of the present invention.

[0076] Figure 3This is a flowchart of the action sampling process of the robot imitation learning strategy based on the diffusion model according to an embodiment of the present invention;

[0077] Figure 4 This is a framework diagram of an image and state generator based on a latent diffusion model according to an embodiment of the present invention. Detailed Implementation

[0078] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0079] like Figure 1 As shown in the figure, this invention provides a robot continuous imitation learning method based on a dual diffusion model. Through data playback and feedback optimization mechanisms, it achieves efficient learning and adaptation capabilities for robots in new task scenarios while retaining knowledge from previous tasks. The system dynamically adjusts loss weights and combines imitation learning strategies with collaborative training of the generator to ensure the model exhibits good generalization ability and stability in complex tasks.

[0080] This embodiment proposes a continuous imitation learning method for robots based on a dual diffusion model. Through data playback and feedback optimization mechanisms, it enables robots to efficiently learn and adapt to new task scenarios while retaining knowledge from previous tasks. The system dynamically adjusts loss weights and combines imitation learning strategies with collaborative training of the generator to ensure the model exhibits good generalization ability and stability in complex tasks.

[0081] (1) System Architecture

[0082] The overall framework of the system is as follows Figure 1 As shown, it mainly includes the following core modules:

[0083] 1. Data Replay Module. This module generates old task data in the latent space using a pre-trained latent diffusion model. Guided by task identifiers and combined with a conditional diffusion generative model, it generates image and state data related to the old task. This data is mixed with real data from the new task to form a hybrid dataset, supporting subsequent training of imitation learning strategies.

[0084] 2. Imitation Learning Strategy Module. This module is built upon a diffusion model, achieving multimodal mapping from state to action through feature encoding, multimodal input fusion, and contextual modeling using a Transformer encoder. The diffusion model generates diverse action sequences while maintaining the smoothness and continuity of the actions through a process of progressively adding noise and reverse denoising. This module is the core of the system, responsible for generating the optimal action strategy based on the input image and state information.

[0085] 3. Image and State Generator Module. This module introduces a latent diffusion generation mechanism, using a pre-trained variational autoencoder (VAE) and a conditional diffusion model to generate high-quality image and state data. Through temporal continuity constraints and semantic consistency constraints, it ensures the temporal consistency of the generated data and the accuracy of global semantics, supporting the robot's vision and state perception in complex tasks.

[0086] 4. Feedback Optimization Module. This module performs feedback optimization training on the latent diffusion model based on a hybrid dataset and a newly trained imitation learning strategy. By dynamically adjusting the loss weights, it ensures that the generator retains knowledge of old tasks while generating new task data, thereby achieving continuous optimization and adaptability of the system.

[0087] The system operation process is as follows:

[0088] 1. Data Preparation. During the learning of a new task, the data replay module uses a latent diffusion model to generate old task data and mixes it with the real data of the new task to form a hybrid dataset.

[0089] 2. Imitation Learning Strategy Training. Based on a hybrid dataset, the imitation learning strategy module learns the mapping relationship from input states to output actions through feature encoding, context modeling, and diffusion model training, generating a strategy suitable for new tasks.

[0090] 3. Generator Optimization. Based on the newly trained imitation learning strategy, the feedback optimization module optimizes the image and state generator, adjusting the generator parameters to improve the quality and consistency of the generated data.

[0091] 4. System Output. Finally, the system outputs an optimized imitation learning strategy and an image and state generator to support the robot's motion planning and visual perception in new task scenarios.

[0092] Through the above framework, the present invention can maintain the importance of old tasks while improving the adaptability to new task objectives when facing an increasing number of new tasks, thereby enabling robots to learn and execute efficiently in complex tasks.

[0093] The entire system adopts a modular design, organically combining robot visual perception, state fusion, motion generation, potential trajectory generation, generated image quality constraints, data playback, and continuous learning.

[0094] The specific modular structure, its functions, and implementation methods are as follows:

[0095] Data playback module. The data playback module utilizes task identifier c task Generate old task data in the latent space. The specific process is as follows: First, generate the latent code sequence {z} using a conditional diffusion generation model. gen}, and then restored to a concatenated variable of image and state via a VAE decoder:

[0096] p gen =VAE Decoder (z gen ),

[0097] Decompose the concatenated variables into images and their corresponding states:

[0098] o gen ,s gen =split(p gen )

[0099] Then, the generated image o gen With the corresponding state data s gen The pre-trained diffusion model's imitation learning policy module is input together, and action output is obtained through action sampling:

[0100] a gen =Imitation learning strategy based on diffusion model (o gen ,s gen ),

[0101] For the specific action sampling process, see section 2.3.

[0102] Ultimately, these generated old task state-action pairs (o gen ,s gen ,a gen ) and real data collected for new tasks Merge into a hybrid dataset

[0103] This module generates image, state, and action data related to the old task. This data is then mixed with real data from the new task to form a hybrid dataset, supporting subsequent training of imitation learning strategies and generators.

[0104] Training a robot imitation learning strategy based on a diffusion model:

[0105] (1) First, the imitation learning dataset (if it is the first task of the entire policy learning process, then the dataset is only the real collected data for that task) is used. If new tasks are learned subsequently, the dataset will be a mixture of old and new tasks. The image information in the image is encoded as a feature, and then spliced ​​and linearly projected with the robot's body state information (such as joint angles and velocities) to generate a unified joint feature representation, which provides multimodal input support for subsequent imitation learning.

[0106] Specifically, it utilizes a pre-trained ResNet18 on the input image. Feature extraction is performed to obtain image features.

[0107]

[0108] Then this feature is compared with the robot's state. After concatenation, the data is mapped to the joint feature space via a linear transformation:

[0109]

[0110] The resulting joint features It retains both visual semantic information and the robot's dynamic state, providing a foundation for subsequent action and trajectory generation.

[0111] By analyzing the historical joint feature sequence {h t-k ,...,h t Use a Transformer encoder to perform context modeling and obtain context vectors.

[0112] c t =TransformerEncoder({h t-k ,...,h t}),

[0113] This process can capture temporal dependencies, ensuring that past state evolution is taken into account when actions are generated.

[0114] (2) Next, we construct an imitation learning strategy based on the diffusion model.

[0115] Traditional methods typically assume a unimodal Gaussian distribution for actions, which is insufficient for the multimodal requirements of complex tasks. This module employs a conditional diffusion model, which learns a multimodal mapping from states to actions through a process of progressively adding noise and reverse denoising. The advantage of the diffusion model is its ability to generate diverse action sequences while maintaining the smoothness and continuity of the actions.

[0116] First, Gaussian noise is gradually added to the expert demonstration action a0 using a forward diffusion process to simulate the diffusion process of the action distribution. The formula is as follows:

[0117]

[0118] in, For Gaussian noise, a t For noise scheduling, the specific parameters are set in the form of a cosine function:

[0119]

[0120] Where T represents the total number of time steps. The reverse denoising stage employs a UNet-based network, using input noise action a. t Time step t and context vector c t To predict noise, i.e.

[0121] ∈ θ (a t ,t,c t ),

[0122] And construct the loss function using the mean squared error:

[0123]

[0124] This process teaches the model how to reconstruct expert actions from noisy actions, thereby generating continuous, smooth, and multimodal action sequences; the overall training process is as follows: Figure 2 As shown.

[0125] Action sampling process for imitation learning strategies based on diffusion models:

[0126] Similar to the process in the data playback module, the input image information is first feature-encoded and then concatenated and linearly projected with the input robot body state information (such as joint angles and velocities) to generate a unified joint feature representation. The historical joint feature sequence is then input into the Transformer encoder for context modeling to obtain the context vector.

[0127] Randomly sampled Gaussian noise random noise a t time step t and context vector c' t Input a trained UNet network and predict noise ∈ θ Perform an iterative denoising process:

[0128] a t-1 =a t -α·∈ θ (a t ,t,c't ),

[0129] Where t∈{1,…,T}. After multiple iterations, the denoised action a0 is finally obtained. The overall action sampling process is as follows: Figure 3 As shown.

[0130] Image and state generator based on latent diffusion model:

[0131] (1) Latent coding and conditional diffusion model

[0132] Meanwhile, to further extend system capabilities and improve the temporal continuity and visual quality of the data, a latent diffusion generation module is introduced. This module first utilizes a pre-trained VAE encoder to generate real images... t With robot state s t The concatenated variable p t Compressing to a low-dimensional latent space yields the latent code representation:

[0133]

[0134] Then, conditional diffusion generation is performed in the latent space. Here, the task identifier c is used... task (After One-Hot encoding and embedding function) is concatenated with time step t (embedded by TimeEmbed) to form a joint condition vector:

[0135] c = Concat(Embed(c) task ),TimeEmbed(t))

[0136] During the forward diffusion stage, noise is added to the latent code:

[0137]

[0138] Inverse denoising is then performed using conditional UNet, with the noise prediction loss being:

[0139]

[0140] (2) Temporal continuity constraint

[0141] To ensure the temporal continuity of the generated latent code sequence, we introduce first-order and second-order smoothness constraints.

[0142] First-order smoothness constraint (encourages smooth changes between adjacent latent codes):

[0143]

[0144] Second-order smoothness constraint (controlling acceleration changes and avoiding drastic jitter):

[0145]

[0146] The final smoothness loss is defined as:

[0147]

[0148] Here, λ1 and λ2 are the prime loss weight hyperparameters for the first-order and second-order smoothness constraints, respectively. Based on the aforementioned first-order and second-order smoothness constraints, the changes in latent variables at adjacent time steps are made smooth, ensuring that the corresponding images or states do not undergo abrupt changes, thereby guaranteeing temporal consistency.

[0149] (3) Semantic consistency constraints

[0150] To further improve the global semantic consistency of the generated images, a feature matching constraint module was embedded in the generation model. Specifically, for the generated image o gen and real images real High-level features were extracted using frozen ResNet18:

[0151] f gen =ResNet18(o gen ),f real =ResNet18(o real ),

[0152] And L2 distance is used as the feature matching loss:

[0153]

[0154] (4) Overall loss of the generative model

[0155] The overall loss of the generative model combines the noise prediction loss, smoothness loss, and feature matching loss:

[0156]

[0157] Where λ smooth and λ feature These are hyperparameters that control the weights of various loss parameters, ensuring that the generated images achieve high standards in both detail and global consistency. The overall framework of the image and state generator based on the latent diffusion model is as follows: Figure 4 As shown.

[0158] Furthermore, in this embodiment of the invention, a robot continuous imitation learning method based on a dual diffusion model is implemented using any one or a combination of the above-mentioned modules. Specifically, it includes the following steps:

[0159] Step 1: By utilizing a pre-trained latent diffusion model, old task data is generated in the latent space. Guided by task identifiers and combined with a conditional diffusion generation model, image and state data related to the old task are generated. The generated data is then mixed with real data from the new task to form a hybrid dataset.

[0160] Step 2: A diffusion model is constructed through feature encoding, multimodal input fusion, and context modeling of the Transformer encoder to achieve multimodal mapping from state to action, generate diverse action sequences, and maintain the smoothness and continuity of the actions.

[0161] Step 3: High-quality image and state data are generated through pre-trained variational autoencoders and conditional diffusion models. Temporal continuity constraints and semantic consistency constraints are used to ensure the temporal consistency and global semantic accuracy of the generated data.

[0162] Step four: Based on the hybrid dataset and the newly trained imitation learning strategy, the latent diffusion model is trained with feedback optimization. By dynamically adjusting the loss weights, it is ensured that the generator retains the knowledge of the old tasks while generating new task data, thereby achieving continuous optimization and adaptability of the system.

[0163] In step one, by utilizing task identifier c task Generate old task data in the latent space, specifically:

[0164] (11) Generate latent code sequence {z} using conditional diffusion generation model. gen}, and then restored to a concatenated variable of image and state via a VAE decoder: p gen =VAE Decoder (z gen ), break down the splicing variables into images and their corresponding states: o gen ,s gen =split(p gen );

[0165] (12) The generated image o gen With the corresponding state data s gen The pre-trained diffusion model's imitation learning policy module is input together, and action output is obtained through action sampling:

[0166] a gen =Imitation learning strategy based on diffusion model (o gen ,s gen ),

[0167] (13) Generate the old task state-action pair (o gen ,s gen ,a gen ) and real data collected for new tasks Merge into a hybrid dataset

[0168] Step two includes the following steps:

[0169] (21) Encode the image information in the real collected data or mixed dataset, and concatenate and linearly project it with the robot's body state information to generate a unified joint feature representation, thereby obtaining the joint feature h. t The joint feature h t It retains both visual semantic information and the robot's dynamic state, through the analysis of historical joint feature sequences {h}. t-k ,…,h t Use a Transformer encoder to perform context modeling and obtain the context vector c. t ;

[0170] (22) The conditional diffusion model is adopted. By gradually adding noise and reverse denoising, the multimodal mapping from state to action is learned. The advantage of the diffusion model is that it can generate diverse action sequences while maintaining the smoothness and continuity of the action.

[0171] Step (21) includes the following steps:

[0172] Using a pre-trained ResNet18 on the input image o t Feature extraction is performed to obtain image features:

[0173]

[0174] Then this feature is compared with the robot's state. After concatenation, the data is mapped to the joint feature space via a linear transformation:

[0175]

[0176] The resulting joint features It retains both visual semantic information and the robot's dynamic state, providing a foundation for subsequent action and trajectory generation.

[0177] Step (22) includes the following steps:

[0178] First, Gaussian noise is gradually added to the expert demonstration action a0 using a forward diffusion process to simulate the diffusion process of the action distribution:

[0179]

[0180] in, For Gaussian noise, a t For noise dispatching,

[0181] The reverse denoising stage uses a UNet-based network, which takes noise input as input and performs actions a. t Time step t and context vector c t To predict noise, and construct a loss function using the mean squared error:

[0182]

[0183] In the formula, This is the loss function.

[0184] Step three, which involves generating high-quality image and state data using a pre-trained variational autoencoder and conditional diffusion model, includes:

[0185] First, a pre-trained VAE encoder is used to process the real image. t With robot state s t The concatenated variable p t Compressing it into a low-dimensional latent space yields the latent code z. t :

[0186]

[0187] Then, conditional diffusion is performed in the latent space to generate a joint conditional vector by concatenating the task identifier with the time step t.

[0188] c = Concat(Embed(c) task ),TimeEmbed(t))

[0189] During the forward diffusion stage, noise is added to the latent code:

[0190]

[0191] Inverse denoising is then performed using conditional UNet, with the noise prediction loss being:

[0192]

[0193] In the formula, z t-1 Let be the latent space variables at time step t-1. For noise scheduling, ∈ represents Gaussian noise. The symbol for expected value indicates that the expected value is calculated with respect to the mean squared loss. φ (z t ,t,c) represents the noise predicted by the network.

[0194] In step three, the temporal continuity constraint includes:

[0195] We introduce first-order and second-order smoothness constraints to construct the smoothness loss function:

[0196]

[0197] In the formula, λ1 and λ2 are the prime loss weight hyperparameters of the first-order smoothness constraint and the second-order smoothness constraint, respectively.

[0198] The first-order smoothness constraint includes:

[0199]

[0200] The second-order smoothness constraint includes:

[0201]

[0202] In step three, the semantic consistency constraints include:

[0203] For the generated image o gen and real images real High-level features were extracted using frozen ResNet18:

[0204] f gen =ResNet18(o gen ),f real =ResNet18(o real ),

[0205] And L2 distance is used as the feature matching loss:

[0206]

[0207] In the formula, f gen To generate image o gen The extracted features, f real To obtain from real images real High-level features extracted from [the data].

[0208] In step three, the overall model loss includes:

[0209]

[0210] Where λ smooth and λ feature These are hyperparameters that control the weights of various losses.

[0211] This embodiment combines data playback, diffusion models, and feedback optimization techniques to enable the robot to effectively retain knowledge from old tasks while continuously learning new ones. This accumulation and reuse of knowledge not only improves the robot's adaptability to new tasks but also makes it possible for it to continuously learn and grow in complex and ever-changing environments.

[0212] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A continuous imitation learning method for robots based on a dual diffusion model, characterized in that, Includes the following steps: Step one involves generating old task data in the latent space using a pre-trained latent diffusion model. Guided by task identifiers and combined with a conditional diffusion generation model, image and state data related to the old task are generated. This generated data is then mixed with real data from the new task to form a hybrid dataset. The step of utilizing task identifiers... Generate old task data in the latent space, specifically: (11) Generate latent code sequences using a conditional diffusion generation model And then, via the VAE decoder, it is restored to a concatenated variable of image and state: Decompose the concatenated variables into images and their corresponding states: ; (12) The generated image With the corresponding state data The pre-trained diffusion model's imitation learning policy module is input together, and action output is obtained through action sampling: (13) Generate the old task state-action pair Data collected in real-world scenarios for new tasks Merge into a hybrid dataset ; Step two involves constructing a diffusion model through feature encoding, multimodal input fusion, and context modeling using a Transformer encoder. This enables multimodal mapping from state to action, generating diverse action sequences while maintaining the smoothness and continuity of the actions. Step two includes the following steps: (21) Encode the image information in the real collected data or mixed dataset, and concatenate and linearly project it with the robot body state information to generate a unified joint feature representation, thereby obtaining the joint feature. This joint feature It retains both visual semantic information and the robot's dynamic state, through the analysis of historical joint feature sequences. Use a Transformer encoder for context modeling to obtain context vectors. ; (22) The conditional diffusion model is adopted. Through the process of gradually adding noise and reverse denoising, the multimodal mapping from state to action is learned. The advantage of the diffusion model is that it can generate diverse action sequences while maintaining the smoothness and continuity of the action. Step (21) includes the following steps: Using pre-trained ResNet18 on the input image Feature extraction is performed to obtain image features: , Then this feature is compared with the robot's state. After concatenation, the data is mapped to the joint feature space via a linear transformation: The resulting joint features It retains both visual semantic information and the robot's dynamic state, providing a foundation for subsequent action generation and trajectory generation; Step (22) includes the following steps: First, the forward diffusion process is used to analyze the expert demonstration actions. Gradually add Gaussian noise to simulate the diffusion process of motion distribution: in, It is Gaussian noise. For noise dispatching, The reverse denoising stage uses a UNet-based network, which operates by processing the input noise. Time step and context vector To predict noise, and construct a loss function using the mean squared error: In the formula, The loss function; Step 3: High-quality image and state data are generated through pre-trained variational autoencoders and conditional diffusion models. Temporal continuity constraints and semantic consistency constraints are used to ensure the temporal consistency and global semantic accuracy of the generated data. Step four: Based on the hybrid dataset and the newly trained imitation learning strategy, the latent diffusion model is trained with feedback optimization. By dynamically adjusting the loss weights, it is ensured that the generator retains the knowledge of the old tasks while generating new task data, thereby achieving continuous optimization and adaptability of the system.

2. The robot continuous imitation learning method based on a dual diffusion model according to claim 1, characterized in that, Step three, which involves generating high-quality image and state data using a pre-trained variational autoencoder and conditional diffusion model, includes: First, a pre-trained VAE encoder is used to process the real image. With robot state spliced ​​variables Compressing into a low-dimensional latent space yields the time step. Latent space variables at time : Then, conditional diffusion generation is performed in the latent space by associating the task identifier with the time step. Concatenate the vectors to form a joint condition vector: During the forward diffusion stage, noise is added to the latent code: Inverse denoising is then performed using conditional UNet, with the noise prediction loss being: In the formula, For time steps Latent space variables at time, For noise dispatching, It is Gaussian noise. The symbol for the expected value indicates that the expected value is calculated with respect to the mean squared error of the loss. The noise predicted by the network.

3. The robot continuous imitation learning method based on a dual diffusion model according to claim 1, characterized in that, In step three, the temporal continuity constraint includes: We introduce first-order and second-order smoothness constraints to construct the smoothness loss function: In the formula, and These are the prime loss weight hyperparameters for the first-order and second-order smoothness constraints, respectively. The first-order smoothness constraint includes: The second-order smoothness constraint includes: 。 4. The robot continuous imitation learning method based on a dual diffusion model according to claim 1, characterized in that, In step three, the semantic consistency constraints include: For the generated image and real images High-level features were extracted using frozen ResNet18: And L2 distance is used as the feature matching loss: In the formula, To generate images The features extracted from them To obtain from real images High-level features extracted from [the data].

5. The robot continuous imitation learning method based on a dual diffusion model according to claim 1, characterized in that, In step three, the overall model loss includes: in, and These are hyperparameters that control the weights of various losses.

6. A robot continuous imitation learning system based on a dual-diffusion model, used to implement the robot continuous imitation learning method based on a dual-diffusion model as described in any one of claims 1-5, characterized in that, include: The data replay module is used to generate old task data in the latent space by utilizing a pre-trained latent diffusion model. Guided by task identifiers and combined with a conditional diffusion generation model, it generates image and state data related to the old task. The generated data is then mixed with real data from the new task to form a hybrid dataset. The Imitation Learning Strategy module is used to build a diffusion model through feature encoding, multimodal input fusion, and context modeling of the Transformer encoder, so as to realize multimodal mapping from state to action, generate diverse action sequences, and maintain the smoothness and continuity of the action. The image and state generator module is used to generate high-quality image and state data through pre-trained variational autoencoders and conditional diffusion models. Through temporal continuity constraints and semantic consistency constraints, the generated data is ensured to be consistent in time and accurate in global semantics. The feedback optimization module is used to perform feedback optimization training on the latent diffusion model based on the hybrid dataset and the newly trained imitation learning strategy. By dynamically adjusting the loss weights, it ensures that the generator retains the knowledge of the old tasks while generating new task data, thereby achieving continuous optimization and adaptability of the system.

Citation Information

Patent Citations

  • Continuous offline reinforcement learning method for double-generation playback based on diffusion

    CN117634647A

  • KR20230008171A