Robot continuous imitation learning method and system based on double diffusion model
Through the robot's continuous imitation learning method based on the potential diffusion model, high-quality, continuous multimodal action sequences are generated, which solves the problems of low image generation quality, poor timing continuity and high computational complexity in the existing methods, and realizes efficient learning and adaptation of the robot in complex tasks.
Patent Information
- Application Number
- CN202510427504.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-08-01
AI Technical Summary
Existing robot continuous imitation learning methods are low in quality, poor timing continuity, high computational complexity and lack of multimodal action support when generating high-resolution images, resulting in poor execution results in complex tasks.
High-quality visual trajectory is generated based on the latent diffusion model (LDM), combined with the imitation learning strategy of the diffusion model, multimodal action distribution is realized through feature coding and Transformer encoder, and a two-way feedback mechanism is introduced to optimize the generated data and strategy learning to ensure the time consistency and global semantic accuracy of the generated data.
The generated action sequence has good coherence and low computational complexity. It can generate high-definition images in real time, supports multimodal actions, and improves the execution quality and adaptability of the robot in complex tasks.
Smart Images

Figure CN120395812A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot learning, and more specifically, relates to a robot continuous imitation learning method and system based on a double diffusion model. Background Technique
[0002] With the rapid development of robot technology, imitation learning has become an important means for robots to master complex tasks. By allowing robots to imitate expert behaviors (such as human operations or pre-programmed strategies), imitation learning can effectively reduce the complexity of task programming and improve the adaptability of robots in real scenarios. However, in practical applications, robots often need to continuously learn multiple tasks (such as grasping, assembly, navigation, etc.), which brings a key challenge: catastrophic forgetting. That is, when learning a new task, robots tend to forget the old tasks they have learned before, resulting in a significant decline in their performance.
[0003] Currently, the common methods to solve this problem are mainly divided into three categories: regularization-based methods, architecture-based methods, and replay-based methods. Among them, the replay-based methods have received much attention because of their easy implementation and robust effects. Replay-based methods mainly consist of two categories, experience replay and generative replay. Experience replay stores a part of the data in the storage memory. When training a new task, ER samples data from the memory and combines it with the training data of the current task. However, this method requires a large amount of storage space, making them impractical and lacking scalability in actual robot tasks. Generative replay, that is, generating pseudo-data of old tasks through a generative model (such as a generative adversarial network GAN) and replaying these data when learning new tasks to help robots retain old tasks, has been widely used to achieve continuous learning capabilities in different fields. However, the existing robot continuous imitation learning methods based on generative replay have the following main problems:
[0004] 1. Low image generation quality: Existing continuous imitation learning methods use conditional GAN or variational autoencoder as generators, which are often unstable when generating high-resolution images, and the generated image quality is poor, making it difficult to provide sufficient visual information to support policy learning, especially in tasks that require fine perception (such as surgical operations, precision assembly, etc.).
[0005] 2. Poor temporal continuity: Existing methods usually generate images or states frame by frame, ignoring the temporal correlation of task trajectories, resulting in fragmented generated trajectories and unable to form a coherent action sequence, affecting the fluency and accuracy of robot task execution.
[0006] 3. High computational cost: Existing methods usually perform generation in a high-dimensional pixel space, with high computational complexity and difficult to meet the real-time requirements, especially in scenarios that require processing high-definition images.
[0007] 4. Lack of support for task diversity: Existing methods are difficult to generate multi-modal action distributions and cannot adapt to multiple solutions that may exist in complex tasks (such as different grasping methods or path planning).
[0008] These problems severely limit the application effect of the generation and playback method in continuous imitation learning of robots. There is an urgent need for a solution that can efficiently generate high-quality, continuous trajectories and support multi-modal actions. Summary of the Invention
[0009] In view of the above defects or improvement requirements of the prior art, the present invention provides a method and system for continuous imitation learning of robots based on a dual diffusion model. It generates high-quality visual trajectories through a latent diffusion model (LDM), combines an imitation learning strategy based on the diffusion model to generate multi-modal action distributions, and introduces a bidirectional feedback mechanism to optimize the coordination of generated data and policy learning, realizing continuous imitation learning of robot vision tasks.
[0010] To achieve the above object, according to one aspect of the present invention, a method for continuous imitation learning of robots based on a dual diffusion model is proposed, including the following steps:
[0011] Step 1, by using a pre-trained latent diffusion model, generate old task data in the latent space. Through the guidance of task identifiers and in combination with a conditional diffusion generation model, generate image and state data related to the old task, and mix the generated data with the real data of the new task to form a mixed data set;
[0012] Step 2, construct a diffusion model through feature encoding, multi-modal input fusion, and context modeling of a Transformer encoder to achieve multi-modal mapping from state to action, generate diverse action sequences, and at the same time maintain the smoothness and continuity of actions;
[0013] Step 3, generate high-quality image and state data through a pre-trained variational autoencoder and a conditional diffusion model, and ensure the temporal consistency and global semantic accuracy of the generated data through temporal continuity constraints and semantic consistency constraints;
[0014] Step 4, based on the mixed data set and the newly trained imitation learning strategy, perform feedback optimization training on the latent diffusion model. By dynamically adjusting the loss weights, ensure that the generator retains the knowledge of the old task while generating new task data, thereby realizing the continuous optimization and adaptation ability of the system.
[0015] As a further preference, in Step 1, generate old task data in the latent space by using the task identifier c task Specifically:
[0016] (i) Use a conditional diffusion generation model to generate a latent code sequence {zgen}, and is restored to the concatenated variable of the image and the state by the VAE decoder: p gen = VAE Decoder (z gen ), split the concatenated variable into the image and the corresponding state: o gen , s gen = split(p gen );
[0017] (12) Input the generated image o gen and the corresponding state data s gen into the pre-trained diffusion model imitation learning strategy module together, and obtain the action output by action sampling:
[0018] a gen = imitation learning strategy based on the diffusion model(o gen , s gen ),
[0019] (13) Combine the generated old task state-action pairs (o gen , s gen , a gen ) with the new task real collected data to form a mixed dataset
[0020] As a further preference, step two includes the following steps:
[0021] (21) Perform feature encoding on the image information in the real collected data or the mixed dataset, splice it with the robot body state information, and perform linear projection to generate a unified joint feature representation, so as to obtain the joint feature h t , and this joint feature h t not only retains the visual semantic information, but also contains the state dynamics of the robot. By using the Transformer encoder to perform context modeling on the historical joint feature sequence {h t-k ,..., h t}, obtain the context vector c t ;
[0022] (22) Adopt a conditional diffusion model, and through the process of gradually adding noise and reverse denoising, learn the multimodal mapping from the state to the action. The advantage of the diffusion model is that it can generate diverse action sequences while maintaining the smoothness and continuity of the actions.
[0023] As a further preference, step (21) includes the following steps:
[0024] Use the pre-trained ResNet18 for the input image o tFeature extraction is performed to obtain image features:
[0025]
[0026] Then this feature is combined with the robot state After splicing, it is mapped to the joint feature space through linear transformation:
[0027]
[0028] Thus, the obtained joint feature not only retains the visual semantic information but also contains the dynamic state of the robot, providing a basis for subsequent action generation and trajectory generation.
[0029] As a further preference, step (22) includes the following steps:
[0030] First, Gaussian noise is gradually added to the expert demonstration action a0 using the forward diffusion process to simulate the diffusion process of the action distribution:
[0031]
[0032] where, is Gaussian noise, a t is the noise schedule,
[0033] In the reverse denoising stage, a UNet-based network is used to predict the noise by inputting the noisy action a t , the time step t, and the context vector c t and construct a loss function with the mean squared error:
[0034]
[0035] In the formula, is the loss function.
[0036] As a further preference, in step three, the generation of high-quality image and state data through the pre-trained variational autoencoder and conditional diffusion model includes:
[0037] First, the pre-trained VAE encoder compresses the concatenated variable p t of the real image o t and the robot state s t into the low-dimensional latent space to obtain the latent code z t :
[0038]
[0039] Then conditional diffusion generation is performed in the latent space. By concatenating the task identifier and the time step t, a joint conditional vector is formed:
[0040] c = Concat(Embed(c task ), TimeEmbed(t))
[0041] In the forward diffusion stage, noise is added to the latent code:
[0042]
[0043] And reverse denoising is performed through the conditional UNet, and its noise prediction loss is:
[0044]
[0045] In the formula, z t-1 is the latent space variable at time step t - 1, is the noise schedule, ∈ is Gaussian noise, is the expected value symbol, indicating taking the expectation of the mean squared error of the loss, ∈ φ (z t , t, c) is the noise predicted by the network.
[0046] As a further optimization, in step three, the temporal continuity constraint includes:
[0047] Introduce first - order and second - order smoothness constraints to construct a smoothness loss function:
[0048]
[0049] In the formula, λ1 and λ2 are the prime loss weight hyperparameters of the first - order smoothness constraint and the second - order smoothness constraint respectively;
[0050] The first - order smoothness constraint includes:
[0051]
[0052] The second - order smoothness constraint includes:
[0053]
[0054] As a further optimization, in step three, the semantic consistency constraint includes:
[0055] For the generated image o gen and the real image o real both use the frozen ResNet18 to extract high - level features:
[0056] f gen = ResNet18(o gen ), f real = ResNet18(o real ),
[0057] And use the L2 distance as the feature matching loss:
[0058]
[0059] In the formula, f gen is the feature extracted from the generated image o gen , and f real is the high-level feature extracted from the real image o real .
[0060] As a further preference, in step three, the overall loss of the model includes:
[0061]
[0062] Among them, λ smooth and λ feature are hyperparameters that control the weights of each loss.
[0063] According to another aspect of the present invention, there is also provided a rectangular electrical connector robot assembly module under point cloud and geometric constraints, including:
[0064] A data playback module, which is used to generate old task data in the latent space by using a pre-trained latent diffusion model, and combine a conditional diffusion generation model through the guidance of a task identifier to generate images and state data related to the old task, and mix the generated data with the real data of the new task to form a mixed data set;
[0065] An imitation learning strategy module, which is used to construct a diffusion model through feature encoding, multi-modal input fusion, and context modeling of a Transformer encoder to achieve multi-modal mapping from state to action, generate diverse action sequences, and maintain the smoothness and continuity of actions at the same time;
[0066] An image and state generator module, which is used to generate high-quality images and state data through a pre-trained variational autoencoder and a conditional diffusion model, and ensure the temporal consistency and global semantic accuracy of the generated data through temporal continuity constraints and semantic consistency constraints;
[0067] A feedback optimization module, which is used to perform feedback optimization training on the latent diffusion model based on the mixed data set and the newly trained imitation learning strategy, and ensure that the generator retains the knowledge of the old task while generating new task data by dynamically adjusting the loss weight, so as to achieve the continuous optimization and adaptability of the system.
[0068] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following technical advantages are mainly possessed:
[0069] 1. The present invention utilizes a latent diffusion model (LDM) to generate high-resolution images in a low-dimensional latent space, significantly improving the generation quality, ensuring that the robot can obtain sufficient visual information to support policy learning, generating continuous task trajectories through a temporal conditional diffusion model, avoiding the trajectory fragmentation problem caused by traditional frame-by-frame generation methods, ensuring the coherence of the action sequence, performing diffusion generation in the low-dimensional latent space, greatly reducing the computational complexity, supporting real-time generation of high-definition images, and meeting the actual deployment requirements; performing diffusion generation in the low-dimensional latent space, greatly reducing the computational complexity, supporting real-time generation of high-definition images, and meeting the actual deployment requirements.
[0070] 2. The imitation learning strategy based on the diffusion model of the present invention can generate diverse action sequences while maintaining the smoothness and continuity of the actions. This enables the robot to perform tasks more naturally and accurately during task execution, effectively improving the quality of task execution. In addition, through the feature matching constraint module, the generated images perform excellently in terms of global semantic consistency, providing more accurate data support for the robot's visual perception.
[0071] 3. The overall system framework of the present invention is reasonably designed, and each module works in coordination, maximizing the utilization of computing resources. The data replay module efficiently generates old task data, the imitation learning strategy module quickly learns action mapping, and the generator module and the feedback optimization module cooperate closely to achieve continuous optimization of the model. This efficient resource utilization method enables the system to complete training and optimization in a short time, improving the overall work efficiency.
[0072] 4. During the model training and optimization process of the present invention, through reasonable module division and efficient algorithm design, the dependence on hardware resources is reduced. For example, the latent diffusion generation module performs data generation in the latent space, reducing the computational complexity of directly operating in the high-dimensional image space, thereby saving energy consumption and computing resources.
[0073] 5. The present invention combines technologies such as data replay, diffusion model, and feedback optimization to enable the robot to effectively retain old task knowledge while continuously learning new tasks. The accumulation and reuse of this knowledge not only improve the robot's adaptability to new tasks but also provide the possibility for it to continuously learn and grow in a complex and changing environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 is a framework diagram of a method for continuous imitation learning of a robot based on a dual diffusion model according to an embodiment of the present invention;
[0075] Figure 2 is a training flowchart of an imitation learning strategy of a robot based on a diffusion model according to an embodiment of the present invention;
[0076] Figure 3It is the action sampling flowchart of the robot imitation learning strategy based on the diffusion model in the embodiment of the present invention;
[0077] Figure 4 It is the framework diagram of the image and state generator based on the latent diffusion model in the embodiment of the present invention. Detailed implementation manners
[0078] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0079] As Figure 1 shown, a robot continuous imitation learning method based on the dual diffusion model provided by the embodiment of the present invention realizes the efficient learning and adaptation ability of the robot in a new task scenario through a data replay and feedback optimization mechanism, while retaining the knowledge of the old task. The system dynamically adjusts the loss weights and combines the collaborative training of the imitation learning strategy and the generator to ensure that the model exhibits good generalization ability and stability in complex tasks.
[0080] This embodiment proposes a robot continuous imitation learning method based on the dual diffusion model, which realizes the efficient learning and adaptation ability of the robot in a new task scenario through a data replay and feedback optimization mechanism, while retaining the knowledge of the old task. The system dynamically adjusts the loss weights and combines the collaborative training of the imitation learning strategy and the generator to ensure that the model exhibits good generalization ability and stability in complex tasks.
[0081] (1) System architecture
[0082] The overall framework of the system is as Figure 1 shown and mainly includes the following core modules:
[0083] 1. Data replay module. This module generates old task data in the latent space by using a pre-trained latent diffusion model. Guided by the task identifier and combined with the conditional diffusion generation model, it generates image and state data related to the old task. These data are mixed with the real data of the new task to form a mixed data set, providing support for the subsequent training of the imitation learning strategy.
[0084] 2. Imitation Learning Strategy Module. This module is built based on the diffusion model and realizes the multimodal mapping from state to action through feature encoding, multimodal input fusion, and context modeling of the Transformer encoder. The diffusion model generates diverse action sequences through a process of gradually adding noise and reverse denoising, while maintaining the smoothness and continuity of the actions. This module is the core of the system and is responsible for generating the optimal action strategy according to the input image and state information.
[0085] 3. Image and State Generator Module. This module introduces a latent diffusion generation mechanism and generates high-quality image and state data through a pre-trained variational autoencoder (VAE) and a conditional diffusion model. Through temporal continuity constraints and semantic consistency constraints, it ensures the temporal consistency of the generated data and the accuracy of the global semantics, providing support for the robot's visual and state perception in complex tasks.
[0086] 4. Feedback Optimization Module. This module performs feedback optimization training on the latent diffusion model based on a mixed dataset and a newly trained imitation learning strategy. By dynamically adjusting the loss weights, it ensures that the generator retains the knowledge of the old tasks while generating new task data, thereby achieving the continuous optimization and adaptability of the system.
[0087] The operation process of the system is as follows:
[0088] 1. Data Preparation. When learning a new task, the data replay module uses the latent diffusion model to generate old task data and mixes it with the real data of the new task to form a mixed dataset.
[0089] 2. Imitation Learning Strategy Training. Based on the mixed dataset, the imitation learning strategy module learns the mapping relationship from the input state to the output action through feature encoding, context modeling, and diffusion model training, and generates a strategy suitable for the new task.
[0090] 3. Generator Optimization. Based on the newly trained imitation learning strategy, the feedback optimization module optimizes the image and state generator, adjusts the parameters of the generator to improve the quality and consistency of the generated data.
[0091] 4. System Output. Finally, the system outputs the optimized imitation learning strategy and the image and state generator, providing support for the robot's action planning and visual perception in the new task scenario.
[0092] Through the above framework, the present invention can, when facing an increasing number of new tasks, not only maintain the emphasis on old tasks but also enhance the adaptability to new task goals, realizing the efficient learning and execution of the robot in complex tasks.
[0093] The entire system adopts a modular design, which organically combines robot vision perception, state fusion, action generation, potential trajectory generation, generated image quality constraints, and data playback and continuous learning.
[0094] The specific modular structure, structural functions, and implementation methods are as follows:
[0095] Data playback module. The data playback module generates old task data in the latent space by using the task identifier c task Specifically, the process is as follows: First, use the conditional diffusion generative model to generate a latent code sequence {z gen}, and restore it to the concatenated variable of the image and state via the VAE decoder:
[0096] p gen = VAE Decoder (z gen ),
[0097] Then, split the concatenated variable into the image and the corresponding state:
[0098] o gen , s gen = split(p gen )
[0099] Subsequently, input the generated image o gen and the corresponding state data s gen into the pre-trained diffusion model imitation learning strategy module, and obtain the action output through action sampling:
[0100] a gen = Imitation learning strategy based on diffusion model (o gen , s gen ),
[0101] See 2.3 for the specific action sampling process.
[0102] Finally, these generated old task state-action pairs (o gen , s gen , a gen ) and the new task real acquisition data are merged into a mixed dataset
[0103] The generation of image, state, and action data related to old tasks is realized through this module. These data are mixed with the real data of new tasks to form a mixed dataset, providing support for subsequent imitation learning strategy training and generator training.
[0104] Train the robot imitation learning strategy based on the diffusion model:
[0105] (1) First, perform feature encoding on the image information in the imitation learning dataset (if it is the first task of the entire policy learning, the dataset is only the real collected data for this task If it is a new task for subsequent learning, then this dataset is a dataset that combines the old and new tasks ) and splice and linearly project it with the robot's own state information (such as joint angles, speeds) to generate a unified joint feature representation, providing multi-modal input support for subsequent imitation learning.
[0106] Specifically, use the pre-trained ResNet18 to extract features from the input image to obtain image features
[0107]
[0108] Then splice this feature with the robot state and map it to the joint feature space through linear transformation after splicing:
[0109]
[0110] Thus, the obtained joint feature not only retains the visual semantic information but also contains the state dynamics of the robot, providing a basis for subsequent action generation and trajectory generation.
[0111] By using the Transformer encoder to perform context modeling on the historical joint feature sequence {h t-k ,...,h t}, a context vector
[0112] c t = TransformerEncoder({h t-k ,...,h t}),
[0113] This process can capture temporal dependencies and ensure that the past state evolution is considered during action generation.
[0114] (2) Next, construct an imitation learning strategy based on the diffusion model.
[0115] Traditional methods usually assume that the action distribution is a unimodal Gaussian distribution, which is difficult to adapt to the multi-modal requirements in complex tasks. This module uses a conditional diffusion model to learn the multi-modal mapping from state to action through the process of gradually adding noise and reverse denoising. The advantage of the diffusion model is that it can generate diverse action sequences while maintaining the smoothness and continuity of actions.
[0116] First, use the forward diffusion process to gradually add Gaussian noise to the expert demonstration action a0 to simulate the diffusion process of the action distribution. The formula is as follows:
[0117]
[0118] where, is Gaussian noise, and a t is the noise schedule, and its specific parameters are set in the form of a cosine function:
[0119]
[0120] where T is the total number of time steps. In the reverse denoising stage, a UNet-based network is used. By inputting the noisy action a t , the time step t, and the context vector c t to predict the noise, that is
[0121] ∈ θ (a t , t, c t ),
[0122] and construct a loss function with the mean square error:
[0123]
[0124] This process enables the model to learn how to restore the expert action from the noisy action, thereby generating a continuous, smooth, and multimodal action sequence; the overall training process is as shown in Figure 2 .
[0125] Action sampling process of the imitation learning strategy based on the diffusion model:
[0126] Same as the process in the data replay module. First, the input image information is feature-encoded, concatenated with the input robot body state information (such as joint angles, speeds), and linearly projected to generate a unified joint feature representation. The historical joint feature sequence is input into the Transformer encoder for context modeling to obtain the context vector
[0127] Randomly sample Gaussian noise Input the random noise a t , the time step t, and the context vector c' t into the trained UNet network to predict the noise ∈ θ , and perform the iterative denoising process:
[0128] a t-1 = a t - α · ∈ θ (a t , t, c't ),
[0129] Among them, \(t\in\{1,\ldots,T\}\). After multiple iterations, the denoised action \(a_0\) is finally obtained. The overall process of action sampling is as Figure 3 shown.
[0130] Image and state generator based on latent diffusion model:
[0131] (1) Latent encoding and conditional diffusion model [[ID=[]]
[0132] Meanwhile, in order to further expand the system capabilities and improve the temporal continuity and visual quality of the data, a latent diffusion generation module is introduced. This module first uses a pre-trained VAE encoder to compress the concatenated variable \(p\) of the real image \(o\) t and the robot state \(s\) t into a low-dimensional latent space to obtain a latent code representation: t
[0133]
[0134] Then, conditional diffusion generation is performed in the latent space. Here, by concatenating the task identifier \(c\) task (after being One-Hot encoded and passed through an embedding function) with the time step \(t\) (embedded by TimeEmbed), a joint conditional vector is formed:
[0135] \(c = Concat(Embed(c\) task ), TimeEmbed(t))
[0136] In the forward diffusion stage, noise is added to the latent code:
[0137] <000?463>
[0138] And reverse denoising is performed through a conditional UNet, and its noise prediction loss is:
[0139]
[0140] (2) Temporal continuity constraint
[0141] To ensure the temporal continuity of the generated latent code sequence, we introduce first-order and second-order smoothness constraints.
[0142] First-order smoothness constraint (encouraging smooth changes between adjacent latent codes):
[0143]
[0144] Second-order smoothness constraint (controlling acceleration changes and avoiding violent jitters):
[0145]
[0145]
[0146] The final smoothness loss is defined as:
[0147]
[0148] where λ1 and λ2 are the prime loss weight hyperparameters of the first-order smoothness constraint and the second-order smoothness constraint respectively. Based on the above first-order and second-order smoothness constraints, the change of latent variables at adjacent time steps is made smooth, ensuring that the corresponding images or states do not undergo sudden changes, thus ensuring temporal consistency.
[0149] (3) Semantic consistency constraint
[0150] To further improve the global semantic consistency of the generated images, a feature matching constraint module is embedded in the generation model. Specifically, for the generated image o gen and the real image o real both use the frozen ResNet18 to extract high-level features:
[0151] f gen = ResNet18(o gen ), f real = ResNet18(o real ),
[0152] and use the L2 distance as the feature matching loss:
[0153]
[0154] (4) Overall loss of the generation model
[0155] The overall loss of the generation model combines the noise prediction loss, the smoothness loss, and the feature matching loss:
[0156]
[0157] where λ smooth and λ feature are hyperparameters that control the weights of each loss, ensuring that the generated images meet high standards in both details and global consistency. The overall framework of the image and state generator based on the latent diffusion model is as Figure 4 shown.
[0158] In addition, in the embodiments of the present invention, a robot continuous imitation learning method based on a double diffusion model is implemented by using any one of the above modules or a combination of multiple modules. Specifically, it includes the following steps:
[0159] Step 1: By using a pre-trained latent diffusion model, generate old task data in the latent space. Guided by the task identifier, combined with a conditional diffusion generation model, generate image and state data related to the old task, and mix the generated data with the real data of the new task to form a mixed dataset;
[0160] Step 2: Construct a diffusion model through feature encoding, multi-modal input fusion, and context modeling of the Transformer encoder to achieve multi-modal mapping from state to action, generate diverse action sequences, and maintain the smoothness and continuity of actions;
[0161] Step 3: Generate high-quality image and state data through a pre-trained variational autoencoder and conditional diffusion model. Ensure the temporal consistency of the generated data and the accuracy of the global semantics through temporal continuity constraints and semantic consistency constraints;
[0162] Step 4: Based on the mixed dataset and the newly trained imitation learning strategy, perform feedback optimization training on the latent diffusion model. By dynamically adjusting the loss weights, ensure that the generator retains the knowledge of the old task while generating new task data, thereby achieving continuous optimization and adaptability of the system.
[0163] In Step 1, generate old task data in the latent space by using the task identifier c task Specifically:
[0164] (11) Use the conditional diffusion generation model to generate a latent code sequence {z gen}, and restore it to the concatenated variable of the image and state via the VAE decoder: p gen = VAE Decoder (z gen ), split the concatenated variable into the image and the corresponding state: o gen , s gen = split(p gen );
[0165] (12) Input the generated image o gen and the corresponding state data s gen into the pre-trained diffusion model imitation learning strategy module, and obtain the action output through action sampling:
[0166] a gen = imitation learning strategy based on the diffusion model (o gen , s gen ),
[0167] [[ID=�0]](13) Combine the generated old task state-action pairs (o gen , s gen , a gen ) with the real data collected from the new task Merge into a hybrid dataset
[0168] Step 2 includes the following steps:
[0169] (21) Perform feature encoding on the image information in the real acquisition data or hybrid dataset, splice and linearly project it with the robot ontology state information to generate a unified joint feature representation, thereby obtaining the joint feature h t , this joint feature h t not only retains the visual semantic information but also contains the state dynamics of the robot. By performing context modeling on the historical joint feature sequence {h t-k ,…,h t} using a Transformer encoder, a context vector c t is obtained;
[0170] (22) Adopt a conditional diffusion model. Through the process of gradually adding noise and reverse denoising, learn the multi-modal mapping from state to action. The advantage of the diffusion model is that it can generate diverse action sequences while maintaining the smoothness and continuity of the actions.
[0171] Step (21) includes the following steps:
[0172] Use the pre-trained ResNet18 to extract features from the input image o t to obtain image features:
[0173]
[0174] Then, after splicing this feature with the robot state , map it to the joint feature space through a linear transformation:
[0175]
[0176] Thus, the obtained joint feature not only retains the visual semantic information but also contains the state dynamics of the robot, providing a basis for subsequent action generation and trajectory generation.
[0177] Step (22) includes the following steps:
[0178] First, gradually add Gaussian noise to the expert demonstration action a0 using the forward diffusion process to simulate the diffusion process of the action distribution:
[0179]
[0180] where is Gaussian noise, a t is the noise schedule,
[0181] In the reverse denoising stage, a UNet-based network is adopted, and by inputting the noisy action a t , the time step t, and the context vector c t to predict the noise, and a loss function is constructed with the mean square error:
[0182]
[0183] In the formula, is the loss function.
[0184] In step three, the generation of high-quality images and state data through the pre-trained variational autoencoder and conditional diffusion model includes:
[0185] First, use the pre-trained VAE encoder to compress the concatenated variable p t of the real image o t and the robot state s t into the low-dimensional latent space to obtain the latent code z t :
[0186]
[0187] Then, perform conditional diffusion generation in the latent space. By concatenating the task identifier and the time step t, a joint conditional vector is formed:
[0188] c = Concat(Embed(c task ), TimeEmbed(t))
[0189] In the forward diffusion stage, noise is added to the latent code:
[0190]
[0191] And perform reverse denoising through the conditional UNet, and its noise prediction loss is:
[0192]
[0193] In the formula, z t-1 is the latent space variable at time step t - 1, is the noise schedule, ∈ is Gaussian noise, is the expected value symbol, indicating taking the expectation of the mean square error of the loss, ∈ φ (z t , t, c) is the noise predicted by the network.
[0194] In step three, the described temporal continuity constraint includes:
[0195] Introduce first-order and second-order smoothness constraints to construct a smoothness loss function:
[0196]
[0197] Wherein, λ1 and λ2 are respectively the prime loss weight hyperparameters of the first-order smoothness constraint and the second-order smoothness constraint;
[0198] The first-order smoothness constraint includes:
[0199]
[0200] The second-order smoothness constraint includes:
[0201]
[0202] In step three, the semantic consistency constraint includes:
[0203] For the generated image o gen and the real image o real Both use the frozen ResNet18 to extract high-level features:
[0204] f gen = ResNet18(o gen ), f real = ResNet18(o real ),
[0205] And use the L2 distance as the feature matching loss:
[0206]
[0207] Wherein, f gen is the feature extracted from the generated image o gen and f real is the high-level feature extracted from the real image o real .
[0208] In step three, the overall loss of the model includes:
[0209]
[0210] Where λ smooth and λ feature are hyperparameters that control the weights of each loss.
[0211] This embodiment combines technologies such as data replay, diffusion model, and feedback optimization to enable the robot to effectively retain the knowledge of old tasks while continuously learning new tasks. This accumulation and reuse of knowledge not only improves the robot's adaptability to new tasks but also provides the possibility for it to continuously learn and grow in a complex and changing environment.
[0212] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A continuous imitation learning method for robots based on a double diffusion model, characterized in that, It includes the following steps: Step 1: By using a pre-trained latent diffusion model, generate old task data in the latent space. Under the guidance of task identifiers, combined with a conditional diffusion generation model, generate image and state data related to the old task, and mix the generated data with the real data of the new task to form a mixed dataset; Step 2: Construct a diffusion model through feature encoding, multi-modal input fusion, and context modeling of the Transformer encoder to achieve multi-modal mapping from state to action, generate diverse action sequences, and at the same time maintain the smoothness and continuity of actions; Step 3: Generate high-quality image and state data through a pre-trained variational autoencoder and a conditional diffusion model. Through temporal continuity constraints and semantic consistency constraints, ensure the temporal consistency of the generated data and the accuracy of global semantics; Step 4: Based on the mixed dataset and the newly trained imitation learning strategy, perform feedback optimization training on the latent diffusion model. By dynamically adjusting the loss weights, ensure that the generator retains the knowledge of the old task while generating new task data, thereby achieving continuous optimization and adaptability of the system.
2. The continuous imitation learning method of a robot based on a double-diffusion model according to claim 1, wherein In step one, by using the task identifier c task Generate old task data in the latent space, specifically: (11) Generate a latent code sequence {z gen} using a conditional diffusion generative model, and restore it to a concatenated variable of an image and a state via a VAE decoder: p gen = VAE Decoder (z gen ), split the concatenated variable into an image and the corresponding state: o gen , s gen = split(p gen ); (12) Input the generated image o gen along with the corresponding state data s gen into the pre-trained diffusion model imitation learning strategy module, and obtain the action output through action sampling: a gen = Imitation learning strategy based on diffusion model (o gen ,s gen ), (13) Combine the generated old task state-action pairs (o gen , s gen , a gen ) with the newly collected real data of the task to form a mixed dataset 3. A continuous imitation learning method for a robot based on a double diffusion model according to claim 1, characterized in that, Step 2 includes the following steps: (21) Feature-encode the image information in the real acquisition data or the mixed dataset, splice it with the robot's own state information, and perform linear projection to generate a unified joint feature representation, thereby obtaining the joint feature h t , this joint feature h t not only retains the visual semantic information but also contains the state dynamics of the robot. By using a Transformer encoder to perform context modeling on the historical joint feature sequence {h t-k , …, h t}, the context vector c t is obtained; (22) Adopt a conditional diffusion model. Through the process of gradually adding noise and reverse denoising, learn the multi-modal mapping from state to action. The advantage of the diffusion model is that it can generate diverse action sequences while maintaining the smoothness and continuity of actions.
4. A continuous imitation learning method for a robot based on a double-diffusion model according to claim 1, characterized in that Step (21) includes the following steps: Use the pre-trained ResNet18 to extract features from the input image o t to obtain image features: Then this feature is combined with the robot state After splicing, it is mapped to the joint feature space through linear transformation: The resulting joint features not only retain the visual semantic information but also contain the dynamic state of the robot, providing a basis for subsequent action generation and trajectory generation.
5. A continuous imitation learning method for a robot based on a double-diffusion model according to claim 1, characterized in that Step (22) includes the following steps: First, use the forward diffusion process to gradually add Gaussian noise to the expert demonstration action a0 to simulate the diffusion process of the action distribution: Among them, is Gaussian noise, and a t is the noise schedule, In the reverse denoising stage, a UNet-based network is adopted, and the noise is predicted by inputting the noisy action a t , the time step t, and the context vector c t to predict the noise, and a loss function is constructed with the mean square error: In the formula, is the loss function.
6. A robotic assembly method for a rectangular electrical connector under point cloud and geometric constraints according to claim 1, characterized in that, In Step 3, the generation of high-quality image and state data through the pre-trained variational autoencoder and the conditional diffusion model includes: First, the pre-trained VAE encoder is used to compress the concatenated variable p of the real image o t and the robot state s t into a low-dimensional latent space to obtain the latent space variable z at time step t t : t : Then perform conditional diffusion generation in the latent space. By concatenating the task identifier with the time step t, form a joint conditional vector: c = Concat(Embed(c task ), TimeEmbed(t)) In the forward diffusion stage, add noise to the latent code: And perform reverse denoising through the conditional UNet, and its noise prediction loss is: where z t-1 is the latent space variable at time step t-1, is the noise schedule, ∈ is Gaussian noise, is the expected value symbol, representing the expectation of the mean squared error of the loss, ∈ φ (z t , t, c) is the noise predicted by the network.
7. A robotic assembly method for a rectangular electrical connector under point cloud and geometric constraints according to claim 1, characterized in that In Step 3, the temporal continuity constraint includes: Introduce first-order and second-order smoothness constraints to construct a smoothness loss function: In the formula, λ1 and λ2 are the prime loss weight hyperparameters of the first-order smoothness constraint and the second-order smoothness constraint respectively; The first-order smoothness constraint includes: The second-order smoothness constraint includes:
8. A method for robot assembly of a rectangular electrical connector under point cloud and geometric constraints according to claim 1, characterized in that In Step 3, the semantic consistency constraint includes: For the generated image o gen and the real image o real Both use the frozen ResNet18 to extract high-level features: f gen = ResNet18(o gen ), f real = ResNet18(o real ), And use the L2 distance as the feature matching loss: where f gen is the feature extracted from the generated image o gen , and f real is the high-level feature extracted from the real image o real .
9. A method for robotic assembly of a rectangular electrical connector under point cloud and geometric constraints according to claim 1, characterized in that, In Step 3, the overall loss of the model includes: where λ smooth and λ feature are hyperparameters that control the weights of the respective losses.
10. A rectangular electrical connector robot assembly module under point cloud and geometric constraints, characterized in that, It includes: A data replay module, used to generate old task data in the latent space by using a pre-trained latent diffusion model. Under the guidance of task identifiers, combined with a conditional diffusion generation model, generate image and state data related to the old task, and mix the generated data with the real data of the new task to form a mixed dataset; An imitation learning strategy module, used to construct a diffusion model through feature encoding, multi-modal input fusion, and context modeling of the Transformer encoder to achieve multi-modal mapping from state to action, generate diverse action sequences, and at the same time maintain the smoothness and continuity of actions; An image and state generator module, which is used to generate high-quality image and state data through a pre-trained variational autoencoder and a conditional diffusion model, and ensure the temporal consistency of the generated data and the accuracy of global semantics through temporal continuity constraints and semantic consistency constraints; A feedback optimization module, which is used to perform feedback optimization training on the latent diffusion model based on a hybrid dataset and a newly trained imitation learning strategy, and ensure that the generator retains the knowledge of old tasks while generating new task data by dynamically adjusting the loss weights, so as to achieve the continuous optimization and adaptability of the system.
Citation Information
Patent Citations
Continuous offline reinforcement learning method for double-generation playback based on diffusion
CN117634647A
Robot welding seam identifying and tracking method based on deep learning
CN118514068A
Industrial robot motion planning method based on diffusion model
CN119217373A
Ice maker and refrigerator
KR1020230018502A
Equivariant trajectory optimization with diffusion models
US20240273261A1
Cited By
Imitation learning method, device and equipment for intelligent agent with body
CN121189522A
Mechanical arm imitation learning method based on teleoperation and smooth constraint diffusion strategy
CN121447649A
Cross-platform humanoid robot non-speech teaching behavior rapid generation system and method
CN121468674A
Method and device for controlling end executing mechanism of robot
CN121589819A
Method and device for learning robot operation strategy from human video
CN122033925A