Robot offline learning method and system based on lightweight diffusion model

Through the combination of lightweight diffusion model and mutual attention mechanism, the problem of limited computing resources in offline learning of robots is solved, the real-time performance and security of robots are improved, and it is suitable for real-time task processing of various environmental changes.

CN120373353APending Publication Date: 2025-07-25NANKAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510230704.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art has problems such as limited computing resources, unstable training and non-convergence in offline robot learning. Especially when the robot has limited computing capabilities, it is difficult to effectively deploy the diffusion model for real-time task processing.

Method used

A lightweight diffusion model is adopted, combining the mutual attention mechanism and the attention mechanism linearization method, a lightweight diffusion model network is designed to reduce the amount of model parameters and delays, and improve the real-time and efficiency of the model on the robot.

Benefits of technology

Through the application of the lightweight diffusion model, the real-time performance and security of the robot in an environment with limited computing resources is improved, battery consumption is reduced, environmental changes are adapted to real-time collaboration and collaborative work are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373353A_ABST
    Figure CN120373353A_ABST
Patent Text Reader

Abstract

The invention discloses a robot offline learning method and system based on a lightweight diffusion model, and belongs to the field of electromechanical system control. And inputting the state information into the trained upper strategy based on the lightweight diffusion model to obtain an action instruction. According to the method, a mutual attention mechanism is introduced to help the model to better capture associated information among different modals, an attention mechanism linearization method is introduced to reduce the calculation complexity of the attention mechanism, and the model parameter quantity and delay of the network are reduced, so that the requirements of the model on calculation resources and storage equipment are reduced; and the operation efficiency of the diffusion model deployed on hardware and used for robot offline learning is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of mechatronic system control, and particularly relates to a robot offline learning method and system based on a lightweight diffusion model. Background Technique

[0002] The statements in this part only mention the background technique related to this application, and do not necessarily constitute the prior art.

[0003] In recent years, the rapid development of reinforcement learning technology has made remarkable progress in the control of robots. Reinforcement learning technology allows robots to automatically learn and adapt in the interaction with the environment, and can continuously improve the strategy according to the feedback information. Secondly, reinforcement learning usually does not require pre-defining complex rules or models, so it is more flexible for complex tasks and environments, which greatly reduces the manual workload of developing robot control systems. However, reinforcement learning usually requires a large number of samples and interactions to train a good strategy, which may not be very practical for actual robot applications, because conducting a large number of interactive experiments in the real world may be expensive or dangerous. In addition, translating the reinforcement learning-based strategy into real-world applications is still an ongoing and challenging task in the field of robotics, because there are various sensor interferences and actuator differences in the real world.

[0004] To solve these problems, many studies focus on recovering effective strategies from previously collected offline data without the need for the agent to interact with the decision-making environment, and the use of this method significantly reduces the cost. However, in offline reinforcement learning, traditional methods usually optimize decisions based on the approximation of the value function or policy function. Since they still follow the traditional temporal difference (TD) learning principle, they face the "fatal triad" challenge in reinforcement learning, which may lead to unstable and non-convergent training. Different from traditional methods, the generative sequence modeling method regards reinforcement learning as a general sequence generation problem and models the joint distribution of state, action, and reward sequences, thus avoiding the "fatal triad" problem and providing an alternative method for offline reinforcement learning. Among them, the diffusion model has a powerful multi-modal distribution capture ability, focuses on the overall accuracy of generating trajectories, and can gracefully handle long-term planning. In addition, the diffusion model models the environmental dynamics and behaviors during the training process, and at the same time allows flexible guidance according to the constraints of different tasks.

[0005] The general implementation of diffusion models usually requires multiple iterations to generate high-quality samples, which poses strict requirements for deployment on robots. The root cause of this challenge lies in the fact that robots are typically limited in computing resources, including aspects such as processor speed, memory, and battery life. In these resource-constrained situations, adopting lightweight deep learning models has significant advantages. Firstly, this approach can effectively avoid problems such as insufficient performance and excessive battery consumption. Secondly, robots need to be able to quickly respond to various changes in the environment, such as obstacle avoidance, navigation, object recognition, and other tasks. Adopting lightweight models can significantly improve the inference speed, contribute to achieving real-time performance, and thus enhance the interactivity and safety of robots. In addition, lightweight models can also reduce the communication bandwidth requirements and communication latency, which is beneficial for realizing real-time collaboration and cooperative work. Therefore, the lightweighting of diffusion models plays a crucial role in robot deployment. It can not only improve the performance, efficiency, and portability of robots but also make them more suitable for various different application scenarios. Summary of the Invention

[0006] To address the deficiencies of the prior art, this application provides a robot offline learning method and system based on a lightweight diffusion model; in the robot offline learning task, a lightweight diffusion model network is designed. The objective of the present invention is to propose a novel lightweight diffusion model, reduce the model parameters and latency of the network, thereby reducing the model's requirements for computing resources and storage devices, and enhancing the real-time performance of the model deployed on the robot.

[0007] In the first aspect, this application provides a robot offline learning method based on a lightweight diffusion model;

[0008] A robot offline learning method based on a lightweight diffusion model includes:

[0009] Obtain state information;

[0010] Input the state information into the trained upper-level policy based on the lightweight diffusion model to obtain an action instruction.

[0011] In the second aspect, this application provides a robot offline learning system based on a lightweight diffusion model;

[0012] A robot offline learning system based on a lightweight diffusion model includes:

[0013] An acquisition module configured to: obtain state information;

[0014] An upper-level policy module configured to: input the state information into the trained upper-level policy based on the lightweight diffusion model to obtain an action instruction.

[0015] In a third aspect, the present application further provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, the one or more computer programs are stored in the memory, and when the electronic device runs, the processor executes the one or more computer programs stored in the memory so that the electronic device executes the method described in the first aspect above.

[0016] In a fourth aspect, the present application further provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, complete the method described in the first aspect.

[0017] In a fifth aspect, the present application further provides a computer program (product), including a computer program, which, when running on one or more processors, is used to implement the method of any item in the foregoing first aspect.

[0018] Compared with the prior art, the beneficial effects of the present application are as follows:

[0019] A novel lightweight diffusion model technology is composed of an encoding module, a bottom-up network, an intermediate module, and a top-down network. The present invention introduces a cross-attention mechanism to help the model better capture the correlation information between different modalities, and introduces a method for linearizing the attention mechanism to reduce the computational complexity of the attention mechanism, greatly improving the running efficiency of the diffusion model for robot offline learning deployed on hardware.

[0020] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The specification drawings constituting a part of the present application are used to provide a further understanding of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application.

[0022] Figure 1 It is a network structure diagram of a new lightweight diffusion model.

[0023] Figure 2 It is a network structure diagram of the encoding module of a new lightweight diffusion model.

[0024] Figure 3 It is an internal network structure diagram of the first basic block of a new lightweight diffusion model.

[0025] Figure 4 It is an internal network structure diagram of the second basic block of a new lightweight diffusion model.

[0026] Figure 5 It is the internal network structure diagram of the sixth basic block of a new lightweight diffusion model. Detailed implementation manners

[0027] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.

[0028] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0029] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0030] Embodiment 1

[0031] This embodiment provides a robot offline learning method based on a lightweight diffusion model;

[0032] As Figure 1 shown, a robot offline learning method based on a lightweight diffusion model includes:

[0033] S101: Obtain state information;

[0034] S102: Input the state information into the trained upper-layer policy based on the lightweight diffusion model to obtain an action instruction;

[0035] As one or more embodiments, the upper-layer policy based on the lightweight diffusion model includes:

[0036] Trajectory initialization and guided sampling of the diffusion model neural network.

[0037] The trajectory initialization, where as an alternative implementation manner, the trajectory representation form is a two-dimensional array formed by concatenating an action sequence and a state sequence, and the actions and states are determined by the action space and the state space.

[0038] As one or more embodiments, the action space is a motion command.

[0039] The initialization noise process of the trajectory is initialized with Gaussian noise;

[0040] The initialized noise is input into the trained diffusion model neural network for multiple guided samplings to obtain a noise-free trajectory;

[0041] As one or more embodiments, such as Figure 1 shown, the lightweight diffusion model network has a network structure including:

[0042] An encoding module and a bottom-up network, an intermediate module, and a top-down network connected in sequence.

[0043] Further, as Figure 2 shown, the structure of the encoding module includes:

[0044] A position encoding layer, a first fully connected layer, a first Mish function layer, and a second fully connected layer connected in sequence;

[0045] The output ends of the second fully connected layer of the encoding module are respectively connected to the input ends of the first module structure, the second module structure, the intermediate module structure, the third module structure, and the fourth module structure.

[0046] Further, the bottom-up network structure includes: a first module structure and a second module structure connected in sequence;

[0047] The first module structure is two first basic blocks and a second basic block connected in sequence;

[0048] The second module structure is two third basic blocks and a fourth basic block connected in sequence;

[0049] Further, the intermediate module network structure includes: three fifth basic blocks, a sixth basic block, and a seventh basic block connected in sequence;

[0050] Further, the top-down network structure includes: a third module structure and a fourth module structure connected in sequence;

[0051] The third module structure is two eighth basic blocks and a ninth basic block connected in sequence;

[0052] The fourth module structure is two tenth basic blocks and an eleventh basic block connected in sequence;

[0053] Among them, the input end of the first basic block is the input trajectory, the input end of the encoding module structure is the diffusion time step, and the output end of the eleventh basic block is the output trajectory;

[0054] Among them, the second basic block is connected to the third basic block, the fourth basic block is connected to the fifth basic block, the outputs of the fourth basic block and the seventh basic block are spliced and connected to the eighth basic block, and the outputs of the second basic block and the ninth basic block are spliced and connected to the tenth basic block;

[0055] Among them, the output ends of the encoding module structure are respectively connected to the input ends of the first basic block, the third basic block, the fifth basic block, the sixth basic block, the seventh basic block, the eighth basic block and the tenth basic block;

[0056] Among them, the internal structures of the first basic block, the third basic block, the fifth basic block, the seventh basic block, the eighth basic block and the tenth basic block are the same;

[0057] Among them, the internal structures of the second basic block, the fourth basic block, the ninth basic block and the eleventh basic block are the same;

[0058] As Figure 3 shown, the internal structure of the first basic block includes:

[0059] The first branch and the second branch in parallel;

[0060] The first branch includes a second Mish function layer and a third fully connected layer connected in sequence;

[0061] The second branch includes a first convolutional layer, a first group normalization layer and a third Mish function layer connected in sequence;

[0062] The data at the output end of the third fully connected layer, the output end of the third Mish function layer and the input end of the first convolutional layer are added and output from the output end of the first basic block.

[0063] The output ends of the encoding module structure are respectively connected to the input ends of the second Mish function layers of the first basic block, the third basic block, the fifth basic block, the seventh basic block, the eighth basic block and the tenth basic block.

[0064] As Figure 4 shown, the internal structure of the second basic block includes:

[0065] The first layer normalization layer and the self-attention layer connected in sequence;

[0066] The data at the input end of the first layer normalization layer and the output end of the self-attention layer are added and output from the output end of the second basic block;

[0067] The output end of the first basic block is connected to the input end of the first normalization layer of the second basic block. The output end of the self-attention layer of the second basic block is connected to the input end of the first convolutional layer of the third basic block. The output end of the third basic block is connected to the input end of the first normalization layer of the fourth basic block. The output end of the self-attention layer of the fourth basic block is connected to the input end of the first convolutional layer of the fifth basic block. The output end of the seventh basic block is connected to the input end of the first normalization layer of the eighth basic block. The output end of the self-attention layer of the eighth basic block is connected to the input end of the first convolutional layer of the ninth basic block. The output end of the ninth basic block is connected to the input end of the first normalization layer of the tenth basic block. The output end of the self-attention layer of the tenth basic block is connected to the input end of the first convolutional layer of the eleventh basic block. The output end of the self-attention layer of the eleventh basic block is the output trajectory.

[0068] As Figure 5 shown, the internal structure of the sixth basic block includes:

[0069] The third branch and the fourth branch arranged in parallel, and the mutual attention layer;

[0070] The third branch includes a fourth Mish function layer and a fourth fully connected layer connected in sequence;

[0071] The fourth branch includes a second normalization layer;

[0072] The output end of the fourth fully connected layer and the output end of the second normalization layer are both connected to the input end of the mutual attention layer;

[0073] The output of the mutual attention layer is added to the data at the input end of the second normalization layer and then output from the output end of the sixth basic block;

[0074] The output end of the fifth basic block is connected to the input end of the second normalization layer of the sixth basic block. The output end of the sixth basic block is connected to the input end of the first convolutional layer of the seventh basic block. The output end of the encoding module is connected to the input end of the fourth Mish function layer of the sixth basic block;

[0075] As one or more embodiments, the working principles of the self-attention and mutual attention layers include:

[0076] Given features from the trajectory and the time step respectively and The calculation of the query matrix Q of the mutual attention layer comes from while the key matrix K and the value matrix V are calculated by as follows:

[0077]

[0078] where W Q,W K ,W V is a learnable mapping matrix, and the query matrix, key matrix, and value matrix of the self-attention layer all come from

[0079] The general form of the attention mechanism is as follows:

[0080]

[0081] where Sim(·,·) represents the similarity function, and O i represents the i-th row of the output matrix. Here, the similarity function is represented in the form of a kernel function, i.e., Sim(Q,K) = φ(Q)φ(K) T , so the form of the attention mechanism can be rewritten as:

[0082]

[0083] Using the associative property of matrix multiplication, it can be further written as:

[0084]

[0085] Here, RuLU is used as the kernel function φ.

[0086] As one or more embodiments, the training steps of the lightweight diffusion model network include:

[0087] Construct a training set, which is an offline data set storing multiple known trajectories;

[0088] Construct a lightweight diffusion model network;

[0089] Input the training set into the constructed lightweight diffusion model network, train the lightweight diffusion model network, and stop training when the loss function of the diffusion model reaches the minimum value to obtain a trained lightweight diffusion model network.

[0090] As one or more embodiments, the working principle of the lightweight diffusion model network includes:

[0091] The encoding module extracts features from the diffusion time step to obtain time features;

[0092] The bottom-up network extracts features from the input trajectory. Each module structure is responsible for extracting one trajectory feature and fusing the time features extracted by the encoding module to obtain two trajectory features in total;

[0093] Input the second trajectory feature into the intermediate module, and the intermediate module processes it and fuses the time features extracted by the encoding module to obtain intermediate features;

[0094] The intermediate features are input into the third module structure of the top-down network for processing. The top-down network processes the intermediate features and the features extracted by the bottom-up network, and fuses the temporal features extracted by the encoding module to obtain the output trajectory.

[0095] As one or more embodiments, the specific working steps of the lightweight diffusion model network include:

[0096] The encoding module extracts features from the diffusion time steps to obtain temporal features;

[0097] The bottom-up network extracts features from the input trajectory. The first module structure fuses the temporal features to extract the first trajectory feature, and the second module structure fuses the temporal features to extract the second trajectory feature;

[0098] The intermediate module processes the second trajectory feature and simultaneously fuses the temporal features to extract the intermediate features;

[0099] The third module structure of the top-down network processes the intermediate features extracted by the intermediate module, and simultaneously fuses the temporal features and the second trajectory feature to extract the third trajectory feature;

[0100] The fourth module structure of the top-down network processes the third trajectory feature and simultaneously fuses the temporal features and the first trajectory feature to obtain the output trajectory.

[0101] Further, as one or more embodiments, the guided sampling process includes:

[0102] Sampling under the gradient guidance of noise-free trajectory estimation, mean estimation, return estimation, and initial state restriction.

[0103] For the noise-free trajectory estimation, the trajectory and the time step are input into the trained lightweight diffusion model network to obtain the estimated noise-free trajectory;

[0104] For the initial state restriction, the initial state of the trajectory is reset using the current state;

[0105] The processes of noise-free trajectory estimation, mean estimation, return estimation under gradient guidance, and initial state restriction are repeated multiple times to obtain the noise-free trajectory, and the first action among them is selected as the action instruction output by the upper-level policy.

[0106] Embodiment 2

[0107] This embodiment provides a robot offline learning system based on a lightweight diffusion model for implementing the above method of the present invention;

[0108] A robot offline learning system based on a lightweight diffusion model includes:

[0109] An acquisition module, which is configured to: acquire status information;

[0110] An upper-layer policy module, which is configured to: input the status information into the trained upper-layer policy based on the lightweight diffusion model to obtain an action instruction.

[0111] It should be noted here that the above acquisition module and the saliency target detection module correspond to steps S101 to S102 in the first embodiment. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the first embodiment. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0112] In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0113] The proposed system can be implemented in other ways. For example, the above-described system embodiments are merely illustrative. For example, the above module division is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0114] Embodiment Three

[0115] This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory so that the electronic device executes the method described in the first embodiment above.

[0116] It should be understood that in this embodiment, the processor can be a central processing unit CPU, and the processor can also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0117] The memory can include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory can also include a non-volatile random memory. For example, the memory can also store information about the device type.

[0118] In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software.

[0119] The method in the first embodiment can be directly implemented by a hardware processor or by a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0120] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0121] Embodiment 4

[0122] This embodiment also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first embodiment is completed.

[0123] The above description is only the preferred embodiment of this application and is not intended to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. A robot offline learning method based on a lightweight diffusion model, characterized in that, Including: S101: Obtain status information; S102: Input the status information into the upper-layer policy based on the lightweight diffusion model after training to obtain an action instruction.

2. The robot offline learning method based on a lightweight diffusion model according to claim 1, characterized in that, The network structure of the lightweight diffusion model includes: An encoding module and a bottom-up network, an intermediate module, and a top-down network connected in sequence.

3. The robot offline learning method based on a lightweight diffusion model according to claim 2, characterized in that, The structure of the encoding module includes: A position encoding layer, a first fully connected layer, a first Mish function layer, and a second fully connected layer connected in sequence; The output ends of the second fully connected layer of the encoding module are respectively connected to the input ends of the first module structure, the second module structure, the intermediate module structure, the third module structure, and the fourth module structure.

4. The robot offline learning method based on a lightweight diffusion model according to claim 2, characterized in that, The structure of the bottom-up network includes: The first module structure and the second module structure connected in sequence; The first module structure is two first basic blocks and a second basic block connected in sequence; The second module structure is two third basic blocks and a fourth basic block connected in sequence; The network structure of the intermediate module includes: Three fifth basic blocks, a sixth basic block, and a seventh basic block connected in sequence; The structure of the top-down network includes: The third module structure and the fourth module structure connected in sequence; The third module structure is two eighth basic blocks and a ninth basic block connected in sequence; The fourth module structure is two tenth basic blocks and an eleventh basic block connected in sequence.

5. The robot offline learning method based on a lightweight diffusion model according to claim 4, characterized in that, The input end of the first basic block is the input trajectory, the input end of the encoding module structure is the diffusion time step, and the output end of the eleventh basic block is the output trajectory; The second basic block is connected to the third basic block, the fourth basic block is connected to the fifth basic block, the outputs of the fourth basic block and the seventh basic block are concatenated and connected to the eighth basic block, and the outputs of the second basic block and the ninth basic block are concatenated and connected to the tenth basic block; The output ends of the encoding module structure are respectively connected to the input ends of the first basic block, the third basic block, the fifth basic block, the sixth basic block, the seventh basic block, the eighth basic block, and the tenth basic block; The internal structures of the first basic block, the third basic block, the fifth basic block, the seventh basic block, the eighth basic block, and the tenth basic block are the same; The internal structures of the second basic block, the fourth basic block, the ninth basic block, and the eleventh basic block are the same; The internal structure of the first basic block includes: A first branch and a second branch in parallel; The first branch includes a second Mish function layer and a third fully connected layer connected in sequence; The second branch includes a first convolutional layer, a first group normalization layer, and a third Mish function layer connected in sequence; The data at the output end of the third fully connected layer, the output end of the third Mish function layer, and the input end of the first convolutional layer are added together and output from the output end of the first basic block; The internal structure of the second basic block includes: A first layer normalization layer and a self-attention layer connected in sequence; The data at the input end of the first layer normalization layer and the output end of the self-attention layer are added together and output from the output end of the second basic block; The internal structure of the sixth basic block includes: A third branch and a fourth branch in parallel, and a cross-attention layer; The third branch includes a fourth Mish function layer and a fourth fully connected layer connected in sequence; The fourth branch includes a second normalization layer; The output end of the fourth fully connected layer and the output end of the second normalization layer are both connected to the input end of the mutual attention layer; The output of the mutual attention layer is added to the data at the input end of the second normalization layer and then output from the output end of the sixth basic block.

6. The robot offline learning method based on a lightweight diffusion model according to claim 5, characterized in that, The self-attention and mutual-attention layer, the working principle includes: Given features from the trajectory and time step respectively and The calculation of the query matrix Q of the mutual attention layer comes from while the key matrix K and value matrix V are calculated by obtained as follows: Among which W Q , W K , W V is a learnable mapping matrix, and the query matrix, key matrix, and value matrix of the self-attention layer all come from The general form of the attention mechanism is as follows: Among them, Sim(·,·) represents the similarity function, and O i represents the i-th row of the output matrix. The similarity function is represented in the form of a kernel function, that is, Sim(Q,K) = φ(Q)φ(K) T , so the form of the attention mechanism can be rewritten as: Utilizing the associative property of matrix multiplication, it is further written as: Here, RuLU is used as the kernel function φ.

7. The offline learning method for a robot based on a lightweight diffusion model according to claim 1, characterized in that The lightweight diffusion model network, the working principle includes: The encoding module extracts features of the diffusion time step to obtain time features; The bottom-up network extracts features of the input trajectory. Each module structure is responsible for extracting one trajectory feature and fusing the time features extracted by the encoding module, resulting in two trajectory features in total; The second trajectory feature is input into the intermediate module. The intermediate module processes it and fuses the time features extracted by the encoding module to obtain intermediate features; The intermediate features are input into the third module structure of the top-down network for processing. The top-down network processes the intermediate features and the features extracted by the bottom-up network and fuses the time features extracted by the encoding module to obtain the output trajectory.

8. A robot offline learning system based on a lightweight diffusion model for implementing the method according to any one of claims 1-7, characterized in that, Including: An acquisition module, which is configured to: acquire status information; An upper-layer policy module, which is configured to: input the status information into the trained upper-layer policy based on the lightweight diffusion model to obtain an action instruction.

9. An electronic device, characterized in that, Including: One or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory so that the electronic device executes the method according to any one of claims 1-7.

10. A computer-readable storage medium for storing computer instructions, characterized in that: When the computer instructions are executed by the processor, the method according to any one of claims 1-7 is completed.