Robot control methods, systems, equipment, and media for diffusion-strategy skip sampling
By using a diffusion strategy and step sampling method, the problems of multimodality and high computational overhead in humanoid robot control are solved, and efficient and real-time control strategy generation is achieved, which is applicable to humanoid robot control systems.
Patent Information
- Application Number
- CN202510403191.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-04-01
AI Technical Summary
Existing reinforcement learning methods struggle to capture multimodality in humanoid robot control, diffusion strategies in online learning are computationally expensive, lack effective training methods, and struggle to balance action optimization and behavior imitation, as well as lack accurate methods for evaluating policy distribution.
A step sampling method based on diffusion strategy is adopted. By initializing network parameters and hyperparameters, calculating the Q function gradient to estimate state sensitivity, adaptively adjusting step parameters, performing DDIM sampling and environmental interaction, combining information theory to evaluate policy distribution, and optimizing Q network and policy network.
It significantly reduces computational complexity and sampling time, improves the efficiency and quality of policy generation, and realizes real-time and flexible control policies, which are particularly suitable for humanoid robot control and can efficiently handle high-dimensional motion spaces.
Smart Images

Figure CN120295188B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot control technology, and more specifically, to robot control methods, systems, devices, and media using diffusion strategy skip sampling. Background Technology
[0002] The development of humanoid robot control systems increasingly highlights the urgent need for complex control strategies. While traditional control methods perform well in structured environments, they often exhibit limited adaptability in dynamic, unstructured scenarios, struggling to cope with environmental changes and task diversity. Reinforcement learning (RL) learns optimal policies through continuous interaction with the environment, providing an effective solution to complex decision-making problems. Compared to traditional methods, reinforcement learning has significant advantages in end-to-end learning, online adaptation, and policy optimization, effectively handling high-dimensional state spaces and continuous action spaces—characteristics crucial for humanoid robot motion control and skill acquisition.
[0003] However, existing technologies have the following problems:
[0004] Traditional reinforcement learning methods typically model policies as conditional Gaussian distributions, making it difficult to capture the inherent multimodality in humanoid robot tasks. For example, in dexterity tasks, a single target may have multiple feasible grasping postures, while bipedal walking involves various effective joint configurations to maintain balance.
[0005] Diffusion models, as an emerging generative framework, can learn more expressive policy distributions, but they face the following challenges in practical applications: In offline reinforcement learning, diffusion policies are limited to fixed datasets and struggle to adapt to new environments; in online learning, there is a lack of effective training methods and high computational overhead; and existing methods lack accurate means of evaluating policy distributions.
[0006] Existing diffusion strategies have the following shortcomings: it is difficult to balance the relationship between action optimization and behavior imitation when optimizing the strategy; gradient-based methods face problems of learning instability and bias accumulation; existing exploration mechanisms generally suffer from high computational overhead and low efficiency, and are difficult to directly guide policy updates; diffusion strategies lack a closed entropy expression, and cannot directly apply the maximum entropy principle for exploration. Summary of the Invention
[0007] The purpose of this invention is to provide a robot control method, system, device, and medium for scattered strategy step sampling to solve the above-mentioned problems in the prior art.
[0008] This invention is achieved through the following technical solution:
[0009] In a first aspect, a robot control method based on diffusion strategy skip sampling is characterized by comprising:
[0010] Initialize network parameters and hyperparameters, including policy network, double-Q network, and diffusion steps;
[0011] Perform state observation, calculate the Q-function gradient to estimate state sensitivity, and adaptively adjust the step parameters based on the magnitude of the gradient sensitivity;
[0012] The execution of actions involves interacting with the environment, including determining the sampling sequence based on the step parameters, performing DDIM sampling, obtaining actions from the policy network, and sending the actions to the simulation environment for interaction to obtain the next state and action reward.
[0013] Preferably, the policy network includes:
[0014]
[0015] In the formula, π θ For parameterized policy functions, For the final clean action generated, s t As the current state, p θ "Policy network parameters" represents the set of parameters used in the parameterized neural network to model the conditional probabilities of the diffusion process. The final action sample is derived from time step T to time step 0 during the diffusion process, and is the output of the entire diffusion sampling process. The distribution follows a standard normal distribution N(0,I). Let T be the initial noise, T be the number of diffusion steps, and k be the step index in the diffusion process. and The intermediate states at steps k-1 and k in the diffusion process represent intermediate results in the gradual denoising process from noise to clean action.
[0016] Preferably, the calculation of the Q-function gradient to estimate the state sensitivity includes:
[0017]
[0018] In the formula, S is the gradient norm of the Q function with respect to the action in the current state. Let a be the action gradient operator, Q() be the Q function, and a t This is the current action.
[0019] Preferably, the adaptive adjustment of the skip step parameters based on the magnitude of the state sensitivity estimated by the Q-function gradient includes:
[0020]
[0021] In the formula, S t To represent the adaptive skip step parameters at time step t, s max s is the maximum allowed skip step value.base Based on the jump coefficient, ∈=1e -8 .
[0022] Preferably, the DDIM sampling includes:
[0023]
[0024] In the formula, This is an intermediate action of ks in the diffusion model. The cumulative noise figure, For the predicted clean action at time t, σ k Let ξ be the noise standard deviation, and ξ be the random noise term.
[0025] Preferred options also include:
[0026] The actions, states, and rewards of each interaction are stored in the experience replay buffer. Several batches of data are randomly sampled from the experience replay buffer to train the Q network and the policy network.
[0027] By sampling several actions and calculating the relationship between them, the information potential of the policy distribution is obtained.
[0028]
[0029] In the formula, Let N be the information potential, and N be the number of samples. For Gaussian kernel function, For the action of the i-th sample at time i, Let j be the action of the j-th sample at time t;
[0030] Use information potential to evaluate the characteristics of policy distribution.
[0031] Preferably, the trained Q-network includes:
[0032]
[0033] In the formula, L is the overall loss function. For the expectation operator, Q ψ Let λ be the action value function, λ be the balance factor, and B be the experience replay buffer.
[0034] Secondly, the present invention also provides a robot control system based on diffusion strategy skip sampling, comprising:
[0035] The initialization module is configured to initialize network parameters and hyperparameters, including policy network, dual network, and diffusion steps.
[0036] The adaptive skip sampling module is configured to perform state observation, calculate the Q-function gradient to estimate state sensitivity, and adaptively adjust the skip parameters according to the magnitude of the gradient sensitivity.
[0037] The DDIM step sampling module is configured to perform actions and interact with the environment, including determining the sampling sequence based on the step parameters, performing DDIM sampling, obtaining actions from the policy network, and sending the actions to the simulation environment for interaction to obtain the next state and action reward.
[0038] The joint optimization module for training the Q network and the policy network error entropy-Q value is configured to store the action, state, and reward of each interaction step into the experience replay buffer, and randomly sample a number of batch data from the experience replay buffer to train the Q network and the policy network.
[0039] The main control module, together with the initialization module, adaptive skip sampling module, DDIM skip sampling module, and error entropy-Q value joint optimization module, is used to execute the above-mentioned robot control method based on diffusion strategy skip sampling.
[0040] Thirdly, an electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the above-described robot control method based on diffusion strategy skip sampling.
[0041] Fourthly, a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described robot control method based on diffusion strategy skip sampling.
[0042] The technical solution of the present invention has at least the following advantages and beneficial effects:
[0043] The method provided by this invention mainly includes initializing network parameters and hyperparameters, calculating the Q-function gradient to estimate state sensitivity based on the current state, adaptively adjusting the step parameters according to the magnitude of the state sensitivity estimated by the Q-function gradient, executing actions to interact with the environment, including determining the sampling sequence based on the step parameters and performing DDIM sampling to obtain actions from the policy network, and sending the actions to the simulation environment for interaction to obtain the next state and action reward. This method significantly reduces computational complexity and sampling time through an innovative deterministic diffusion process step sampling technique, while maintaining high-quality policy generation. This invention is particularly suitable for humanoid robot control systems, capable of efficiently handling complex high-dimensional action spaces and realizing real-time and flexible control strategies. Attached Figure Description
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0045] Figure 1 This is a schematic diagram of the process of the present invention;
[0046] Figure 2 This is an experimental effect diagram of the present invention;
[0047] Figure 3 This is the adaptive skip-step comparison process of the present invention. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0049] The module division in this application is a logical division. In actual application, there may be other division methods. For example, multiple modules may be combined into or integrated into another system, or some features may be ignored or not executed.
[0050] The independently described modules or sub-modules may or may not be physically separated; they may be implemented in software or hardware, and some modules or sub-modules may be implemented in software, with the processor calling the software to implement the function of these modules or sub-modules, while other modules or sub-modules may be implemented in hardware, such as through hardware circuits. Furthermore, some or all of the modules can be selected to achieve the purpose of this application's solution according to actual needs.
[0051] like Figure 1 and Figure 3 As shown, the present invention provides a robot control method based on diffusion strategy skip sampling, comprising:
[0052] S101: Initialize network parameters and hyperparameters, including policy network, dual-Q network, and diffusion steps;
[0053] The dual-Q network comprises two independent Q-networks: a main network responsible for selecting the current optimal action, and a target network used to evaluate the long-term value of the selected action. For example, when updating the Q-value, the main network selects an action, while the target network calculates the corresponding Q-value, avoiding overestimation that can occur when a single network performs both selection and evaluation simultaneously.
[0054] S102: Perform state observation, calculate the Q-function gradient to estimate state sensitivity, and adaptively adjust the step parameters according to the magnitude of the gradient sensitivity;
[0055] The system evaluates the sensitivity of the current state in the following ways: calculating the gradient of the Q function with respect to the action (a large gradient indicates higher sensitivity), analyzing the rate of change of the state (rapidly changing states require more precise control), and considering task characteristics, such as the need for more precise actions in balancing tasks under unstable states.
[0056] The core contribution of this invention is the skip sampling mechanism, which significantly reduces the computation time required for action generation. Compared with traditional diffusion models, the computational complexity is greatly reduced, and the action generation time is reduced from tens of milliseconds required by traditional methods to a level that meets the requirements of real-time control. The diffusion steps required for each decision are also reduced from tens of steps in traditional methods to single digits, while memory usage is also significantly reduced, enabling the algorithm to run on resource-constrained embedded systems.
[0057] These improvements enable the diffusion model to be truly applied to real-time control systems for the first time, especially in high-dimensional motion spaces such as humanoid robot control. The improved sampling efficiency has revolutionary significance for the practical application of the system.
[0058] In this context, state observation involves obtaining the robot's state vector st∈Rn at the current time t, where n is the dimension of the state space. The state vector includes, but is not limited to, physical quantities such as the robot's position coordinates, joint angles, angular velocity, and end effector pose. The obtained state vector st is standardized to scale the features of each dimension to the interval [-1, 1].
[0059] S103: Execute actions and interact with the environment, including determining the sampling sequence based on the step parameters, performing DDIM sampling, obtaining actions from the policy network, and sending the actions to the simulation environment for interaction to obtain the next state and action reward.
[0060] The actions, states, and rewards of each interaction are stored in the experience replay buffer. Several batches of data are randomly sampled from the experience replay buffer to train the Q network and the policy network.
[0061] The method provided by this invention mainly includes initializing network parameters and hyperparameters, including a policy network, a dual network, and a diffusion step count; performing state observation, calculating the Q-function gradient to estimate state sensitivity, and adaptively adjusting the skip step parameters based on the magnitude of the gradient sensitivity; determining the sampling sequence based on the skip step parameters and executing the actions obtained from DDIM sampling; executing the actions to interact with the environment and updating the Q-network. This method significantly reduces computational complexity and sampling time through an innovative deterministic diffusion process skip sampling technique, while maintaining high-quality policy generation. This invention is particularly suitable for humanoid robot control systems, capable of efficiently handling complex high-dimensional action spaces and realizing real-time and flexible control strategies.
[0062] Despite significantly reducing computational complexity, the skip sampling mechanism of this invention maintains high-quality generated actions. Actions generated through skip sampling are virtually indistinguishable in quality from those generated by the complete diffusion process. In high-precision tasks, skip-sampled actions can achieve the same control accuracy as standard diffusion. In some tasks, the action quality is even improved due to reduced accumulated errors.
[0063] The invention focuses on optimizing the Q-value maximization process, achieving significant results. Compared with traditional methods, the Q-value convergence speed is significantly improved, and the Q-value of the ultimately learned policy is also significantly enhanced. Through the use of double Q-network technology, the overestimation problem is effectively suppressed, the stability of Q-value evaluation is improved, and fluctuations are significantly reduced.
[0064] In real-time control scenarios, this invention offers significant advantages. The control frequency is greatly increased, and latency is significantly reduced, enabling the system to respond to environmental changes more quickly. In the balance control of humanoid robots, reaction time is drastically shortened, allowing the system to perform complex control tasks with limited computing resources.
[0065] The method of this invention exhibits good generalization ability, especially in dynamic environments. In unseen environments, the task success rate is significantly improved, and the resistance to external interference is enhanced. In transfer learning scenarios, the speed of adapting to new tasks is accelerated, robustness to parameter changes is enhanced, and overall stability is improved.
[0066] In one exemplary embodiment of the present invention, the core idea is to significantly improve the computational efficiency of the diffusion strategy through an innovative skip sampling technique, while optimizing the Q-value to achieve a high-performance control strategy. Firstly, regarding policy representation, the present invention employs an improved deterministic diffusion process to generate actions. At each time step t, the entire policy generation process can be represented as a Markov chain, and at each time step t, the policy generation process can be represented as:
[0067]
[0068] In the formula, π θ For parameterized policy functions, For the final clean action generated, s t As the current state, p θ "Policy network parameters" represents the set of parameters used in the parameterized neural network to model the conditional probabilities of the diffusion process. The final action sample is derived from time step T to time step 0 during the diffusion process, and is the output of the entire diffusion sampling process. The distribution follows a standard normal distribution N(0,I). Let T be the initial noise, T be the number of diffusion steps, and k be the step index in the diffusion process. and The intermediate states at steps k-1 and k in the diffusion process represent the intermediate results in the gradual denoising process from noise to clean action.
[0069] In one exemplary embodiment of the present invention, calculating the state sensitivity of the Q-function gradient estimation includes:
[0070]
[0071] In the formula, S is the gradient norm of the Q function with respect to the action in the current state. Here, Q is the action gradient operator, Q() is the Q function, and a t This is the current action.
[0072] Specifically, the adaptive adjustment of the skip step parameters based on the magnitude of the state sensitivity estimated by the Q-function gradient includes:
[0073]
[0074] In the formula, S t To represent the adaptive skip step parameters at time step t, s max s is the maximum allowed skip step value. base The basic jump coefficient, ∈ is a small constant, ∈=1e -8 To prevent division by zero errors.
[0075] This adaptive strategy can use large step values to improve efficiency when the task is simple or the state is stable, while using small step values to ensure the quality of the action when the task is complex or requires precise control.
[0076] In one exemplary embodiment of the present invention, DDIM sampling includes:
[0077]
[0078] In the formula, This is an intermediate action of ks in the diffusion model. The cumulative noise figure, For the predicted clean action at time t, σ k Let ξ be the noise standard deviation, and ξ be the random noise term.
[0079] This skip-sampling mechanism is a groundbreaking innovation of this invention. It allows the system to skip intermediate steps during the diffusion process, moving directly from step t to step ts, instead of the traditional step-by-step iteration. This method reduces computational complexity by approximately 70%, and the generation time for each action is reduced from tens of milliseconds to a few milliseconds, enabling the diffusion model to be truly applied to real-time control scenarios.
[0080] This invention further investigates the step sampling parameter s, finding that when s is between 2 and 4, computational efficiency can be maximized while maintaining action quality. Specifically, this invention proposes an adaptive step strategy that dynamically adjusts the step parameter s based on task complexity and state uncertainty.
[0081] An exemplary embodiment of the present invention further includes:
[0082] By sampling several actions and calculating the relationship between them, the information potential of the policy distribution is obtained.
[0083]
[0084] In the formula, Let N be the information potential, and N be the number of samples. For Gaussian kernel function, For the action of the i-th sample at time i, The action of the j-th sample at time t. This evaluation method does not require an analytical expression for the policy distribution, providing an effective approach for handling complex multimodal distributions;
[0085] Regarding the introduction of an information theory framework, using information potential to evaluate the properties of policy distributions, specifically including:
[0086] Calculate the mutual information between the strategy and the target variable (such as environmental feedback or payoff function) to screen for high-information potential strategies that have a significant impact on the target.
[0087] By introducing conditional mutual information, we can evaluate the independent contribution of a policy given the knowledge of other policies, thus avoiding information potential overlap caused by redundant policies.
[0088] Modeling dynamic information potential:
[0089]
[0090] In the formula, K(D) is the information structure of the policy distribution, K(D+ΔD) is the information structure of the updated policy distribution, and ΔD is the update amount.
[0091] By monitoring the update degree of ΔD during the policy iteration process, it is possible to determine whether the information potential has reached the threshold of qualitative change (i.e., the qualitative information potential dominates the structural reconstruction), and to evaluate the characteristics of the current policy distribution, i.e. whether it has caused the information potential to break through the critical value, thereby determining whether the overall reconstruction of the policy distribution has been triggered.
[0092] Information potential provides a dynamic and quantitative perspective for policy distribution evaluation. By quantifying distribution characteristics through entropy and mutual information, and combining this with an information field model to reveal evolutionary patterns, it ultimately guides policy optimization. This method has broad application potential in reinforcement learning, multi-agent systems, and other fields.
[0093] Based on the above policy representation and evaluation methods, this invention employs an innovative dual-objective optimization framework to train the policy. The optimization objective combines traditional Q-value maximization with an information-theoretic exploration mechanism. Updating the Q-network includes:
[0094]
[0095] In the formula, L is the overall loss function. For the expectation operator, Q ψ Let λ be the action value function, λ be the balance factor, and B be the experience replay buffer.
[0096] This optimization approach considers both the performance improvement of the strategy and ensures sufficient state space exploration by maximizing the information potential.
[0097] This invention also provides a robot control system based on diffusion strategy skip sampling, comprising:
[0098] An initialization module is configured to initialize network parameters and hyperparameters, including policy network, dual-Q network, and diffusion steps.
[0099] The adaptive skip sampling module is configured to perform state observation, calculate the Q-function gradient to estimate state sensitivity, and adaptively adjust the skip parameters according to the magnitude of the gradient sensitivity.
[0100] The sampling efficiency is intelligently adjusted based on the complexity of the current task and the uncertainty of the state, achieving an optimal balance between computational efficiency and action quality. The number of hops is reduced to improve accuracy at critical decision-making moments or in complex states, while the number of hops is increased to improve efficiency in stable states or simple tasks. This dynamic adjustment mechanism can achieve optimal performance-efficiency balance in different scenarios.
[0101] The DDIM step sampling module is configured to perform actions and interact with the environment, including determining the sampling sequence based on the step parameters, performing DDIM sampling, obtaining actions from the policy network, and sending the actions to the simulation environment for interaction to obtain the next state and action reward.
[0102] This module achieves efficient action generation through an improved deterministic diffusion process:
[0103] It employs a deterministic back-diffusion process; implements innovative skip sampling to significantly improve computational efficiency (reducing computation time by approximately 70%); optimizes sampling quality and speed by dynamically adjusting skip parameters; and ensures the generation of high-quality actions with a small number of diffusion steps.
[0104] The error entropy-Q value joint optimization module is configured to train the Q network and the policy network. The error entropy-Q value joint optimization module is configured to store the action, state and reward of each interaction into the experience replay buffer D, and randomly sample a number of batch data from the experience replay buffer D to train the Q network and the policy network.
[0105] This module innovatively integrates error entropy and Q-value optimization:
[0106] We use Renyi entropy as the information theory-based measure of strategy distribution characteristics; we design a joint objective function to simultaneously optimize error entropy and Q-value; we dynamically balance exploration (error entropy) and utilization (Q-value); we use double-Q network technology to reduce overestimation bias; and we guide the direction of Q-value optimization through error entropy to achieve a more stable learning process.
[0107] The core innovation of this invention lies in the organic combination of computational efficiency (skip sampling) and learning performance (joint optimization of error entropy and Q-value), creating a diffusion reinforcement learning paradigm that is both efficient and high-performance. In particular, the fusion optimization framework of error entropy and Q-value solves the problem of traditional diffusion strategies struggling to balance exploration and utilization, while the skip sampling technique overcomes the computational bottleneck of diffusion models in real-time control.
[0108] The main control module, together with the initialization module, adaptive skip sampling module, DDIM skip sampling module, and error entropy-Q value joint optimization module, is used to execute the above-mentioned robot control method based on diffusion strategy skip sampling.
[0109] In this invention, a specific example is provided to further explain the above solution:
[0110] To verify the effectiveness of this invention, it was systematically tested on the Humanoid task in the MuJoCo physics simulation environment. This task involves a humanoid robot model with 17 joints. The state space includes multi-dimensional information such as joint angles and angular velocities, body posture and position, contact force information, and center of mass velocity. The motion space corresponds to the torque control of the 17 joints.
[0111] State observation and sensitivity assessment, calculation of Q-function gradient The system employs adaptive step calculation and DDIM step sampling execution. During the calculation process, it calculates the Q-function gradient based on state observations to assess state sensitivity and determine adaptive step parameters. Then, it generates actions based on the DDIM policy representation module. For the robot's state st at time t, the system generates actions through a diffusion process. Initial noise is sampled from a standard Gaussian distribution and undergoes a deterministic back-diffusion process to gradually generate precise joint control signals. In this process, the step sampling mechanism significantly improves the efficiency of action generation.
[0112] In the motion generation phase, the deterministic diffusion process employed in this invention is particularly well-suited for handling the high-dimensional motion space of humanoid robots. Through a carefully designed noise scheduling mechanism, the system can generate high-quality motions within a relatively small number of diffusion steps (10 steps), which is crucial for humanoid robot tasks requiring real-time control.
[0113] In terms of policy evaluation, this invention innovatively introduces an information theory framework. For each state, the system samples 64 action samples, and evaluates the characteristics of the policy distribution by calculating the information potential between these samples. This method is particularly suitable for evaluating the multimodal action distribution of humanoid robots, such as when multiple effective action choices may exist during balance recovery.
[0114] During the optimization phase, this invention employs a dual-objective optimization framework. On one hand, it maximizes the Q-value of actions to improve task performance; on the other hand, it maintains policy diversity by maximizing the information potential. This optimization approach enables the system to learn a motion control policy that is both stable and flexible.
[0115] In actual training, the parameters of this invention are configured as follows:
[0116] The experience replay buffer size is set to 1 million to store rich motion experience, the batch size is 256 to make full use of historical data with each update, the neural network adopts a three-layer structure (400-300-output dimension), the learning rate is set to 3e-4, and the Adam optimizer is used.
[0117] During the training process, the systematic learning process can be divided into the following stages:
[0118] Phase 1 (Initial Exploration):
[0119] At the beginning, the system mainly relies on the guidance of information potential for exploration. It uses DDIM to generate diverse actions to explore the state space and accumulate basic motion experience data.
[0120] Phase Two (Skill Acquisition):
[0121] The system begins to learn basic balance and movement skills, the influence of Q-value begins to increase, the strategy gradually focuses on high-reward areas, and at the same time maintains sufficient exploration through information potential.
[0122] Phase 3 (Strategy Optimization):
[0123] The system further optimizes the acquired motor skills, improves the quality of motion generation by dynamically adjusting the number of diffusion steps, balances exploration and utilization, and achieves stable performance improvement.
[0124] After 3 million steps of training, this invention demonstrates significant advantages in the following aspects:
[0125] Regarding strategy expression:
[0126] It successfully learns diverse movement patterns, generates precise joint control signals, and effectively handles the challenges of high-dimensional motion spaces.
[0127] In terms of learning efficiency:
[0128] Compared to the SAC algorithm, it converges faster; compared to the TD3 algorithm, it has higher sample utilization efficiency and a more stable training process.
[0129] These experimental results fully demonstrate the effectiveness of the information-theoretic-based diffusion reinforcement learning method proposed in this invention for handling complex humanoid robot control tasks. In particular, this invention exhibits significant technical advantages in areas such as exploration of high-dimensional action spaces, learning of multimodal policies, and optimization of motion performance. Figure 2 The training curve is shown.
[0130] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0131] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. This computer software product, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0132] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A robot control method based on diffusion strategy skip sampling, characterized in that, include: Initialize network parameters and hyperparameters, including policy network, double-Q network, and diffusion steps; The state sensitivity is estimated by calculating the gradient of the Q-function based on the current state, and the step parameters are adaptively adjusted according to the magnitude of the state sensitivity estimated by the gradient of the Q-function. The execution of actions involves interacting with the environment, including determining the sampling sequence based on the step parameters, performing DDIM sampling, and obtaining actions from the policy network. The action is fed into the simulation environment for interaction, and the next state and action reward are obtained. The policy network includes: In the formula, For parameterized policy functions, For the final clean action, This is the current state. "Policy network parameters" represents the set of parameters used in the parameterized neural network to model the conditional probabilities of the diffusion process. The final action sample is derived from time step T to time step 0 during the diffusion process, and is the output of the entire diffusion sampling process. The distribution is a standard normal distribution N(0,I). Initial noise, For diffusion steps, An index of the steps in the diffusion process. and The intermediate states at steps k-1 and k in the diffusion process represent intermediate results in the gradual denoising process from noise to clean action.
2. The robot control method based on diffusion strategy skip sampling according to claim 1, characterized in that, The calculation of the Q-function gradient estimation state sensitivity includes: In the formula, Let Q be the gradient norm of the Q-function with respect to the action in the current state. For the action gradient operator, For the Q function, This is the current action.
3. The robot control method based on diffusion strategy skip sampling according to claim 2, characterized in that, The adaptive adjustment of the step parameters based on the magnitude of the state sensitivity estimated by the Q-function gradient includes: In the formula, To represent the adaptive skip step parameters at time step t, This is the maximum allowed number of skip steps. Based on the jump coefficient, .
4. The robot control method based on diffusion strategy skip sampling according to claim 3, characterized in that, The DDIM sampling includes: In the formula, This is an intermediate action of ks in the diffusion model. This is the cumulative noise figure. For the predicted clean action at time t, The standard deviation of noise. This is a random noise term.
5. The robot control method based on diffusion strategy skip sampling according to claim 4, characterized in that, Also includes: The actions, states, and rewards of each interaction are stored in the experience replay buffer. Several batches of data are randomly sampled from the experience replay buffer to train the Q network and the policy network. By sampling several actions and calculating the relationship between them, the information potential of the policy distribution is obtained, and the characteristics of the policy distribution are evaluated using the information potential. Where, Let N be the information potential, and N be the number of samples. For Gaussian kernel function, For the action of the i-th sample at time t, Let be the action of the j-th sample at time t.
6. The robot control method based on diffusion strategy skip sampling according to claim 5, characterized in that, The trained Q-network includes: Where, For the overall loss function, For the expectation operator, For action value functions, As a balance factor, This serves as a buffer for experience replay.
7. A robot control system based on a diffusion strategy for skip-step sampling, characterized in that, include: The initialization module is configured to initialize network parameters and hyperparameters, including policy network, dual network, and diffusion steps. The adaptive skip sampling module is configured to perform state observation, calculate the Q-function gradient to estimate state sensitivity, and adaptively adjust the skip parameters according to the magnitude of the gradient sensitivity. The DDIM step sampling module is configured to perform actions and interact with the environment, including determining the sampling sequence based on the step parameters, performing DDIM sampling, obtaining actions from the policy network, and sending the actions to the simulation environment for interaction to obtain the next state and action reward. The joint optimization module for training the Q network and the policy network error entropy-Q value is configured to store the action, state, and reward of each interaction step into the experience replay buffer D, and randomly sample a number of batch data from the experience replay buffer D to train the Q network and the policy network. The main control module, together with the initialization module, the adaptive skip sampling module, the DDIM skip sampling module, and the error entropy-Q value joint optimization module, is used to execute the robot control method based on diffusion strategy skip sampling as described in any one of claims 1-6.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the robot control method based on diffusion strategy skip sampling as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when executed by a processor, the computer program implements the robot control method based on diffusion strategy skip sampling as described in any one of claims 1-6.
Citation Information
Patent Citations
Optimal control method of gait of humanoid robot based on deep Q network
CN110764416A