A robot control method based on conservative model reinforcement learning

By building multiple real environment estimation models and control policy networks, and using conservative environment estimation models to generate simulation data, the problem of avoiding high uncertainty areas in traditional methods is solved, and the stable optimization of robot control strategies is achieved.

CN119260713BActive Publication Date: 2025-07-22UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411416100.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-11
Publication Date
2025-07-22
Estimated Expiration
2044-10-11

AI Technical Summary

Technical Problem

Traditional model-based reinforcement learning methods fail to effectively avoid high uncertainty areas, resulting in the generation of low-quality samples and hinder the learning of the optimal control strategy network, especially in multi-step prediction, the extrapolation error is serious.

Method used

Build multiple real environment estimation models and control policy networks, generate simulation data through the interaction between conservative environment estimation models and real data, optimize Q-value and control policy networks, ensure that conservative approximation estimation models are selected in each learning step to punish overestimation predictions, and achieve balance of model reinforcement learning.

Benefits of technology

Through conservative model reinforcement learning methods, high-quality multi-step model simulation samples are generated to improve sample efficiency, ensure the stability and optimization of the control policy network, and are suitable for robot control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119260713B_ABST
    Figure CN119260713B_ABST
Patent Text Reader

Abstract

The invention discloses a robot control method based on conservative model reinforcement learning, which relates to the technical field of machine learning. The robot control method based on conservative model reinforcement learning randomly selects an estimation model with conservative approximation from an integrated probability model in each model learning step. It appears in the form of a set of probability estimation models, but includes a mechanism for penalizing overestimation or overly optimistic predictions. This ensures a balance between conservativeness and generalization of the model-based reinforcement learning algorithm, thereby solving the problem that multi-step model simulation samples generated in the simulated environment in model-based reinforcement learning deviate seriously from the real environment data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and particularly to a robot control method based on conservative model reinforcement learning. Background Art

[0002] Traditional model-based reinforcement learning methods are limited to passively exploiting low-uncertainty model simulation samples by quantifying model uncertainty or using conservative approximations. Therefore, they fail to actively prevent trajectories induced by the current policy from entering high-uncertainty regions. This inability to avoid high-uncertainty regions leads to the generation of low-quality samples, thereby providing unstable learning objectives, which hinders the learning of the optimal control policy network. This problem becomes more prominent when using an estimated model for multi-step prediction. The reason for the problem is that the control policy network even makes suboptimal decisions using small model errors, which in turn prompts the estimated model to generate more extrapolation errors in multi-step prediction. Therefore, model-based reinforcement learning algorithms still face the challenge of generating high-quality multi-step model simulation samples to improve sample efficiency. Summary of the Invention

[0003] In order to at least overcome the above deficiencies in the prior art, the purpose of this application is to provide a robot control method based on conservative model reinforcement learning.

[0004] This application provides a robot control method based on conservative model reinforcement learning, and the method includes:

[0005] Step 1: Construct multiple real environment estimation models, multiple Q-values, a control policy network, a real data buffer pool, and a simulated data buffer pool;

[0006] Step 2: When the robot interacts with the real environment according to the control policy network, when a state transition occurs after executing an action, store the interaction trajectory of the state transition into the real data buffer pool;

[0007] Step 3: Construct an optimization objective for the conservative environment estimation model through the multiple real environment estimation models;

[0008] Step 4: Optimize the conservative environment estimation model using the data in the real data buffer pool;

[0009] Step 5: Generate prediction data through multi-step interaction trajectory prediction between the control policy network and the conservative environment estimation model, and store the obtained data into the simulated data buffer pool;

[0010] Step 6: Optimize the Q-value and the control policy network using the data in the simulated data buffer pool;

[0011] Step 7: Continuously iterate and optimize the conservative environment estimation model, Q-value, and control policy network until the performance of the current control policy network meets the expected requirements;

[0012] Step 8: Control the robot's movement according to the final control policy network.

[0013] Furthermore, the specific method of step 1 is as follows:

[0014] Use multiple probabilistic neural networks to represent the real environment estimation model, that is where N represents the number of probabilistic estimation models, and each probabilistic neural network models the transition probability density as a Gaussian model, whose mean and diagonal covariance are given by the neural network:

[0015]

[0016] where θ represents the optimizable parameters of the neural network, s represents the current state, a represents the action executed in state s, s′ and r represent the state reached after the environment transition occurs after executing the action and the reward value received by the agent, and represent the mean and diagonal covariance of the nth estimation model, represents the normal distribution with mean and covariance ;

[0017] The Q-value function is represented by multiple fully connected neural networks, denoted as Q i , the control policy network consists of multiple layers of fully connected neural networks, and the real data buffer pool and the simulated data buffer pool are storage areas in memory.

[0018] Furthermore, the trajectory data obtained after the robot interacts with the real environment according to the control policy network in step 2 includes:

[0019] Input the current state data s of the robot in the real environment into the control policy network and receive the action data a output by the control policy network; the control policy network samples the mean and variance of the multi-dimensional Gaussian distribution output by the last fully connected layer of the control policy network;

[0020] Control the robot to run in the real environment with the action data a, and obtain the state data s′ of the robot at the next moment and the reward value r after executing the action data a;

[0021] Record the state data, action data, state data at the next moment, and reward value as (s, a, s′, r) and use it as the trajectory data.

[0022] Furthermore, the specific method of step 3 is as follows:

[0023] Randomly select 2 real environment estimation models from multiple real environment estimation models and Calculate the average value μ of the means of the selected real environment estimation models z and the average value σ of the variances z , that is, the following formula:

[0024]

[0025]

[0026] According to the calculated μ z and σ z Construct the upper bound μ u and the lower bound μ l of the selected real environment estimation models:

[0027] μ u = μ z + σ z

[0028] μ l = μ z - σ z

[0029] Then, let the upper bound μ u of the selected real environment estimation models approach the lower bound μ l :

[0030]

[0031] where μ l does not propagate gradient information during optimization;

[0032] Obtain the optimization objective of the conservative environment estimation model according to the following formula

[0033]

[0034] wherein, represents the expected value of sampling interaction trajectory samples from the real buffer, represents the loss of optimizing the estimation model using the maximum likelihood function, s t+1 represents the state at the next moment in the real environment, s t+1 contains the return value r, η is an adjustable hyperparameter, represents the squared value of the covariance of the estimation model.

[0035] Furthermore, the specific method of step 5 is:

[0036] Step 5.1: Randomly sample state data and action data in the real data buffer pool, denoted as (s, a);

[0037] Step 5.2: Input the sampled (s, a) into the conservative environment estimation model to generate the simulated state at the next moment and the return value

[0038] Step 5.3: Input the simulated state at the next moment generated by the conservative environment estimation model into the control policy network and receive the action data a′ output by the control policy network;

[0039] Step 5.4: Input the data into the conservative environment estimation model again to generate subsequent model simulation samples;

[0040] Step 5.5: Store the data as trajectory data in the simulated data buffer pool.

[0041] Furthermore, the specific method of Step 6 is as follows:

[0042] Optimize the Q value using the data in the simulated data buffer pool, and the optimization objective is

[0043]

[0044] In the formula, φ i represents the optimizable parameter corresponding to the Q value function, represents the simulated data buffer pool, represents the expected value of sampling samples from the simulated data buffer, represents the state value predicted by the estimation model, i represents the index value of the Q value function set, represents the Q value function with index value i and parameter φ, represents the return value predicted by the estimation model, Q i (.) represents the Q value with index value i for state-action evaluation, and Q′1(.), Q′2(.) respectively represent the corresponding Q values at the next moment;

[0045] The update of the control policy network is completed by maximizing the Q value, that is, the following formula:

[0046]

[0047] In the formula, represents the loss function of the policy function of, represents the control policy network function, are the optimizable parameters of the control policy network function.

[0048] The robot control method based on conservative model reinforcement learning according to the present invention randomly selects an estimation model with conservative approximation from the integrated probability model in each model learning step. It appears in the form of a set of probability estimation models, but includes a mechanism for penalizing overestimation or overly optimistic predictions. This ensures a balance between conservativeness and generalization of the model-based reinforcement learning algorithm, and further solves the problem that the multi-step model simulation samples generated in the simulation environment in model-based reinforcement learning deviate seriously from the real environment data. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is the overall learning schematic diagram of the robot control method based on conservative model reinforcement learning;

[0050] Figure 2 is the schematic diagram of the method steps of the embodiment of the present application;

[0051] Figure 3 is the optimization process of the conservative estimation model of the embodiment of the present application;

[0052] Figure 4 is the learning process of the Q-value function and the control policy network of the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] Please refer to Figure 2 , which is the flow schematic diagram of the robot control method based on conservative model reinforcement learning provided by the embodiment of the present invention. The robot control method based on conservative model reinforcement learning may specifically include the content described in the following steps S1 - step S8.

[0054] S1: Construct multiple real environment estimation models, multiple Q-values, a control policy network, a real data buffer pool, and a simulation data buffer pool;

[0055] S2: When the control policy network interacts with the real environment, after executing an action, perform a state transition, and record the interaction trajectory of the state transition and store it in the real data buffer pool;

[0056] S3: Construct an optimization objective for conservative estimation through the multiple real environment estimation models;

[0057] S4: Optimize the conservative environment estimation model using the data in the real data buffer pool;

[0058] S5: Generate prediction data through multi-step interaction trajectory prediction between the control policy network and the conservative environment estimation model, and store the obtained data in the simulation data buffer pool;

[0059] S6: Optimize the Q value and the control policy network with the data in the simulated data buffer pool;

[0060] S7: Continuously iteratively optimize the conservative environment estimation model, the Q value, and the control policy network until the performance of the current control policy network meets the expected requirements;

[0061] S8: Control the movement of the robot according to the final control policy network.

[0062] In existing model - based reinforcement learning techniques, there are post - processing of model simulation samples and pessimistic policy updates. In the method of post - processing model simulation samples, one technique involves selecting low - uncertainty interaction samples according to the minimum uncertainty score. They explicitly quantify the uncertainty of the Q value by measuring the difference between set - probability estimation models. Another technique is to correct these model simulation samples with real samples. In the method of pessimistic update policy, they select model simulation samples from the conservative approximation of the set - model Q value for policy optimization. However, these objectives mainly focus on minimizing one - step model differences, while empirical observations show that using multi - step model simulation samples deviates significantly from real - world environment data, and these objectives may lead to poor asymptotic performance. This includes using the low - uncertainty regions where the model is accurate while avoiding the high - uncertainty regions where the model is inaccurate, so as to achieve reliable decision - making based on the estimated model. The estimated model may generate model simulation samples with extrapolation errors in high - uncertainty regions, resulting in unstable or subsequent collapse of sub - optimal policies developed through model development. Although there are inherent inaccuracies in these regions, seeking the optimal policy often requires venturing into these regions. Overly conservative exploration, mainly sticking to low - uncertainty regions, will hinder the discovery of superior policies. Therefore, it is crucial to balance conservatism and generalization in model - based reinforcement learning.

[0063] In the embodiments of this application, please refer to Figure 1 , every iterative update of the conservative estimation model requires controlling the robot to run in the real environment using the policy in the policy model, and the generated trajectory data is used to update the conservative estimation model; then, the conservative estimation model and the control policy network perform multi - step interactive trajectory prediction to generate a large number of model simulation samples as training samples to update the control policy network model, so as to ensure that after iteration, the control policy network can become the policy algorithm for robot control for subsequent testing. The embodiments of this application can provide a conservative estimation model for real - robot interaction, which is learned through real - historical interaction trajectories. The loss function of the conservative estimation model in the learning process has a direct impact on the conservatism and generalization of the estimated environment. The conservative estimation model of the embodiments of this application corresponds to the simulated environment, providing safety and efficiency for robot interaction.

[0064] In a possible implementation, the robot control method based on conservative model reinforcement learning includes constructing multiple real environment estimation models, multiple Q-values, a control policy network, a real data buffer pool, and a simulated data buffer pool, including:

[0065] Using multiple probabilistic neural networks to represent the real environment estimation models, that is where n represents the number of probability estimation models, and each probabilistic neural network models the transition probability density as a Gaussian model, whose mean and diagonal covariance are given by the neural network:

[0066]

[0067] where θ represents the optimizable parameters of the neural network, s represents the current state, a represents the action executed in state s, s′ and r represent the state reached after the environment transition occurs after executing the action and the reward value received by the agent, and represent the mean and diagonal covariance of the nth estimation model;

[0068] The Q-value function is represented by using multiple fully connected neural networks as Q i , the control policy network is obtained from a multi-layer fully connected neural network, and the real data buffer pool and the simulated data buffer pool are storage areas in the memory.

[0069] In a possible implementation, obtaining trajectory data by interacting the control policy network with the real environment includes:

[0070] Inputting the current state data s of the robot in the real environment into the control policy network, and receiving the action value a output by the control policy network; the control policy network samples the mean and variance of the multi-dimensional Gaussian distribution output by the last fully connected layer of the control policy network;

[0071] Controlling the robot to run in the real environment with the action data, and obtaining the state data s′ of the robot at the next moment and the reward value r after executing the action data;

[0072] Taking the state data, action data, state data at the next moment, and reward value (s, a, s′, r) as the trajectory data.

[0073] In a possible implementation, constructing a conservative estimation optimization target through the multiple real environment estimation models includes:

[0074] Randomly selecting 2 estimation models from multiple estimation models and Calculate the average of the mean and variance of the selected estimation model, i.e., the following formula:

[0075]

[0076]

[0077] According to the calculated μ z and σ z Construct the upper bound μ u and lower bound μ l of the selected estimation model:

[0078] μ u = μ z + σ z

[0079] μ l = μ z - σ z

[0080] Then, let the upper bound μ u of the selected estimation model approach the lower bound μ l :

[0081]

[0082] where μ l does not propagate gradient information during optimization.

[0083] In a possible implementation, optimize the conservative environment estimation model using the data in the real data buffer pool, and obtain the optimization objective of the conservative environment estimation model according to the following formula:

[0084]

[0085] where s t+1 contains the return value r, and η is an adjustable hyperparameter.

[0086] In a possible implementation, generate prediction data through multi-step interactive trajectory prediction between the control policy network and the conservative environment estimation model, and store the obtained data in the simulated data buffer pool, including:

[0087] Randomly sample the state and action pair (s, a) from the real data buffer pool;

[0088] Input the sampled (s, a) into the conservative environment estimation model to generate the simulated state at the next moment and the return value

[0089] Input the simulated state at the next moment generated by the model Input it into the control policy network and receive the action value a' output by the control policy network;

[0090] Input the new data into the conservative environment estimation model again to generate subsequent model simulation samples;

[0091] Take the data and store it as the trajectory data in the simulation data buffer pool.

[0092] In a possible implementation manner, optimizing the Q value and the control policy network through the data in the simulation data buffer pool includes:

[0093] Use the data in the simulation data buffer pool to optimize the Q value, using the optimization objective:

[0094]

[0095] where φ i represents the optimizable parameter of the Q value function, represents the simulation data buffer pool,

[0096] The update of the control policy network is completed by maximizing the Q value, that is, the following formula:

[0097]

[0098] where represents the control policy network function, is the optimizable parameter of the control policy network function.

[0099] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0100] The robot control method based on conservative model reinforcement learning of the present invention randomly selects an estimation model with a conservative approximation from the integrated probability model in each model learning step. It appears in the form of a set of probability estimation models, but includes a mechanism for penalizing overestimation or overly optimistic predictions. This ensures a balance between conservatism and generalization of the model-based reinforcement learning algorithm, and further solves the problem that the multi-step model simulation samples generated by the simulation environment in the model-based reinforcement learning deviate seriously from the real environment data.

[0101] For example, please refer to Figure 3 , which shows the specific scheme of the optimization process of the conservative estimation model in the embodiment of the present application, where the conservative estimation model It includes 7 integrated neural network models. Each neural network model is sequentially provided with 4 layers of fully connected layers and 3 activation functions. Each layer of fully connected layers is provided with 200 hidden layers, and the activation function used is Swish.

[0102] For examples, please refer to Figure 4 , which shows a specific solution for the learning process of the Q-value function and the control policy network of the embodiments of the present application. Among them, the Q-value network is sequentially provided with 4 layers of fully connected layers and 3 activation functions. Each layer of fully connected layers is provided with 256 hidden layers, and the activation function used is ReLU; the control policy network is sequentially provided with 4 layers of fully connected layers and 3 activation functions. Each layer of fully connected layers is provided with 256 hidden layers, and the activation function used is ReLU.

[0103] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0104] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices or units, and can also be electrical, mechanical or other forms of connection.

[0105] The units described as separate components may or may not be physically separated. Obviously, those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0106] In addition, each functional unit in various embodiments of the present invention may be integrated into one processing unit, may exist physically as individual units, or two or more units may be integrated into one unit. The above integrated units may be implemented in the form of hardware or in the form of software functional units.

[0107] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a grid device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0108] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A robot control method based on conservative model reinforcement learning, characterized in that, The method includes: Step 1: Construct multiple real - environment estimation models, multiple Q - values, a control policy network, a real - data buffer pool, and a simulated - data buffer pool; Step 2: When the robot interacts with the real environment according to the control policy network, during the state transition after executing an action, store the interaction trajectory of the state transition into the real - data buffer pool; Step 3: Construct the optimization objective of the conservative environment estimation model through the multiple real - environment estimation models; Randomly select 2 real environment estimation models from multiple real environment estimation models and used to calculate the average value μ of the means of the selected real environment estimation models z and the average value σ of the variances z , that is, the following formula: Among them, and represent the mean and diagonal covariance of the i-th estimation model, s represents the current state, a represents the action executed in state s. Based on the calculated μ z and σ z construct the upper bound μ u and the lower bound μ l : μ u = μ z + σ z μ l = μ z - σ z Then, let the upper bound μ of the selected true environment estimation model u approach the lower bound μ l closer: where μ l does not propagate gradient information during optimization; Obtain the optimization objective of the conservative environment estimation model according to the following formula In the formula, represents the expected value of sampling interaction trajectory samples from the real buffer, represents the loss of optimizing the estimation model using the maximum likelihood function, s t+1 represents the state at the next moment in the real environment, s t+1 contains the return value r, and η is an adjustable hyperparameter, represents the square value of the covariance of the estimation model; Step 4: Optimize the conservative environment estimation model using the data in the real - data buffer pool; Step 5: Generate prediction data through multi - step interaction trajectory prediction between the control policy network and the conservative environment estimation model, and store the obtained data into the simulated - data buffer pool; Step 6: Optimize the Q - values and the control policy network using the data in the simulated - data buffer pool; Step 7: Continuously iteratively optimize the conservative environment estimation model, Q - values, and the control policy network until the performance of the current control policy network meets the expected requirements; Step 8: Control the robot's movement according to the final control policy network.

2. A robot control method based on conservative model reinforcement learning as claimed in claim 1, characterized in that, The specific method of Step 1 is: Use multiple probabilistic neural networks to represent the true environment estimation model, that is where N represents the number of probabilistic estimation models, and each probabilistic neural network models the transition probability density as a Gaussian model, whose mean and diagonal covariance are given by the neural network: where θ represents the optimizable parameters of the neural network, s represents the current state, a represents the action executed in state s, s ′ and r represent the state reached after the environmental transition occurs after the execution of the action and the reward value received by the agent, and represent the mean and diagonal covariance of the n-th estimation model, represents a normal distribution with mean and covariance ; The Q-value function is represented by multiple fully-connected neural networks as Q i , the control policy network consists of multiple layers of fully-connected neural networks, and the real data buffer pool and the simulated data buffer pool are storage areas in the memory.

3. A robot control method based on conservative model reinforcement learning as claimed in claim 1, characterized in that, The trajectory data obtained when the robot interacts with the real environment according to the control policy network in Step 2 includes: Input the current state data s of the robot in the real environment into the control policy network and receive the action data a output by the control policy network; the control policy network samples the mean and variance of the multi - dimensional Gaussian distribution output by the last fully - connected layer of the control policy network; Control the robot to run in the real environment with action data a, and obtain the state data s of the robot at the next moment ′ and the return value r obtained after executing the action data a; Record the state data, action data, next - moment state data, and reward value as (s, a, s′, r) and use it as the trajectory data.

4. A robot control method based on conservative model reinforcement learning as claimed in claim 1, characterized in that, The specific method of Step 5 is: Step 5.1: Randomly sample state data and action data in the real - data buffer pool and record them as (s, a); Step 5.2: Input the sampled (s,a) into the conservative environment estimation model to generate the simulated state at the next moment and the return value Step 5.3: Input the simulated state at the next moment generated by the conservative environment estimation model into the control policy network and receive the action data a′ output by the control policy network; Step 5.4: Input the data into the conservative environment estimation model again to generate subsequent model simulation samples; Step 5.5: Store the data as trajectory data in the simulated data buffer pool.

5. A robot control method based on conservative model reinforcement learning as shown in claim 1, characterized in that, The specific method of Step 6 is: Optimizing the Q value using the data in the analog data buffer pool, with the optimization objective being Where φ i represents the optimizable parameter corresponding to the Q-value function, represents the simulated data buffer pool, represents the simulated state at the next moment, a′ represents the action data, and a represents the executed action data, represents the expected value of sampling samples from the simulated data buffer, represents the state value predicted by the estimation model, and i represents the index value of the Q-value function set, represents the Q-value function with index value i and parameter φ, represents the return value predicted by the estimation model, Q i (.) represents the Q-value with index value i for state-action evaluation, Q ′ 1(.), Q ′ 2(.) respectively represent the corresponding Q-values at the next moment; The update of the control policy network is completed by maximizing the Q - value, that is, the following formula: In the formula, represents the loss function of the policy function π θ , where π θ represents the control policy network function, and θ is the optimizable parameter of the control policy network function.

Citation Information

Patent Citations

  • Reinforcement learning robot control method based on consistency constraint modeling and system thereof

    CN113485107A

  • Unmanned equipment control method for model-based high-sample-rate deep reinforcement learning

    CN115293334A