Re-parameterization reinforcement learning system and learning method based on beta probability distribution

The mean and deviation parameters of beta distribution are output through the neural network, and the problems of unboundary nature of Gaussian strategy and the correlation between the shape parameters of traditional beta strategy in reinforcement learning are solved, achieving more efficient sample utilization and learning efficiency.

CN120068986APending Publication Date: 2025-05-30AERONAUTICS RES INST OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411966793.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

During the action sampling process of existing reinforcement learning algorithms, the unbounded nature of Gaussian strategy leads to bias in policy gradient calculations, while the shape parameter correlation of traditional beta strategy makes policy gradient optimization complex.

Method used

The mean and deviation parameters of the beta distribution are output by neural networks. By analyzing the relationship between the mean and shape parameters of the beta distribution, the deviation parameters are calculated in segments, which solves the problem of deviation output of beta distribution.

Benefits of technology

It realizes more efficient sample utilization in on-orbit and off-orbit reinforcement learning, optimizes the strategy gradient calculation and learning process, improves learning efficiency and obtains higher returns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068986A_ABST
    Figure CN120068986A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence, and relates to a re-parameterization reinforcement learning system and learning method based on beta probability distribution. According to the invention, environment state parameters are input into a neural network through a neural network subsystem; and the beta distribution parameter calculation subsystem obtains the mean value and the deviation parameter of beta distribution after being processed by the full connection layer and the nonlinear transformation layer, and calculates the shape parameter of beta distribution based on the mean value and the deviation parameter of beta distribution by the beta distribution parameter calculation subsystem. And the beta distribution sampling subsystem constructs beta probability distribution according to the shape parameters of beta distribution and samples actions according to the obtained beta probability distribution. The environment subsystem is trained through reinforcement learning, return is obtained after interaction with the environment, neural network parameters in the neural network subsystem are updated by adopting an optimization algorithm, and a better strategy is obtained through exploration. The method can be used for on-orbit and off-orbit reinforcement learning, the sample utilization efficiency is improved, and a higher return is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and relates to a reparameterized reinforcement learning system and learning method based on beta probability distribution. Background Art

[0002] Reinforcement learning obtains good policies through the interaction and learning between an agent and an environment. The policies obtained by the agent need to meet many requirements. First of all, reinforcement learning should be able to complete targeted exploration to avoid the agent learning high-reward policies with potential safety hazards; secondly, various costs are incurred during the interaction process between reinforcement learning and the environment, and the sample efficiency of reinforcement learning has a direct impact on the learning cost. Deep reinforcement learning using neural networks has been successfully used to solve various complex control decision-making tasks. However, data collection in the real world may be difficult and expensive, so reinforcement learning must effectively utilize limited data. Combining these requirements is not easy because stability and sample efficiency often conflict with each other.

[0003] To address the above problems, the present invention starts from the probability distribution used in the action sampling of the reinforcement learning agent. By analyzing the Gaussian distribution and the beta distribution, it can be found that there are advantages and disadvantages in applying these two distributions to action sampling. Further analyzing the advantages and defects of directly outputting the two shape parameters of the beta distribution using a neural network, a method of using a neural network to output the mean and deviation parameters of the beta distribution is proposed. In a specific implementation, the characteristics of the beta distribution are analyzed, and by bisecting the value range of the beta distribution, a reliable way to calculate the shape parameters of the beta distribution is obtained.

[0004] In the process of reinforcement learning to solve stochastic continuous control problems, the Gaussian distribution is generally used to represent its action distribution, and function approximators such as deep neural networks are used to calculate its mean and variance. The strategy of using a neural network to output Gaussian distribution parameters is called the Gaussian strategy. In the process of using the Gaussian strategy, the agent calculates the gradient of the strategy with respect to the mean and deviation of the Gaussian distribution, and optimizes the network through backpropagation and mini-batch stochastic gradient changes to obtain a reasonable action output. However, the Gaussian probability distribution is an unbounded function, and the range of sample actions is not restricted during the sampling process, which is inconsistent with the control actions in actual control problems. The action spaces of the vast majority of control applications that require strategies are bounded. Due to the physical limitations of components, the actions of the control system can only take values within a limited interval. For example, the control surface of an aircraft rudder can only move within a certain angle range, and control actions outside this range have no practical significance. Although after Gaussian sampling, the output of the Gaussian strategy can be restricted within the range of the actual control surface through action clipping, biases will be introduced during the gradient calculation of the neural network, resulting in problems in the strategy learning process.

[0005] The beta distribution is a bounded probability distribution with a probability interval of [0; 1]. The beta distribution can be described by two shape parameters. The beta distribution is the conjugate prior of the Bernoulli distribution, binomial distribution, negative binomial distribution, and geometric distribution. The finite probability interval of the beta distribution is consistent with the requirements in practical control problems. Therefore, two shape parameters can be simply obtained by using the output of a neural network to obtain a simple beta strategy. However, it is difficult to match the shape parameters directly output by the beta distribution with the actual control parameters, resulting in problems in the calculation of policy gradients during the learning process. The two parameters of the beta distribution jointly affect its mean, and the deviation range of the beta distribution is also associated with the two shape parameters simultaneously. Therefore, in the process of reinforcement learning, directly using a neural network to output the two shape parameters of the beta distribution faces many problems: in the process of calculating policy gradients, the gradient calculations of different dimensions are strongly correlated, which is not conducive to the implementation of the optimization process.

[0006] In view of the many problems existing in directly using a neural network to output the shape parameters of the beta distribution, the present invention proposes to use a neural network to output the mean and deviation parameters of the beta distribution. The mean and deviation of the beta distribution directly describe the shape and relative position of the distribution, and there is a direct relationship with the output of control actions. Therefore, outputting the mean and deviation of the beta distribution through a neural network is beneficial to the calculation of policy gradients and the optimization process of reinforcement learning.

[0007] In the calculation process, based on the analysis of the relationship between the mean of the beta distribution and the two shape parameters, the deviation of the beta distribution can be calculated in segments. When the mean of the beta distribution is less than or equal to 0.5, the deviation parameter of the beta distribution corresponds to the left shape parameter; when the mean of the beta distribution is greater than 0.5, the deviation parameter of the beta distribution corresponds to the right shape parameter. This method solves the problem of outputting the deviation of the beta distribution. It can be widely applied to various deep reinforcement learning algorithms.

[0008] Reinforcement learning can be divided into on-orbit and off-orbit reinforcement learning. During the interaction between the agent and the environment in reinforcement learning, the agent needs to provide a stochastic policy to explore the environment, so as to obtain learning samples. The agent can include two types of policies: the target policy and the behavior policy. Among them, the target policy is the policy that reinforcement learning currently hopes to obtain and needs to be updated according to the learning samples; the behavior policy is the policy that the agent currently executes and provides the behavior calculation results for the interaction between the agent and the environment. Among them, the target policy and the behavior policy of on-orbit reinforcement learning are the same. During the on-orbit reinforcement learning process, the agent calculates the learning samples according to the behavior policy and updates the policy according to the learning samples, so as to obtain a new policy. In off-orbit reinforcement learning, the agent uses the behavior policy to collect learning samples, and sample collection is a separate task of off-orbit reinforcement learning. The behavior policy is specifically responsible for obtaining learning data and has a certain degree of randomness, aiming to obtain more reasonable actions. The target policy improves its own performance by means of the samples collected by the behavior policy and the policy improvement method, and finally becomes the optimal policy.

[0009] The unbiased problem of beta probability distribution reparameterization proposed by the present invention can be applied to on-orbit and off-orbit reinforcement learning, and the learning efficiency is significantly higher than that of the Gaussian policy and the beta policy, and higher rewards can be obtained. Summary of the Invention

[0010] The Gaussian policy and the traditional beta policy can use a neural network to provide a policy for the agent. The sampling space of the Gaussian policy has no boundary, resulting in possible bias in the policy gradient calculation process. The traditional beta policy uses the characteristics of the beta distribution to represent a bounded policy, but the two parameters in the policy sampling process have a strong correlation with the policy output, resulting in relatively complex adjustment of the policy gradient. An improved method needs to be proposed for the traditional beta policy so that it can be applied to on-orbit and off-orbit reinforcement learning, improve the learning efficiency and obtain higher rewards.

[0011] For this purpose, the technical solution of the present invention: a reinforcement learning system based on beta probability distribution reparameterization, the system includes a neural network subsystem 100, a beta distribution parameter calculation subsystem 200, a beta distribution sampling subsystem 300, and a reinforcement learning training environment subsystem 400; the neural network subsystem 100 includes an environmental state input layer 101, a fully connected layer 102, a non-linear transformation layer 103, a beta distribution mean layer 104, and a beta distribution deviation layer 105; the environmental state parameters are input into the neural network through the environmental state input layer 101; through the fully connected layer 102 and the non-linear transformation layer 103, the input environmental state parameters are linearly and non-linearly transformed to extract valuable feature vectors;

[0012] The eigenvectors obtained from the mean layer 104 of the beta distribution through linear and nonlinear transformations are used to calculate the mean parameter of the beta distribution; the beta distribution deviation layer 105 calculates the deviation parameter of the beta distribution based on the eigenvectors obtained from linear and nonlinear transformations and outputs it using a piecewise method; the beta distribution parameter calculation subsystem 200 calculates the shape parameter of the beta distribution based on the mean and deviation parameters of the beta distribution; the beta distribution sampling subsystem 300 constructs a beta probability distribution according to the shape parameter of the beta distribution and samples according to the obtained beta probability distribution to obtain each action value sampled according to the current beta probability distribution; the strategy output by this beta distribution is the beta probability distribution reparameterization strategy; the reinforcement learning training environment subsystem 400 inputs the actions obtained based on the beta probability distribution reparameterization strategy into the reinforcement learning training environment; after interacting with the environment, rewards and new environmental state parameters are obtained, and the neural network parameters in the neural network subsystem 100 are updated using an optimization algorithm to obtain a new mean policy neural network, and a better strategy is obtained through exploration.

[0013] Further, the input environmental state parameters are input into the neural network input layer 101 in vector form. First, a dot product is performed with the weight vector of the neural network in the fully connected layer 102 and added to the bias vector to complete a linear transformation; then, it is input into the activation function in the nonlinear transformation layer 103 to complete the nonlinear transformation; by continuously updating the values of the weight vector and bias vector of the neural network, the extracted eigenvectors are gradually optimized and output as the mean and deviation of the beta distribution. Among them, the activation function is the tanh function, and its functional form is as follows:

[0014]

[0015] In the formula, x is the result after linear transformation; the number of neurons in one fully connected layer of the neural network is 64.

[0016] Further, the method for constructing the beta probability distribution is that the beta distribution is a parametric probability distribution defined on the interval (0,1) and has two positive shape parameters, denoted by a β and b β respectively. For x ∈ (0, 1), the probability density function of the beta distribution can be defined as:

[0017]

[0018] Among them, x is a random variable defined on the interval (0, 1); a β and b β are the two shape parameters of the beta distribution; B(a β , b β ) is the normalization constant and can be directly calculated using the gamma function:

[0019]

[0020] Among them, the gamma function is defined as: Thus, the probability density function of the beta distribution can be calculated by the following formula:

[0021]

[0022] Furthermore, the method for outputting the segmentation method of the beta distribution is to perform segmented output according to the expected value m of the beta distribution β Specifically: when 0 < m β ≤ 0.5, At this time, the neural network outputs the mean value m β and the shape parameter a β , and ensure that a β > 1; when 0.5 < m β < 1, At this time, let the neural network output the mean value m β and the shape parameter b β , and ensure that b β ≥ 1.

[0023] Furthermore, the calculation method of the shape parameter of the beta distribution is: when 0 < m β ≤ 0.5, b β can be calculated based on the values of m β and a β :

[0024]

[0025] When 0.5 < m β < 1, a β can be calculated based on the values of m β and b β :

[0026]

[0027] Furthermore, the method for obtaining each action value of the beta probability distribution sampling is that the expected value m of the random variable x ~ Beta(a, b) subject to the beta distribution β can be calculated by integrating over this interval:

[0028]

[0029] It can be adopted to represent the random policy and sample each action value according to the probability distribution.

[0030] Furthermore, when reparameterizing the beta probability distribution for on-orbit reinforcement learning, an optimized clipping algorithm is used to limit the magnitude of policy updates generated by the beta probability distribution. Multiple steps (usually mini-batches) of stochastic gradient descent are taken to maximize the following formula, and the policy θ k+1 The iterative calculation formula is:

[0031]

[0032] In the formula, L is the objective function, defined as:

[0033]

[0034] In the formula, π θ (a|s) is the probability distribution of action a under the given state s; is expressed as the estimation of the advantage function under the policy π θ ; ∈ is a hyperparameter that roughly represents how far the new policy is allowed to deviate from the old policy. The optimized clipping algorithm acts as a regularizer by eliminating the incentive for large policy changes, and the hyperparameter ∈ corresponds to the distance between the new policy and the old policy while still benefiting the objective.

[0035] Furthermore, when reparameterizing the beta probability distribution for off-orbit reinforcement learning, to prevent the policy generated by the beta probability distribution from prematurely converging to a poor local optimum, an entropy regularization method is adopted, and the optimized objective function is:

[0036]

[0037] In the formula, γ is the discount factor for each step of action; R(s t , a t s t+1 ) is the expected return of the next step after the given state and action; H(P) = E x~P [-logP(x)], entropy represents the randomness of a random variable; and α > 0 is the weight coefficient.

[0038] The state value function V π is changed to include the entropy reward for each time step:

[0039]

[0040] Q π is changed to include the entropy reward for each time step except the first time step:

[0041]

[0042] Advantages of the present invention:

[0043] ①The beta probability distribution reparameterized reinforcement learning system and learning method of the present invention can represent the beta distribution using a neural network, and calculate the mean and deviation parameters of the beta distribution according to the environmental state parameters input to the neural network.

[0044] ②The beta probability distribution reparameterized reinforcement learning system and learning method of the present invention can use the mean of the beta distribution output by the neural network to represent the action mean of reinforcement learning, and use the deviation parameter of the beta distribution output by the neural network to represent the deviation of the distribution.

[0045] ③The beta probability distribution reparameterized reinforcement learning system and learning method of the present invention represents the shape parameter of the beta distribution output by the neural network in the form of a piecewise function. When the mean of the beta distribution is less than or equal to 0.5, the deviation parameter of the beta distribution corresponds to the left shape parameter; when the mean of the beta distribution is greater than 0.5, the deviation parameter of the beta distribution corresponds to the right shape parameter.

[0046] ④The beta probability distribution reparameterized reinforcement learning system and learning method of the present invention can apply the beta probability distribution reparameterization to on-orbit and off-orbit reinforcement learning.

[0047] ⑤The beta probability distribution reparameterized reinforcement learning system and learning method of the present invention can improve the sample utilization efficiency, optimize the interaction process between the reinforcement learning algorithm and the environment, improve the learning efficiency and obtain higher rewards. Description of the Drawings

[0048] Figure 1 Schematic diagram of deep reinforcement learning using beta probability distribution reparameterization;

[0049] Figure 2 Schematic diagram of beta distributions with different deviations. (Example of a reparameterized beta distribution: when the mean of the beta distribution is 0.5, different shape parameters represent different probability distributions) Specific Embodiments

[0050] Refer to Figure 1 、 Figure 2 。 Figure 1 Schematic diagram of deep reinforcement learning using beta probability distribution reparameterization; Figure 2 Schematic diagram of beta distributions with different deviations.

[0051] A reinforcement learning system based on beta probability distribution reparameterization, the system includes a neural network subsystem 100, a beta distribution parameter calculation subsystem 200, a beta distribution sampling subsystem 300, and a reinforcement learning training environment subsystem 400; the neural network subsystem 100 includes an environmental state input layer 101, a fully connected layer 102, a non-linear transformation layer 103, a beta distribution mean layer 104, and a beta distribution deviation layer 105; environmental state parameters are input into the neural network through the environmental state input layer 101; through the fully connected layer 102 and the non-linear transformation layer 103, the input environmental state parameters are linearly and non-linearly transformed to extract valuable feature vectors; the beta distribution mean layer 104 calculates the mean parameter of the beta distribution according to the feature vectors obtained from the linear and non-linear transformations; the beta distribution deviation layer 105 calculates the deviation parameter of the beta distribution according to the feature vectors obtained from the linear and non-linear transformations and outputs it using a piecewise method; the beta distribution parameter calculation subsystem 200 calculates the shape parameter of the beta distribution based on the mean and deviation parameters of the beta distribution; the beta distribution sampling subsystem 300 constructs a beta probability distribution according to the shape parameter of the beta distribution and samples according to the obtained beta probability distribution to obtain each action value sampled according to the current beta probability distribution; the policy output by this beta distribution is the beta probability distribution reparameterization policy; through the reinforcement learning training environment subsystem 400, the actions obtained based on the beta probability distribution reparameterization policy are input into the reinforcement learning training environment; after interacting with the environment, rewards and new environmental state parameters are obtained, and the neural network parameters in the neural network subsystem 100 are updated using an optimization algorithm to obtain a new mean policy neural network, and a better policy is obtained through exploration.

[0052] Further, the input environmental state parameters are input into the neural network input layer 101 in vector form, first dot-multiplied with the weight vector of the neural network in the fully connected layer 102 and added to the bias vector to complete a linear transformation; then input into the activation function in the non-linear transformation layer 103 to complete the non-linear transformation; by continuously updating the values of the weight vector and bias vector of the neural network, the extracted feature vectors are gradually optimized and output as the mean and deviation of the beta distribution. Among them, the activation function is the tanh function, and its function form is as follows:

[0053]

[0054] In the formula, x is the result of the linear transformation; the number of neurons in one fully connected layer of the neural network is 64.

[0055] Further, the method for constructing the beta probability distribution is that the beta distribution is a parametric probability distribution defined on the interval (0, 1) and has two positive shape parameters, denoted by a β and b βRepresentation. For \(x\in(0, 1)\), the probability density function of the beta distribution can be defined as:

[0056]

[0057] where \(x\) is a random variable defined on the interval \((0, 1)\); \(a\) β and \(b\) β are two shape parameters of the beta distribution; \(B(a\) β , \(b\) β ) is the normalization constant, which can be directly calculated using the gamma function:

[0058]

[0059] where the gamma function is defined as: Thus, the probability density function of the beta distribution can be calculated by the following formula:

[0060]

[0061] Furthermore, the method for outputting the segmentation of the beta distribution is to output the segmentation according to the expected value \(m\) β of the beta distribution, specifically: when \(0 \lt m\) β \(\leq 0.5\), At this time, the neural network outputs the mean \(m\) β and the shape parameter \(a\) β , and ensures that \(a\) β \gt 1\); when \(0.5 \lt m\) β \lt 1\), At this time, let the neural network output the mean \(m\) β and the shape parameter \(b\) β , and ensures that \(b\) β \geq 1\).

[0062] Furthermore, the method for calculating the shape parameters of the beta distribution is that when \(0 \lt m\) β \(\leq 0.5\), \(b\) β can be calculated based on the values of \(m\) β and \(a\) β :

[0063]

[0064] When \(0.5 \lt m\) β \lt 1\), \(a\) β can be calculated based on the values of \(m\) β and \(b\) β :

[0065]

[0066] Further, the method for obtaining each action value by sampling from the beta probability distribution is the expected value m of a random variable x ~ Beta(a, b) that follows the beta distribution. β It can be calculated by integrating over this interval:

[0067]

[0068] Can adopt Represents a stochastic policy and samples each action value according to the probability distribution.

[0069] Embodiment

[0070] Apply the beta probability distribution reparameterization method to reinforcement learning for dealing with Markov decision process problems. The tuple Is used to describe the Markov decision process, where Is the state space observed by the agent, Represents the action space, Describes the transition probability kernel function, Is the reward function defined for the application, r t ∈[-r max , r max is a reward for a certain uniform constraint, the initial state distribution is d 0 , γ is the discount factor that reduces the impact of the reward according to the time scale.

[0071] Define the policy π ∈ Π as the probability distribution of actions in a given state, π(s, a) = Pr(a t = a|s t = s), where Π is the set consisting of all possible policies of the agent. The policy can choose different probability distributions, such as Gaussian distribution, Cauchy distribution, traditional beta distribution, etc. The present invention adopts the beta mean distribution. The policy is described by the parameters of the probability distribution , where μ θ (s, a) = Pr(a t = a|s t = s, θ t = θ), where t represents the time step.

[0072] The state value function V π Describes the expected value of future cumulative rewards in the state. The mapping from the state to the sum of expected discounted returns is defined as where T is the time step. The state-action value function characterizes the expected value of future cumulative rewards in the state-action and is defined by the Q function:

[0073] The deep reinforcement learning algorithm represents the policy using a neural network, and the policy is output by the neural network, π(s, a) = Pr(a t= a|s t = s); At the same time, the neural network can also be used to approximate the state value function. The representation ability of the deep neural network has enabled a leap in the ability of deep reinforcement learning.

[0074] Deep reinforcement learning uses deep neural networks to represent the policy function and the value function. The complex representation of deep neural networks has greatly expanded the scope of problems that reinforcement learning can handle. However, the Gaussian policy is an unbounded distribution and is inconsistent with the operating range of the actual control surface. The traditional beta distribution is somewhat committed to solving the problems existing in the Gaussian policy, but the shape parameters of the two beta distributions output by its neural network are strongly correlated, which is not conducive to the optimization of parameters during the learning process.

[0075] The following specifically analyzes the characteristics of the beta distribution. The beta distribution is a parametric probability distribution defined on the interval (0, 1) and has two positive shape parameters, denoted by a β and b β respectively. The beta distribution is the conjugate prior of the Bernoulli distribution, binomial distribution, negative binomial distribution, and geometric distribution. For x ∈ (0, 1), the probability density function of the beta distribution can be defined as:

[0076]

[0077] where B(a β , b β ) is the normalization constant and can be directly calculated using the gamma function:

[0078]

[0079] where the gamma function is defined as: Thus, the probability density function of the beta distribution can be calculated by the following formula:

[0080]

[0081] Since the interval of the beta distribution is [0, 1], the expected value of the random variable x ~ Beta(a, b) following the beta distribution can be calculated by integrating over this interval:

[0082]

[0083] If the shape parameters of the beta distribution are directly output by the neural network, then can be used to represent the stochastic policy.

[0084] Since the beta distribution has a finite value range and no probability density falls outside the boundary, the beta strategy can be consistent with the control range in the actual control process, which is beneficial to the progress of deep reinforcement learning. However, the traditional beta strategy directly uses a neural network to output the two shape parameters of the beta distribution. The two shape parameters of the beta distribution are highly correlated, and it is difficult to coordinate the changes of the shape parameters with the changes of the actual policy gradient. Therefore, although the traditional beta strategy has achieved performance improvement, there are still problems in policy optimization. To solve this problem, the present invention proposes a reparameterization of the beta probability distribution.

[0085] In the reparameterization of the beta probability distribution, the beta distribution is represented by a neural network. The input of the neural network is the state of the environment, and its output is the mean and deviation of the beta distribution. During the output process of the control policy, the beta distribution needs to be a unimodal distribution. Therefore, the two shape parameters of the beta distribution need to be greater than 1.

[0086] The mean of the beta distribution is:

[0087]

[0088] The above formula can be transformed into

[0089]

[0090] Therefore, the mean of the beta distribution actually depends on the ratio of the two shape parameters. The domain of the beta distribution can be divided into two intervals, the left interval and the right interval, with different ratios of shape parameters:

[0091] 0 < m β ≤ 0.5;

[0092] 0.5 < m β < 1.

[0093] When the mean m of the beta distribution β is determined, only by giving any one of the shape parameters a β , b β , a β + b β can all the parameters of the beta distribution be given. The traditional method generally uses a β + b β as the deviation value. However, when using the beta distribution to output the policy, generally the beta distribution needs to be a unimodal distribution, that is, it is required that a β ≥ 1, and b β ≥ 1. If a β + b βAs the deviation value of the beta distribution, it is impossible to control the output of the neural network during the actual calculation process. Even if the neural network is forced to output a β +b β ≥1, it cannot be guaranteed that a β ≥1, and b β ≥1.

[0094] During the actual calculation process, the neural network outputs the beta mean and deviation in segments. Based on the characteristics of the beta distribution, when 0 < m β ≤0.5, At this time, the neural network outputs the mean m β and the shape parameter a β , and it is guaranteed that a β > 1. Since b β > 1 necessarily holds. b β can be calculated based on the values of m β and a β :

[0095]

[0096] Similarly, when 0.5 < m β < 1, At this time, the neural network is made to output b β , and it is guaranteed that b β ≥1. Since a β ≥1 necessarily holds. The neural network outputs m β and b β (b β ≥1), and a β can be calculated based on the values of m β and b β :

[0097]

[0098] The method of segmental output is adopted to ensure that the output of the neural network conforms to the characteristics of the beta distribution.

[0099] The reparameterization of the beta probability distribution can be used for action sampling in the reinforcement learning process. Both on-orbit reinforcement learning and off-orbit reinforcement learning can adopt the reparameterization of the beta probability distribution. Specific embodiments include the application of the reparameterization of the beta probability distribution in on-orbit and off-orbit reinforcement learning.

[0100] The reparameterization of the beta probability distribution is used for on-orbit reinforcement learning

[0101] The proximal policy optimization algorithm is an important method for on-orbit reinforcement learning. The reparameterization of the beta probability distribution can be used in the proximal policy optimization algorithm.

[0102] The Proximal Policy Optimization (PPO) algorithm can be used in environments with discrete or continuous action spaces. The PPO algorithm aims to solve the problem of how to take the largest possible improvement step for the policy using the currently available samples, while also preventing performance collapse due to unexpected policy changes. The PPO algorithm uses a series of first-order methods to address this issue. The PPO algorithm employs some techniques to ensure that the new policy does not deviate too far from the old policy.

[0103] According to the different adjustment methods adopted, the PPO algorithm can be divided into two types: KL constraint and clipping constraint. Among them, the KL constraint approximately solves the constrained update, but only penalizes the KL divergence in the objective function rather than making it a hard constraint, and automatically adjusts the penalty coefficient during training to obtain an appropriate scaling. The clipping constraint does not have a KL divergence term in the optimization objective and relies on specialized clipping in the objective function to eliminate the rewards for new policies that are far from the old policies.

[0104] The policy update of the Proximal Policy Optimization Clipping (PPO Clip) algorithm is carried out by optimizing the following formula. The iterative formula for the policy θ can be expressed as:

[0105]

[0106] Typically, multiple steps (usually in mini-batches) of stochastic gradient descent are taken to maximize the objective. Here, the objective function L is defined as:

[0107]

[0108] In the formula, π θ (a|s) is the probability distribution of action a under the given state s; denotes the estimate of the advantage function under the policy π θ ; ∈ is a hyperparameter that roughly indicates how far the new policy is allowed to deviate from the old policy.

[0109] The above formula can be simplified to:

[0110]

[0111] where

[0112] g(∈, A) = (1 + ∈)A A ≥ 0

[0113] g(∈, A) = (1 - ∈)A A ≤ 0

[0114] When the advantage function A is positive: Assuming that the advantage of the state-action pair is positive, in this case, its contribution to the objective is reduced to

[0115]

[0116] Because the advantage function is positive, if the action π θ (a|s) has a higher probability, then the objective function will increase. However, the minimum function in the above equation limits how much the objective can increase. When At this time, the minimum function takes In this way, the new policy does not benefit from being far from the old policy.

[0117] When the advantage function is negative, assume that the advantage value of this state-action pair is negative. In this case, its contribution to the objective is reduced to

[0118]

[0119] Since the advantage function is negative, if the action becomes less likely, that is, if π θ (a|s) has a lower probability, the objective will increase. However, the maximum function in the above equation limits how much the objective can increase. When At this time, the maximum function takes The maximum function limits the reduction of the action probability until Therefore, the new policy does not benefit from being far from the old policy.

[0120] Clipping acts as a regularizer by eliminating the incentive for large policy changes, and the hyperparameter ∈ corresponds to the distance between the new policy and the old policy while still benefiting the objective. Although this clipping helps a lot in ensuring reasonable policy updates, it is still possible to end up with a new policy that is very far from the old policy.

[0121] Algorithm: Proximal Policy Optimization Algorithm Based on Reparameterization of Beta Probability Distribution

[0122]

[0123] Reparameterization of Beta Probability Distribution for Off-Policy Reinforcement Learning

[0124] SAC is an important off-policy reinforcement learning method. Here, the application of reparameterization of beta probability distribution in the SAC algorithm is taken as an example of its application in off-policy reinforcement learning. SAC is an algorithm that optimizes a stochastic policy in an off-policy manner, and its core feature is entropy regularization. The policy is trained to maximize the trade-off between the expected return and entropy, where entropy is a measure of the randomness in the policy. This is closely related to the exploration-exploitation trade-off: increasing entropy leads to more exploration, which can accelerate future learning. The entropy regularization process of SAC can also prevent the policy from converging prematurely to a bad local optimum.

[0125] Key Equation

[0126] SAC belongs to the entropy-regularized reinforcement learning method. In the SAC algorithm, the agent obtains a reward proportional to the policy entropy at each time step. The optimization objective function of SAC is:

[0127]

[0128] where γ is the discount factor for each step of the action; R(s t , a t s t+1 ) is the expected return of the next step given the state and action; H(P) = E x~P [-logP(x)], the entropy represents the randomness of the random variable; and α > 0 is the weight coefficient.

[0129] The state value function V π is changed to include the entropy reward at each time step:

[0130]

[0131] Q π is changed to include the entropy reward at each time step except the first time step:

[0132]

[0133] V π (s) and Q π (s, a) are derived as follows:

[0134] V π (s) = E τ~π [Q π (s, a)] + αH(π(.|s t ))).

[0135] Q π 's Bellman equation is

[0136] Q π (s, a) = E τ~π [R(s, a, s′) + y(Q π (s′, a′) + H(π(.|s′)))]),

[0137] Q π (s, a) = E τ~π [R(s, a, s′) + γV π (s′)].

[0138] SAC simultaneously learns a policy π θ and two Q functions and where π θReparameterized by the beta probability distribution.

Claims

1. A reinforcement learning system based on beta probability distribution reparameterization, characterized in that: The invention comprises a neural network subsystem (100), a beta distribution parameter calculation subsystem (200), a beta distribution sampling subsystem (300) and a reinforcement learning training environment subsystem (400); the neural network subsystem (100) comprises an environment state input layer (101), a fully connected layer (102), a nonlinear transformation layer (103), a beta distribution mean layer (104) and a beta distribution deviation layer (105); the environment state parameters are input into the neural network through the environment state input layer (101); the input environment state parameters are linearly and nonlinearly transformed through the fully connected layer (102) and the nonlinear transformation layer (103) to extract the environment state parameters. The beta distribution mean layer (104) calculates the mean parameter of the beta distribution according to the characteristic vector obtained by linear and nonlinear transformation; the beta distribution deviation layer (105) calculates the deviation parameter of the beta distribution according to the characteristic vector obtained by linear and nonlinear transformation, and outputs it in a segmented manner; the beta distribution parameter calculation subsystem (200) calculates the shape parameter of the beta distribution based on the mean and deviation parameters of the beta distribution; the beta distribution sampling subsystem (300) constructs a beta probability distribution according to the shape parameter of the beta distribution, and samples according to the obtained beta probability distribution to obtain the action values ​​sampled according to the current beta probability distribution; The strategy output by this beta distribution is a beta probability distribution reparameterization strategy; through the reinforcement learning training environment subsystem (400), the action obtained based on the beta probability distribution reparameterization strategy is input into the reinforcement learning training environment; after interacting with the environment, rewards and new environmental state parameters are obtained, and the optimization algorithm is used to update the neural network parameters in the neural network subsystem (100) to obtain a new mean strategy neural network, and a better strategy is obtained through exploration.

2. The beta probability distribution based reparameterized reinforcement learning system according to claim 1, characterized in that: The input environmental state parameters are input into the neural network input layer (101) in the form of vectors, first multiplied by the weight vector of the neural network in the fully connected layer (102), and added to the bias vector to complete a linear transformation; then input into the activation function in the nonlinear transformation layer (103) to complete the nonlinear transformation; by continuously updating the values ​​of the weight vector and bias vector of the neural network, the extracted feature vector is gradually optimized and output as the mean and bias of the beta distribution. Among them, the activation function is the tanh function, and its function form is as follows: Where x is the result after linear transformation; the number of neurons in a fully connected layer of the neural network is 64.

3. The beta probability distribution based reparameterized reinforcement learning system according to claim 1, characterized in that: The method of constructing the Beta probability distribution is that the Beta distribution is a parametric probability distribution defined on the interval (0,1) with two positive shape parameters, a and β and b β Indicates; for x∈(0,1), the probability density function of the Beta distribution can be defined as: Among them, x is a random variable defined on the interval (0,1); a β and b β are the two shape parameters of the Beta distribution; B(a β ,b β ) is the normalization constant, which can be directly calculated using the gamma function: Among them, the gamma function is defined as: Thus, the probability density function of the Beta distribution can be calculated as follows:

4. The beta probability distribution based reparameterized reinforcement learning system according to claim 1, characterized in that: The output method of the segmented method of Beta distribution is as follows: according to the expected value m of Beta distribution β Perform segmented output, specifically: when 0 <m β When ≤0.5, At this time, the neural network outputs the mean value m β and shape parameter a β , and ensure a β >1;0.5 <m β <1 hour, At this time, let the neural network output mean m β and shape parameter b β , and ensure that b β ≥1.

5. The beta probability distribution based reparameterized reinforcement learning system according to claim 1, characterized in that: The shape parameter of the Beta distribution is calculated as follows: <m β When ≤0.5, b β Can be based on m β and a β The value of is calculated: 0.5 <m β <1 hour, a β Can be based on m β and b β The value of is calculated:

6. The beta probability distribution based reparameterized reinforcement learning system according to claim 1, characterized in that: The method for obtaining the action values ​​of the Beta probability distribution sampling is to obtain the expected value m of the random variable x~Beta(a,b) that obeys the Beta distribution β It can be calculated by integrating over this interval: Can be used Represents a random strategy and samples each action value according to the probability distribution.

7. The learning method based on the beta probability distribution reparameterized reinforcement learning system according to any one of claims 1 to 6, characterized in that: When reparameterizing the Beta probability distribution for on-orbit reinforcement learning, an optimized clipping algorithm is used to limit the policy update amplitude generated by the Beta probability distribution; multiple steps of stochastic gradient descent are taken to maximize the following realization, then the policy θ k+1 The iterative calculation formula is: Where L is the objective function, defined as: In the formula, π θ (a|s) is the probability distribution of action a given state s; Denoted as strategy π θ is an estimate of the advantage function under the assumption that ∈ is a hyperparameter that roughly represents how far the new policy is allowed to be from the old policy; the optimized clipping algorithm acts as a regularizer by removing the incentive for large policy changes, and the hyperparameter ∈ corresponds to how far the new policy can be from the old policy while still benefiting the objective.

8. The learning method based on the beta probability distribution reparameterized reinforcement learning system according to any one of claims 1 to 6, characterized in that: When reparameterizing the Beta probability distribution for off-track reinforcement learning, in order to prevent the strategy generated by the Beta probability distribution from converging to a bad local optimum too early, the entropy regularization method is used, and the optimization objective function is: In the formula, γ is the discount factor for each action; R(s t ,a t s t+1 ) is the expected reward for the next step after a given state and action; H(P) = E x~P [-logP(x)], entropy represents the randomness of the random variable; and α>0 is the weight coefficient. State value function V π Changed to include an entropy bonus at each time step: Q π Change to include an entropy bonus for every time step except the first: