An attitude control device, system, and method based on reinforcement learning

By combining the neural network Lyapunov function and control obstacle function, the posture control method of reinforcement learning is improved, the security problem in the online learning process is solved, and the stable control of the system within the safe state is realized.

CN114638346BActive Publication Date: 2025-07-08INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210333060.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-30
Publication Date
2025-07-08
Estimated Expiration
2042-03-30

AI Technical Summary

Technical Problem

The existing reinforcement learning methods are difficult to ensure the safety of the system during attitude control, especially during online learning, which cannot effectively ensure the safety of the interaction between the agent and the environment.

Method used

The Lyamanov function based on neural network is used to fit the global security information of the learning system, combined with the model-free reinforcement learning and control obstacle function (CBF) controller, and the Lyamanov function of the neural network is used to guide the execution of the action to provide security guarantees.

Benefits of technology

Without intervention in the absence of model-free strategy and CBF learning, the security and learning efficiency of the system are improved, the forward invariance of the system within the safe state range is ensured, and the stability and security of the control algorithm are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114638346B_ABST
    Figure CN114638346B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides an attitude control device, system and method based on reinforcement learning. The attitude control device based on reinforcement learning includes: a model-free reinforcement learning controller configured to obtain an initial action control amount according to a current state; a historical CBF controller configured to obtain a historical safety action control amount according to the current state; a CBF controller configured to obtain a current safety action control amount according to the initial action and the historical safety action; a neural network Lyapunov learner configured to determine the safety of a comprehensive safety action control amount and provide the comprehensive safety action control amount to a dynamics model when it is confirmed that the comprehensive safety action control amount is safe, wherein the comprehensive safety action control amount is determined according to the initial action control amount, the historical safety action control amount and the current safety action control amount.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent decision-making. Specifically, embodiments of this application relate to an attitude control device, system, and method based on reinforcement learning. Background Art

[0002] In recent years, many difficult problems in machine learning and artificial intelligence have made rapid progress in multiple fields such as computer vision, video games, autonomous vehicles, and Go. These advancements have excited people about the positive potential of artificial intelligence in transforming medicine, science, and transportation, while also raising concerns about the privacy, security, fairness, economy, and military aspects of autonomous systems, as well as concerns about the long-term impact of powerful artificial intelligence.

[0003] As a machine learning method, reinforcement learning has been widely studied since the 1980s. Compared with supervised learning, reinforcement learning does not require experts to classify each sample. Instead, it simply collects immediate rewards while moving and optimizes for the expected long-term rewards (while supervised learning optimizes for immediate rewards). Its key advantage is that the design of rewards is usually much simpler than classifying all data samples. However, in many cases, the safety of the agent is crucial, especially on some expensive control platforms. Researchers not only focus on maximizing long-term rewards but also pay great attention to avoiding damage. Therefore, in the process of learning the policy, how to ensure the safety of the system during the attitude control process is a very critical issue.

[0004] Safe learning and safe exploration are two directions for solving the safety problems of reinforcement learning. They have similar characteristics to the trade-off between exploration and exploitation. Traditional reinforcement learning uses the value function and cumulative rewards as the learning objectives, while safe learning will increase the learning of safety information about the environment and use it as the basis for policy execution to guide and improve the policy. This method is mainly for offline reinforcement learning. Safe exploration, on the other hand, focuses on the safety during the interaction and learning process between the agent and the environment. It guides the exploration interaction process by adding additional models or expert information. This method is an improvement for online learning. Obviously, by learning the global information of the environment, the safety of the system can be better guaranteed, but the safety during the learning process cannot be guaranteed. And restricting the learning process cannot traverse the global safety state.

[0005] Therefore, how to improve the existing reinforcement learning process has become a technical problem to be solved urgently. Summary of the Invention

[0006] The purpose of the embodiments of the present application is to provide a posture control device, system and method based on reinforcement learning. Through the neural network Lyapunov function of the embodiments of the present application, the global safety information of the learning system is fitted by the neural network, providing guidance for action execution, and ensuring the safety of the system without intervening in the model-free policy and CBF learning.

[0007] In a first aspect, some embodiments of the present application provide a posture control device based on reinforcement learning. The device includes: a model-free reinforcement learning controller configured to obtain an initial action control amount according to the current state; a historical CBF controller configured to obtain a historical safety action control amount according to the current state; a CBF controller configured to obtain a current safety action control amount according to the initial action and the historical safety action; and a neural network Lyapunov learner configured to determine the safety of the comprehensive safety action control amount and provide the comprehensive safety action control amount to the dynamics model when it is confirmed that the comprehensive safety action control amount is safe, where the comprehensive safety action control amount is determined according to the initial action control amount, the historical safety action control amount, and the current safety action control amount.

[0008] Some embodiments of the present application use a neural network to fit and learn the global safety information of the system, providing guidance for action execution, and ensuring the safety of the system without intervening in the model-free policy and CBF learning.

[0009] In some embodiments, the CBF controller is configured to obtain the current safety action control amount through the following formula (1);

[0010]

[0011] where ∈ is a relaxation variable under safe conditions, and K ∈ is a large constant that penalizes safety violations. The obtained a t by solving the above formula is the current safety action control amount, and ||a t ||2 refers to the sum of the squares of all values of the action vector a t , and argmin represents the operation of taking the minimum value.

[0012] The CBF controller of the present application uses the above formula to obtain the output action control amount. Compared with the existing formula, this CBF controller can be combined with the model-free reinforcement learning controller and can learn and explore with higher safety assurance and efficiency.

[0013] In some embodiments, the formula (1) is solved through the following constraint conditions:

[0014]

[0015] For all \(i = 1, 2, \ldots, M\)

[0016] where \(P\) T represents the transpose of the control barrier function control parameter vector, \(q\) is the control parameter, is the initial action control quantity obtained by the model-free reinforcement learning controller according to the current state, is the safe action control quantity obtained by the CBF controller during the \(j\)-th loop iteration, \(\mu\) d (\(s\) t ) and \(\sigma\) d (\(s\) t ) are the mean and variance of the system state calculated by the Gaussian process, \(k\) δ is the confidence interval parameter satisfying Gaussian regression, \(\eta\) represents the parameter for adjusting the control barrier condition, refers to the model-free reinforcement learning controller, refers to the upper bound of the control quantity of control dimension \(i\), refers to the lower bound of the control quantity of control dimension \(i\), \(|p|\) T is the transpose of the vector composed of the absolute values of all terms of vector \(p\).

[0017] The above constraints of the present application can calculate the safe action of the system by using the environment models \(f\) and \(g\), ensuring that the output control quantity will not put the system in a dangerous state. That is to say, the above constraints are control barrier conditions, which guarantee the forward invariance of the system by modeling the system state, provide safety limitations for each action output of the system, and ensure that the system will not enter a dangerous state.

[0018] In some embodiments, the historical CBF controller obtains the historical safe action control quantity according to the following formula:

[0019]

[0020] where to are calculated by the CBF controller each time a loop is performed, \(k\) is used to record the number of loops and the value range of \(k\) is greater than or equal to 1.

[0021] Some embodiments of the present application obtain the output of the CBF controller through the above formula, can utilize the historical information of the CBF controller, and ensure the exploration efficiency and stability of the algorithm during the interaction and learning process.

[0022] In some embodiments, the comprehensive safe action control quantity is the sum of the initial action control quantity, the historical safe action control quantity, and the current safe action control quantity.

[0023] In some embodiments of the present application, the output of the model-free reinforcement learning controller is safely compensated by the output of the CBF controller and the historical CBF controller. When the output of the model-free controller is too high, a negative value is compensated, and when its output is too low, a positive value is compensated to ensure the safety of the system.

[0024] In a second aspect, some embodiments of the present application provide an attitude control system based on reinforcement learning. The system includes: a dynamic model configured to provide the current state or receive an input comprehensive safe action control amount; and an attitude control device based on reinforcement learning as described in any embodiment of the first aspect.

[0025] In a third aspect, some embodiments of the present application provide an attitude control method based on reinforcement learning. The attitude control method includes: obtaining an initial action control amount through a model-free reinforcement learning algorithm, and obtaining a historical safe action control amount through a historical CBF controller; inputting the initial action control amount and the historical safe action control amount into the CBF controller to obtain the current safe action control amount; determining the safety of the comprehensive safe action control amount through the Lyapunov algorithm, and sending the comprehensive safe action control amount to the dynamic model when it is confirmed that the comprehensive safe action control amount is safe.

[0026] In some embodiments, before obtaining the initial action control amount through the model-free reinforcement learning algorithm and obtaining the historical safe action control amount through the historical CBF controller, the control method further includes: initializing the parameters of the model-free reinforcement learning controller and the historical CBF controller.

[0027] In some embodiments, the control algorithm adopted by the model-free reinforcement learning includes PPO or SAC.

[0028] In some embodiments, the parameters corresponding to the model-free reinforcement learning controller include: a memory pool D and an action pool A. Among them, the memory pool D is used for interactive data (s, u, r, s'), where s is the current state, u is the executed action, r is the immediate reward, and S' is the state at the next moment. The memory pool is used to learn and update the parameter θ of the model-free reinforcement learning controller. The action pool is used to store action data (s, u), where the action pool is used for supervised learning of the historical CBF controller. Description of the Drawings

[0029] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following accompanying drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related accompanying drawings can also be obtained based on these drawings.

[0030] Figure 1 The architecture diagram of the attitude control device provided by the related technology;

[0031] Figure 2 The architecture diagram of the attitude control device based on reinforcement learning provided by the embodiments of the present application;

[0032] Figure 3 、 Figure 4 and Figure 5 The process of training the parameterized Lyapunov candidate function and the extended safety set provided by the embodiments of the present application;

[0033] Figure 6 One of the flowcharts of the attitude control method based on reinforcement learning provided by the embodiments of the present application;

[0034] Figure 7 Another flowchart of the attitude control method based on reinforcement learning provided by the embodiments of the present application. Detailed implementation manners

[0035] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.

[0036] It should be noted that similar reference numerals and letters denote similar items in the following accompanying drawings. Therefore, once an item is defined in one accompanying drawing, it does not need to be further defined and explained in subsequent accompanying drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.

[0037] Please refer to Figure 1 , Figure 1 The architecture diagram of the attitude control device provided by the related technology, that is, the initial algorithm structure using the control barrier function (CBF), which includes: a dynamic model, a model-free reinforcement learning controller, and a CBF controller.

[0038] Different from Figure 1 is that the attitude control device architecture of the embodiments of the present application Figure 2 is shown. It should be noted that the present application Figure 2 provides an alternative to the CBF-based control method (that is, an alternative to Figure 1The structure of the architecture) mainly increases the control limit within the control loop through the CBF controller. It can be understood that Figure 1 directly corrects the action output of the model-free reinforcement learning controller and executes it into the environment or the dynamics model, relatively lacking exploration data of global information. However, the device of the embodiment of the present application uses the neural network Lyapunov function to fit and learn the global safety information of the learning system through the neural network, providing guidance for action execution, and ensuring the safety of the system without interfering with the model-free policy and CBF learning. That is to say, the embodiment of the present application passes the environmental model s t+1 = f(s t ) + g(s t )a t + d(s t ) in the CBF controller to learn the safety information of the environment and model the global information of the environment.

[0039] The embodiment of the present application aims to use the model-based reinforcement learning method to mathematically model the uncertainty of the environment. The basic framework of this structure is based on the Markov decision process with an infinite discount rate, which is composed of (S, A, f, g, d, r, ρ0, γ). A is a set of actions, f: S → S is the nominal non-driven dynamics, g: S → R n,m is the nominal driving dynamics, and d: S → S is the unknown system dynamics model. The way the entire system evolves over time is given by the following formula:

[0040] s t+1 = f(s t ) + g(s t )a t + d(s t ).

[0041] Among them, s t ∈ S, a t ∈ A, f and g form a known nominal dynamics model, and d represents the unknown model. In fact, the nominal model may be quite poor and a better dynamics model must be learned through data.

[0042] Combined with the knowledge of cybernetics, it is used to realize the safety exploration and learning of automatic control problems. The embodiment of the present application aims at the safety problem of reinforcement learning attitude control. From the perspective of the control barrier function (CBF), combined with the model-based reinforcement learning control algorithm and the neural network Lyapunov function, it learns and restricts the safe range of the state-action space of the control problem, and realizes a safe learning process while maintaining the learning sample efficiency of the system. Specifically:

[0043]

[0044] where h is the structure of the control barrier function, and P T is the transpose of the control parameter vector of the control barrier function, which contains n parameters, and n is the dimension of the system state; s is the system state, and q is also a control parameter. The control barrier function is solved by the formula for solution.

[0045] The main module structure of the method of some embodiments of this application is as Figure 1 shown. During the simulation experiment process, the state information of the system is obtained from the dynamic model. Then, the corresponding output action control quantities corresponding to the respective states are obtained in the model-free reinforcement learning controller and the historical CBF controller respectively. Finally, a safe output action control quantity is synthesized in the CBF controller, output to the neural network Lyapunov controller for judgment, and finally output to the simulation or real physical environment for execution, forming a stable control closed-loop system.

[0046] The control barrier function (CBF) provides sufficient guarantee for system safety by using Lyapunov-like parameters, ensuring the forward invariance of the system within the safe set. Therefore, it is a tool that can ensure safety throughout the learning process and can provide a control framework for the synthesized control system. This structure enables a large number of existing model-free reinforcement learning algorithms to be combined into the CBF-based framework to achieve safety control. This has important practical value for applying the reinforcement learning algorithms trained in the simulation environment to the actual physical environment.

[0047] Next, in combination with Figure 2 exemplarily elaborate on the reinforcement learning-based attitude control device of the embodiments of this application.

[0048] As Figure 2 shown, some embodiments of this application provide a reinforcement learning-based attitude control device, which includes: a model-free reinforcement learning controller configured to obtain an initial action control quantity according to the current state (i.e., Figure 2 the initial action u1 of Figure 2 ); a historical CBF controller configured to obtain a historical safe action control quantity according to the current state (i.e., Figure 2 the historical safe action control quantity u2 of Figure 2The safety action u'), where the comprehensive safety action control quantity is determined according to the initial action control quantity, the historical safety action control quantity, and the current safety action control quantity.

[0049] In some embodiments of the present application, the CBF controller is configured to obtain the current safety action control quantity through the following formula (1);

[0050]

[0051] where ∈ is a slack variable under safety conditions, and K ∈ is a large constant that punishes safety violations. The a obtained by solving the above formula t is the current safety action control quantity.

[0052] In some embodiments of the present application, the formula (1) is solved through the following constraint conditions:

[0053]

[0054] and

[0055] For all i = 1,…, M.

[0056] where f, g, s, and a refer to the quantities in the system model, and s t+1 = f(s t ) + g(s t )a t + d(s t )

[0057] is the structure of the control barrier function. P T is the transpose of the control barrier function control parameter vector, which contains n parameters. n is the dimension of the system state. q is also a control parameter and is a constant. is the initial action control quantity obtained by the model-free reinforcement learning controller according to the current state. is the safety action control quantity obtained by the CBF controller during the jth loop iteration. μ d (s t ) and σ d (s t ) are the mean and variance of the system state calculated by the Gaussian process. k δ is the confidence interval parameter that satisfies Gaussian regression. σ d represents the variance of the error model d. μ d represents the mean of the error model d. η represents the parameter for adjusting the control barrier condition, and the value of this parameter is specified manually. refers to the model-free reinforcement learning controller. Refers to the upper bound of the control quantity for controlling dimension i, Refers to the lower bound of the control quantity for control dimension i, u RL (S) is used to represent a model-free reinforcement learning controller, |p| T Is the transpose of the vector composed of the absolute values of all terms of vector p.

[0058] In some embodiments of the present application, the historical CBF controller obtains the historical safe action control quantity according to the following formula:

[0059]

[0060] Wherein, To Indicates that it is calculated by the CBF controller each time a loop is performed. k is used to record the number of loops, and the value range of k is greater than or equal to 1.

[0061] In some embodiments of the present application, the comprehensive safe action control quantity is the sum of the initial action control quantity, the historical safe action control quantity, and the current safe action control quantity.

[0062] The relevant controller is elaborated below in combination with the algorithm Figure 2 of

[0063] Figure 2 The basic framework of n,m is based on the Markov decision process with an infinite discount rate, which is composed of (S, A, f, g, d, r, ρ0, γ). A is a set of actions, f: S → S is the nominal non-driven dynamics, g: S → R

[0064] s t+1 = f(s t ) + g(s t )a t + d(s t ) (1)

[0065] Where s t ∈ S, representing the set composed of all states, a t ∈ A, representing the set composed of all actions. f and g form a known nominal dynamics model, and d represents an unknown model. In fact, a better dynamics model must be learned through data. The structure of CBF is defined as:

[0066]

[0067] Among them, η represents the degree to which the control barrier function changes the state within the safety set (if η = 0, the gradient barrier condition simplifies to the Lyapunov condition). The form of the affine barrier function is defined as h = p T s + q, (p ∈ R n , q ∈ R). After using the Gaussian process to increase the parameter confidence interval of the system state uncertainty model, the problem of finding the boundary conditions of the safety set can be transformed into a quadratic programming problem:

[0068]

[0069] Among them, ∈ is a slack variable under safety conditions, K ∈ is a large constant that penalizes safety violations, and P is the absolute value of the vector P. The optimization is insensitive to the K ∈ parameter as long as it is large enough (e.g., 10e12), that is, serious penalties will be imposed when safety restrictions are violated. is an arbitrary model-free reinforcement learning controller. refers to the sum of the CBF controllers of all previous steps. Measuring the dynamic uncertainty through the Gaussian process model enables the embodiments of the present application to prove the system safety even when using a poor nominal model. μ d and σ d are the mean and variance of the Gaussian process respectively. The Gaussian process is a very commonly used non-parametric regression method and has a good effect on error fitting. Here, M refers to the dimension of the action space, and refer to the constraints that should be satisfied in each dimension, which are generally formulated by humans. The action obtained by formula (3) is the safe action control quantity output by the CBF controller, that is, the formula 3 provided by the embodiments of the present application adds an output limit to the solution of the safe action control quantity, providing higher safety protection.

[0070] To improve the computational efficiency of the algorithm during real-time interaction, the embodiments of the present application adopt a historical CBF controller to optimize the solution method of this quadratic programming problem. The historical CBF can be expressed as:

[0071]

[0072] Among them to represent the output actions of the CBF controller during each loop, and each is solved using the above formula (3). This method combined with any model-free method can improve the safety of the agent system under the conditions of the model-free reinforcement learning control algorithm. Compared with the previous CBF method (such as Figure 1As shown in the figure, the method of the embodiment of the present application adds the guidance of the neural network Lyapunov function, that is, the model-free method adopted by the embodiment of the present application is the Lyapunov neural network function, and the Lyapunov neural network function is used to achieve safe and stable control.

[0073] The neural network Lyapunov function is a scalar function that measures the environmental safety. It uses the fitting performance of the neural network to learn and fit the safe and stable set of the environment, and can learn the maximum safe domain of the complex nonlinear system to ensure that the system will not move to a dangerous state.

[0074] In the process of the embodiment of the present application learning the Lyapunov function, a scalar function that measures the environmental safety is found through the neural network Lyapunov function. This method assumes the candidate Lyapunov function as a multi-layer feedforward network with a tanh activation function. Since the function here needs to satisfy the Lyapunov asymptotic convergence condition, a smooth tanh function is required as the activation function instead of ReLu. The defined form of the Lyapunov risk function is as follows:

[0075]

[0076] The Lyapunov risk function here measures the degree of violation of the following Lyapunov condition. Among them, That is, the Lyapunov function is constructed through the neural network, and the training of the objective function ensures that the value of is positive, and the value of the Lie derivative is negative, and the value of is zero. Intuitively, the Lyapunov control design problem is to minimize the Lyapunov risk of the control. The purpose of this method is to find a function that can fit the global safety set so that the execution process of the entire strategy can be carried out under such a safety metric standard. As Figures 3 - 5 shown, the safety level set (ellipsoid) is expanded towards the true maximum candidate set The schematic diagram of the expansion. Forward simulate the system states in the gap G between the small ellipsoid and the large ellipsoid to determine the area of the safety level set that can be expanded, and iteratively make the safety set fit the shape of (as Figure 5 shown).

[0077] Through a neural network Lyapunov learner, the safety information of the environment can be learned. Utilizing the generalization ability of the neural network, as much global safety information of the environment as possible can be fitted. Safe states are labeled as positive values, and dangerous states are labeled as negative values. Before the action u is executed, safety guidance is provided for the system output to predict the probability of future danger occurrence, ensuring that only safe actions can be executed and deployed.

[0078] Some embodiments of the present application provide an attitude control system based on reinforcement learning. The system includes: a dynamic model configured to provide the current state or receive the input comprehensive safe action control amount; and an attitude control device based on reinforcement learning as described in the above embodiments.

[0079] The following combines Figure 6 Exemplarily elaborates on the attitude control method based on reinforcement learning provided by some embodiments of the present application. The attitude control method includes: S101, obtaining an initial action control amount through a model-free reinforcement learning algorithm and obtaining a historical safe action control amount through a historical CBF controller; S102, inputting the initial action control amount and the historical safe action control amount into the CBF controller to obtain the current safe action control amount; S103, determining the safety of the comprehensive safe action control amount through the Lyapunov algorithm, and sending the comprehensive safe action control amount to the dynamic model when it is confirmed that the comprehensive safe action control amount is safe.

[0080] In some embodiments of the present application, before obtaining the initial action control amount through the model-free reinforcement learning algorithm and obtaining the historical safe action control amount through the historical CBF controller, the control method further includes: initializing the parameters of the model-free reinforcement learning controller and the historical CBF controller.

[0081] In some embodiments of the present application, the control algorithms adopted by the model-free reinforcement learning include PPO or SAC.

[0082] In some embodiments of the present application, the parameters corresponding to the model-free reinforcement learning controller include: a memory pool D and an action pool A. Among them, the memory pool D is used for interaction data (s, u, r, s'), where s is the current state, u is the executed action, r is the immediate reward, and S' is the state at the next moment. The memory pool is used to learn and update the parameter θ of the model-free reinforcement learning controller. The action pool is used to store action data (s, u), and the action pool is used to perform supervised learning on the historical CBF controller.

[0083] The following combines Figure 7 Exemplarily elaborates on the attitude control method based on reinforcement learning provided by the embodiments of the present application, that is, the safety control algorithm based on CBF includes the following steps:

[0084] 1: Initialize each learner, memory pool, and action pool. For example, initialize the model-free reinforcement learning controller π RL , memory pool D, and action pool A.

[0085] 2: When k < maximum number of loops, update the model-free reinforcement learning controller, and use the memory pool D to update the control parameters of the model-free reinforcement learning controller .

[0086] 3. When k < maximum number of loops, update the historical CBF controller, specifically including: using the action pool A to update the control parameters of the fitted historical CBF controller .

[0087] 4: Confirm that the time t = 1, 2, … T (T is the maximum number of time steps per episode) is less than the maximum number of time steps.

[0088] 5. Obtain actions from the historical CBF controller and the model-free reinforcement learning controller: Obtain the action according to the current state s

[0089] 6. Solve the solution of the CBF controller using the above formula (3)

[0090] 7. Confirm that the Lyapunov condition is satisfied, that is, confirm that satisfies the Lyapunov safety condition

[0091] 8. Execute the control action to the dynamic model

[0092] 9. Otherwise: Return to step 10 of updating the model-free reinforcement learning controller. Store the action-state pair into the memory pool A

[0093] 11. Store the state reward into the memory pool D

[0094] 12. Collect the total reward for the episode

[0095] 13. Update the neural network Lyapunov model using formula (5) and the memory pool D

[0096] 14. k = k + 1

[0097] 15. Return the model-free controller, historical CBF controller, and CBF controller

[0098] It should be noted that the parameters of each controller are initialized. Among them, the model-free reinforcement learning controller can be any model-free reinforcement learning control algorithm (such as PPO or SAC). The memory pool D stores the interaction data (s, u, r, s'), where s is the current state, u is the executed action, r is the immediate reward defined for a specific control problem, and s' is the state at the next moment. This memory pool is used to learn and update the parameters θ of the model-free reinforcement learning controller. The action pool stores the action data (s, u), and this action pool is used to perform supervised learning on the historical CBF controller, and generally updates the control parameter φ by means of gradient descent. Then, the corresponding control quantities u are obtained by the two controllers according to the corresponding formulas respectively, and are input into the neural network Lyapunov controller for judgment. If it is judged to be safe, it is input into the dynamic model for execution, otherwise, the repeated steps are returned.

[0099] Some embodiments of the present application provide a basic framework for combining the model-free reinforcement learning method with the CBF control barrier function. By providing guidance on model danger information during the exploration process of the model-free method through the CBF and the neural network Lyapunov function, safe control is achieved. While minimizing interference in the policy learning process, it increases the learning of safety information and safety restrictions during action execution, which can improve the control stability of the system and the convergence speed of the algorithm.

[0100] Some embodiments of the present application improve the reinforcement learning attitude control algorithm based on the control barrier function, combine the model-free method with the control barrier function during the policy learning process, and improve the convergence speed and the interaction speed during the exploration process. Some embodiments of the present application provide a proof based on probabilistic reachability for this method, providing a theoretical guarantee for the safety of the control process. The migration algorithm framework of some embodiments of the present application improves the implementation performance of the algorithm.

[0101] In several embodiments provided by the present application, it should be understood that the disclosed apparatus and method can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the apparatus, method, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0102] In addition, in each embodiment of the present application, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0103] If the above functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0104] The above are only embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application. It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0105] As described above, these are only the specific implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0106] It should be noted that in this text, relative terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent in such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

Claims

1. A posture control device based on reinforcement learning, characterized in that The device includes: A model-free reinforcement learning controller configured to obtain an initial action control amount according to the current state; A historical CBF controller configured to obtain a historical safety action control amount according to the current state; A CBF controller configured to obtain a current safety action control amount according to the initial action and the historical safety action; A neural network Lyapunov learner configured to determine the safety of the comprehensive safety action control amount and provide the comprehensive safety action control amount to the dynamic model when it is confirmed that the comprehensive safety action control amount is safe, where the comprehensive safety action control amount is determined according to the initial action control amount, the historical safety action control amount, and the current safety action control amount.

2. The attitude control device according to claim 1, wherein The CBF controller is configured to obtain the current safety action control amount through the following formula (1); (1) Among them, is a slack variable under safe conditions, is a constant for punishing safety violations, and the obtained by solving the above formula is the current safety action control quantity, refers to the action vector which is the sum of the squares of all values of, and argmin represents the operation of taking the minimum value.

3. The attitude control device according to claim 2, wherein Solve the formula (1) through the following constraint conditions: and For all Among them, represents the transpose of the control parameter vector of the control barrier function, q is the control parameter, is the initial action control amount obtained by the model-free reinforcement learning controller according to the current state, is the safe action control amount obtained by the CBF controller during the j-th loop iteration, and are the mean and variance of the system state calculated by the Gaussian process, is the confidence interval parameter satisfying Gaussian regression, represents the parameter for adjusting the control barrier condition, refers to the upper bound of the control amount of the control dimension i, refers to the lower bound of the control amount of the control dimension i, is the transpose of the vector composed of the absolute values of all terms of the vector p.

4. The attitude control device according to claim 1, characterized in that, The historical CBF controller obtains the historical safety action control amount according to the following formula: Among them, to are calculated by the CBF controller each time a loop is performed. k is used to record the number of loops, and the value range of k is greater than or equal to 1.

5. The attitude control device according to any one of claims 1 to 4, characterized in that The comprehensive safety action control amount is the sum of the initial action control amount, the historical safety action control amount, and the current safety action control amount.

6. A posture control method based on reinforcement learning, characterized in that, The attitude control method includes: Obtaining an initial action control amount through a model-free reinforcement learning algorithm and obtaining a historical safety action control amount through a historical CBF controller; Inputting the initial action control amount and the historical safety action control amount into a CBF controller to obtain a current safety action control amount; Determining the safety of the comprehensive safety action control amount through a Lyapunov algorithm and sending the comprehensive safety action control amount to the dynamic model when it is confirmed that the comprehensive safety action control amount is safe; The comprehensive safety action control amount is the sum of the initial action control amount, the historical safety action control amount, and the current safety action control amount.

7. The attitude control method according to claim 6, characterized in that, Before obtaining the initial action control amount through the model-free reinforcement learning algorithm and obtaining the historical safety action control amount through the historical CBF controller, the control method further includes: Initializing the parameters of the model-free reinforcement learning controller and the historical CBF controller.

8. The attitude control method according to claim 7, characterized in that, The control algorithm adopted by the model-free reinforcement learning includes Proximal Policy Optimization (PPO) or Soft Actor-Critic (SAC) algorithm.

9. The attitude control method according to claim 7, characterized in that, The parameters corresponding to the model-free reinforcement learning controller include: a memory pool D and an action pool A. Among them, the memory pool D is used for interaction data , s is the current state, u is the executed action, r is the immediate reward, s’ is the state at the next moment, and the memory pool is used to learn and update the parameters of the model-free reinforcement learning controller , and the action pool is used to store action data , where the action pool is used for supervised learning of the historical CBF controller.

Citation Information

Patent Citations

  • Mobile phone satellite self-adaptive attitude control method

    CN106019950A

  • Automobile self-adaptive cruise control method under model uncertainty

    CN111897213A