Disturbance rejection controller design method, device and storage medium
By replacing the nonlinear error feedback control law with a neural network in the active disturbance rejection controller (ADRC), and combining deep reinforcement learning and delayed gradient descent strategies, the training of the ADRC is optimized. This solves the problems of numerous parameters in the ADRC and the limitations of the nonlinear error feedback control law in complex environments, thereby improving control performance.
Patent Information
- Application Number
- CN202211558443.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-06
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2042-12-06
AI Technical Summary
Active disturbance rejection controllers have a large number of internal parameters, and their internal nonlinear state error feedback control law has limitations in complex environments.
An active disturbance rejection controller (ADRC) model based on feedforward field weakening control is adopted. The nonlinear error feedback control law is replaced by a neural network. By combining Markov decision process and deep reinforcement learning, a dual-delay deep deterministic gradient descent strategy is designed to optimize the training process of the ADRC.
The number of internal parameters of the active disturbance rejection controller is reduced, which improves its control performance in complex environments, avoids the limitations of nonlinear state error feedback control law, and improves the overall performance of the controller.
Smart Images

Figure CN115903510B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electromechanical technology, specifically to a method, device, and storage medium for designing an active disturbance rejection controller. Background Technology
[0002] Currently, due to the increasing demand for air transport, heightened environmental awareness, and restrictions on carbon emissions, multi-electric and all-electric aircraft, which differ from traditional hybrid aircraft, are receiving considerable attention. The core component of multi-electric / all-electric aircraft is the electric motor, and permanent magnet synchronous motors (PMSMs) are suitable for use as aircraft motors due to their simple structure and high power density. Feedforward field weakening control is commonly used in multi-electric / all-electric aircraft motors to achieve a wide speed range, but its drawback is its reliance on parameter accuracy and susceptibility to disturbances in the motor's internal time-varying parameters. External disturbances from wind on aircraft motors are also another influencing factor.
[0003] In existing technologies, Active Disturbance Rejection Control (ADRC) is a popular motor control technology that effectively suppresses both internal and external disturbances in the control system and is widely used in various fields of automation control and industrial production. However, its drawbacks include the large number of internal parameters of the ADRC and the limitations of its internal nonlinear state error feedback control law in complex environments. Summary of the Invention
[0004] (a) Technical problems to be solved
[0005] To address the shortcomings of existing technologies, this invention provides a design method, device, and storage medium for an active disturbance rejection controller, which solves the problems of a large number of internal parameters in the active disturbance rejection controller and the limitations of its internal nonlinear state error feedback control law in complex environments.
[0006] (II) Technical Solution
[0007] To achieve the above objectives, the present invention provides the following technical solution:
[0008] Firstly, a method for designing an active disturbance rejection controller is provided, the method comprising:
[0009] A motor control system based on feedforward field weakening control was established, and an active disturbance rejection controller model was built and used as the speed loop of the motor control system based on feedforward field weakening control.
[0010] The nonlinear error feedback control law of the active disturbance rejection controller model is replaced by a neural network;
[0011] A deep reinforcement learning model is built by combining Markov decision processes with active disturbance rejection controller models.
[0012] Train a deep reinforcement learning agent. The deep reinforcement learning agent learns deep neural network optimization methods autonomously by continuously interacting with the environment based on the state and reward.
[0013] We design a dual-delay deep deterministic gradient descent strategy for training an active disturbance rejection controller (ADRC) model, thereby forming an ADRC.
[0014] Preferably, the active disturbance rejection controller model includes a tracking differentiator, an extended state observer, and a nonlinear state error feedback;
[0015] The tracking differentiator is used to track the input variable and arrange its transition process to speed up the control process and reduce overshoot.
[0016] The extended state observer is used to observe the output and analyze its differential state, and to estimate and analyze internal disturbances and external interferences of the controller.
[0017] The nonlinear state error feedback is used to perform nonlinear control on the error signals of the reference input and the extended state.
[0018] Preferably, the formula for the tracking differentiator is:
[0019]
[0020] In the formula: Z 11 It is the state of the input value after being arranged by TD. This represents the differential state of the input value, ω* is the input of the TD (Transmission Tunneling), which in this scheme is the given speed, e1 is the error between the input value and the input value state arranged by the TD, r is the gain coefficient, which is a constant, and the fal function represents the nonlinear function.
[0021]
[0022] In the formula: x is the input of the nonlinear function, α is the nonlinear factor coefficient, and δ is the filter factor coefficient. The factor coefficients of the fal function used by different components of the active disturbance rejection controller are different, and are distinguished by adding subscripts. Taking the tracking differentiator as an example, the parameters are α1 and δ1.
[0023] The extended state observer model is as follows:
[0024]
[0025] In the formula, ω is the actual velocity observed by ESO, and Z 21 It is the observation output after ESO processing, e2 is Z 21 The error of ω; Z 22The total disturbance, β, is the total measured quantity of both internal and external disturbances within the system. 21 ,β 22 These are the gain coefficients, It is the differential state of the observation error, b yes The compensation coefficient, u, is the output of the control system.
[0026] Preferably, the nonlinear error feedback control law of the active disturbance rejection controller model is replaced by a neural network, specifically as follows:
[0027] The input is the state Z, which tracks the output of the differentiator. 11 The observation Z output by the extended state observer 21 The error value is output as the reference torque T. ref The speed reference value is input into the active disturbance rejection controller (ADRC), and the tracking differentiator obtains the state variables of the input value based on this value. The extended state observer observes the actual speed value and processes the state variables of each order that yield the actual speed value. The input value state Z output by the tracking differentiator... 11 The observation Z output by the extended state observer 21 The error value is input into the deep neural network to obtain the corresponding reference torque output, which serves as the input required for field weakening control.
[0028] Preferably, the step of combining the Markov decision process with the active disturbance rejection controller model to build a deep reinforcement learning model specifically includes:
[0029] Using a motor control system based on feedforward field weakening control as the environment, and the motor's operating state as the state, rewards are set based on the steady-state performance of the speed and its anti-interference capability. Motor speed is set as the state for calculating the training effect of the deep reinforcement learning model. The motor control system based on feedforward field weakening control is set as the environment for interaction with the deep reinforcement learning model, obtaining the corresponding state after the neural network output. The output of the deep neural network is defined as the action, and the adjustment process of the action enables the deep reinforcement learning model to autonomously learn and explore strategies. Rewards are given by evaluating the new state given by the environment after taking an action. The initial reward function is set as follows:
[0030]
[0031] In the formula: e os It is the error between the observed and reference states of rotational speed, e l It is the speed error during sudden load changes, e ts It is the error between the observed torque and the reference torque; t ss and t sl These are the response time and the recovery time for suppressing disturbances, respectively, which can be calculated after the disturbance occurs; s 1-5These are the standardization coefficients because the dimensions of the optimization objectives are different; r 1-5 The weight coefficients can be changed according to different application requirements; the optimal neural network is obtained after the value of R converges to the minimum value.
[0032] Preferably, the training of the deep reinforcement learning agent involves the deep reinforcement learning agent autonomously learning deep neural network optimization methods based on states and rewards through continuous interaction with the environment, specifically as follows:
[0033] The deep reinforcement learning agent is trained using the actor-critic algorithm (AC). The AC algorithm comprises: an actor (constructing an actor policy network that outputs different actions under different conditions, and whose network can be adjusted and optimized by the deep reinforcement learning agent; in this invention, the deep neural network of the active disturbance rejection controller is consistent with the actor network); and a critic (using the reward value R obtained after an action as the critic's evaluation criterion); the deep reinforcement learning agent prompts the critic to evaluate the action and its effect, and changes the actor's action policy based on the critic's evaluation. The agent is trained through continuous interaction and repetition of the actor-critic policy process until the reward value converges and the agent finds the optimal action policy.
[0034] After the actor takes an action, the intelligent system adds noise to the action to prevent the actor's behavior strategy from falling into a local optimum.
[0035] a t =μ(s) t |θ μ )+Noise
[0036] In the formula: μ(s|θ) μ ) is the action strategy network, θ μ It is a hyperparameter of network μ; s t and a t represents the state and action at time t, respectively. The current action is selected by the agent based on the current state and policy; Noise is the noise added to the action.
[0037] Reward r t and the next state s t+1 Data s can be obtained after the action is completed. t ,a t ,r t ,s t+1 The dataset will be stored in a database, and some datasets randomly selected from the database will be used to train the method. The commenter network will evaluate the previous action based on the new state and reward generated.
[0038] Preferably, the design of the dual-delay deep deterministic gradient descent strategy method, used for training the active disturbance rejection controller model to form the active disturbance rejection controller, specifically includes:
[0039] Q(s,a|θ Q ) is used as the evaluation network for commentators; by updating θ π and θ Q To optimize the method;
[0040] Two commentator network parameter update systems are constructed. Each time the commentator network parameters are updated, parameters that can obtain smaller evaluation values are used to alleviate overestimation.
[0041] Two gradient descent networks were constructed. The actual network was updated in real time, while the target network to be used was updated with a delay. That is, a relationship was established between the network and the actor π(s|θ). π ) and the evaluator Q(s,a|θ Q The two policy networks have identical structures, but their parameters are represented as π'(s|θ) π’ ) and Q'(s,π'(s|θ) π’ )|θ Q’ The network of actors and evaluators is called the target network of actors and evaluators;
[0042] The update formula is as follows: The target network and the update rate coefficient τ≤1 are used to slowly update the policy network parameters.
[0043]
[0044] The evaluation value y of the target Q network calculated by two reviewers i And the loss L between the target network and the original network, expressed by the formula below. Obtain;
[0045] With minimizing L as the optimization objective, the Q network parameters θ are optimized. Q Optimize according to the formula; Descent gradient update θ of a randomly sampled dataset π To maximize the evaluation value of Q;
[0046] Each correction will be penalized or rewarded based on its effect, denoted as R. a Through formula Adjustments are made to the formula R = -(α) Obs R Obs +α a R a The final total evaluation value is calculated.
[0047] In the formula: α Obs α a These are the corresponding weights for observation rewards and adjustment rewards.
[0048] Secondly, a device is provided, comprising:
[0049] One or more processors;
[0050] Memory, used to store one or more programs.
[0051] When the one or more programs are executed by the one or more processors, the one or more processors execute the active disturbance rejection controller design method.
[0052] Thirdly, a computer-readable storage medium is provided that stores a computer program, which, when executed by a processor, implements the aforementioned active disturbance rejection controller design method.
[0053] (III) Beneficial Effects
[0054] Compared with the prior art, the present invention provides a design method, device and storage medium for an active disturbance rejection controller. The present invention has the following significant advantages: by designing and optimizing the novel active disturbance rejection controller, the number of internal parameters of the active disturbance rejection control is reduced, the limitations of the original nonlinear state error feedback control law in the face of complex situations are avoided, and the control performance of the active disturbance rejection controller is improved. Attached Figure Description
[0055] Figure 1 This is a schematic diagram of the design method for the active disturbance rejection controller of the present invention;
[0056] Figure 2 This is a control structure diagram in an embodiment of the present invention;
[0057] Figure 3 This is a schematic diagram of the active disturbance rejection controller structure in an embodiment of the present invention;
[0058] Figure 4 This is a schematic diagram of the speed loop active disturbance rejection control structure in the feedforward field weakening control system for a multi-electric aircraft motor in an embodiment of the present invention.
[0059] Figure 5 This is a schematic diagram of the active disturbance rejection controller and deep neural network optimization structure based on deep reinforcement learning in an embodiment of the present invention;
[0060] Figure 6 This is a schematic diagram of the structure of the dual-delay deep deterministic gradient strategy algorithm in an embodiment of the present invention. Detailed Implementation
[0061] The technical solutions in the embodiments of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0062] Please see Figure 1 One embodiment of the present invention provides a design method for an active disturbance rejection controller (ADRC) based on deep reinforcement learning, comprising: establishing a velocity loop ADRC model and establishing a deep neural network to replace the nonlinear error feedback control law; constructing a new ADRC deep reinforcement learning model by combining a Markov decision process; using an actor-commentator algorithm to enable the deep reinforcement learning agent to continuously interact with the environment and autonomously learn the deep neural network optimization method; designing a dual-delay deep deterministic gradient descent strategy to converge the deep neural network optimization process, thereby completing the training of the new ADRC optimization model based on deep reinforcement learning. The control structure diagram is shown below. Figure 2 As shown. The main steps are as follows:
[0063] Step 1: Establish a velocity loop active disturbance rejection controller model; establish a deep neural network to replace the nonlinear state error feedback control law.
[0064] Specifically, a traditional first-order active disturbance rejection controller (ADRC) consists of three parts: a tracking differentiator (TD), an extended state observer (ESO), and a nonlinear state error feedback. The tracking differentiator (TD) tracks the input variable and manages its transient response to accelerate the control process and reduce overshoot. The extended state observer (ESO) observes the output and analyzes its differential state, estimating and analyzing internal disturbances and external interferences. The nonlinear state error feedback (NLSEF) primarily performs nonlinear control on the error signals from the reference input and the extended state.
[0065] The formula for a first-order tracking differentiator is:
[0066]
[0067] In the formula: Z 11 It is the state of the input value after being arranged by TD. This represents the differential state of the input value. ω* is the input to the TD, which in this scheme is the given speed; e1 is the error between the input value and the input value state arranged by the TD; r is the gain coefficient, which is a constant; and the fal function represents the nonlinear function.
[0068]
[0069] In the formula: x is the input of the nonlinear function, α is the nonlinear factor coefficient, and δ is the filter factor coefficient. The factor coefficients of the fal function used in different components of the active disturbance rejection controller are different, and are distinguished by adding subscripts. Taking the tracking differentiator as an example, the parameters are α1 and δ1.
[0070] The first-order extended state observer model is as follows:
[0071]
[0072] In the formula: ω is the actual velocity observed by ESO, Z 21 It is the observation output after ESO processing, e2 is Z 21 The error of ω; Z 22 It is the total measurement of both internal and external disturbances within the system, referred to as total disturbance.
[0073] Movement, β 21 ,β 22 These are the gain coefficients, It is the differential state of the observation error, b yes The compensation coefficient, u, is the output of the control system.
[0074] In this invention, the nonlinear state error feedback part of the active disturbance rejection control is replaced by a deep neural network, whose input is the input value Z of the tracking differentiator output. 11 The observation Z output by the extended state observer 21 The error value is output as the reference torque T. ref The novel active disturbance rejection controller structure is as follows: Figure 3 As shown. The structure of the novel active disturbance rejection controller applied to the feedforward field weakening control system of a multi-electric aircraft motor is as follows. Figure 4 As shown. The speed reference value is input into the active disturbance rejection controller (ADRC), and the tracking differentiator obtains the state variables of the input value based on it. The extended state observer observes the actual speed value and processes the state variables of each order that yield the actual speed value. The input value state Z output by the tracking differentiator is shown. 11 The observation Z output by the extended state observer 21 The error value is input into the deep neural network to obtain the corresponding reference torque output, which serves as the input required for field weakening control.
[0075] Step 2: Combine Markov decision processes with active disturbance rejection control environments as the basis for building deep reinforcement learning models.
[0076] Specifically, the motor speed is set as the state, used to calculate the training effect of the deep reinforcement learning model; the motor control system is set as the environment, used to interact with the deep reinforcement learning model and obtain the corresponding state after the neural network outputs; the output of the deep neural network is defined as the action, and the adjustment process of the action allows the algorithm agent to autonomously learn exploration strategies; the reward is used to evaluate the new state given by the environment after taking the action. The initial reward function is set as follows:
[0077]
[0078] In the formula: e os It is the error between the observed and reference states of rotational speed, e l It is the speed error during sudden load changes, e ts It is the error between the observed torque and the reference torque; t ss and t sl These are the response time and the recovery time for suppressing disturbances, respectively, which can be calculated after the disturbance occurs; s 1-5 These are the standardization coefficients because the dimensions of the optimization objectives are different; r 1-5 These weights can be adjusted according to the different needs of the application. The optimal neural network can be obtained after the value of R converges to its minimum.
[0079] Step 3: The process of training the agent using the Actor-Critic algorithm.
[0080] Specifically, the Actor-Critic AC algorithm comprises: an Actor, which constructs an actor policy network to output different actions under different conditions, and the Actor network can be adjusted and optimized through a deep reinforcement learning agent. In this invention, the deep neural network of the active disturbance rejection controller is consistent with the actor network; and a Critic, which uses the reward value R obtained after the action as the criterion for evaluation. The agent can obtain evaluations from the criterion based on the action and its effect, and change the agent's action policy based on the criterion's evaluation. The agent is trained through continuous interaction and repetition of the Actor-Critic policy process until the reward value converges and the agent finds the optimal action policy.
[0081] After an actor takes an action, the intelligent system adds noise to the action to prevent the actor's behavioral strategy from falling into a local optimum.
[0082] a t =μ(s) t |θ μ )+Noise (5)
[0083] In the formula: μ(s|θ) μ ) is the action strategy network, θ μ It is a hyperparameter of network μ. t and a t Let represent the state and action at time t, respectively. The current action is selected by the agent based on the current state and policy. Noise is the noise added to the action.
[0084] Reward r t and the next state s t+1 Data (s) can be obtained after the action is completed. t ,a t ,r t ,s t+1 The dataset will be stored in a database, and some randomly selected datasets from the database will be used to train the proposed method TD3-ADRC. The reviewer network will evaluate the previous action based on the generated new state and reward. The proposed deep reinforcement learning-based active disturbance rejection controller and deep neural network optimization structure are as follows: Figure 5 As shown.
[0085] Step 4: The Actor-Critic algorithm is trained and converged using a dual-delay deep deterministic gradient descent strategy, completing the optimized design of the novel active disturbance rejection controller. After each round of learning by the agent, the target network is used to perform gradient descent on the Actor-Critic algorithm, ensuring convergence of the reward value and the actor network optimization. The structure of the dual-delay deep deterministic gradient strategy algorithm is as follows: Figure 6 As shown.
[0086] Specifically, Q(s,a|θ Q ) was used as the evaluation network for commentators. By updating θ π and θ Q To optimize the proposed method, two commentator network parameter update systems were constructed. Each time the commentator network parameters were updated, parameters that yielded smaller evaluation values were used to mitigate overestimation. Two gradient descent networks were constructed: the actual network was updated in real-time, while the target network to be used was updated with a delay. This involved establishing a connection with the actor π(s|θ). π ) and the evaluator Q(s,a|θ Q The two policy networks have identical structures, but their parameters are represented as π'(s|θ) π’ ) and Q'(s,π'(s|θ) π’ )|θ Q’ The network consisting of the actor and the evaluator is called the target network. To allow the agent sufficient learning time, the policy network parameters are slowly updated through the target network and the update rate coefficient τ≤1. The update formula is:
[0087]
[0088] The evaluation value y of the target Q network calculated by two reviewers i The loss L between the target network and the original network can be obtained through equation (9). With minimizing L as the optimization objective, the Q network parameters θ are optimized. Q Optimization is performed. θ is updated based on the descent gradient of the randomly sampled dataset in equation (10). π To maximize the evaluation value of Q.
[0089]
[0090]
[0091] Each correction will be penalized or rewarded based on its effect, denoted as R. a The final total evaluation value can be adjusted using equation (11). Equation (12) is used to calculate the final total evaluation value.
[0092]
[0093] R = -(α) Obs R Obs +α a R a (12)
[0094] In the formula: α Obs α a These are the corresponding weights for observation rewards and adjustment rewards.
[0095] As another embodiment of the present invention, an apparatus is provided, comprising:
[0096] One or more processors;
[0097] Memory, used to store one or more programs.
[0098] When the one or more programs are executed by the one or more processors, the one or more processors execute a self-disturbance rejection controller design method as described in the above embodiments.
[0099] As another embodiment of the present invention, a computer-readable storage medium storing a computer program is provided, which, when executed by a processor, implements a method for designing an active disturbance rejection controller as described in the above embodiments.
[0100] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
Claims
1. A method for designing an active disturbance rejection controller, characterized in that, The method includes: A motor control system based on feedforward field weakening control was established, and an active disturbance rejection control system model was built and used as the speed loop of the motor control system based on feedforward field weakening control. An active disturbance rejection control system model is formed by replacing the nonlinear error feedback control law of the active disturbance rejection controller model with a neural network; A deep reinforcement learning model is built by combining Markov decision processes with active disturbance rejection controller models. Train a deep reinforcement learning agent. The deep reinforcement learning agent learns deep neural network optimization methods autonomously by continuously interacting with the environment based on the state and reward. We design a dual-delay deep deterministic gradient descent strategy for training an active disturbance rejection controller (ADRC) model, thereby forming an ADRC. The replacement of the nonlinear error feedback control law of the active disturbance rejection controller model with a neural network is specifically as follows: The input is the state of the input value that tracks the output of the differentiator. Z 11 Observations output by the extended state observer Z 21 The error value is output as the reference torque. T ref The speed reference value is input into the active disturbance rejection controller (ADRC), and the tracking differentiator obtains the state variables of the input value based on this value. The extended state observer observes the actual speed value and processes the state variables of each order that yield the actual speed value. The input value state is output by the tracking differentiator. Z 11 Observations output by the extended state observer Z 21 The error value is input into the deep neural network to obtain the corresponding reference torque output, which serves as the input required for field weakening control.
2. The active disturbance rejection controller design method according to claim 1, characterized in that: The active disturbance rejection controller model includes a tracking differentiator, an extended state observer, and a nonlinear state error feedback. The tracking differentiator is used to track the input variable and arrange its transition process to speed up the control process and reduce overshoot. The extended state observer is used to observe the output and analyze its differential state, and to estimate and analyze internal disturbances and external interferences of the controller. The nonlinear state error feedback is used to perform nonlinear control on the error signals of the reference input and the extended state.
3. The active disturbance rejection controller design method according to claim 2, characterized in that: The formula for the tracking differentiator is: In the formula: Z 11 It is the state of the input value after being arranged by TD. Z· 11 It is the differential state of the input value state. ω The input to TD is the given speed in this scheme. e 1 represents the error between the input value and the input state arranged by TD, r is the gain coefficient, which is a constant, and the fal function represents a nonlinear function. In the formula: x It is the input of a nonlinear function. α These are nonlinear factor coefficients. δ These are the filter factor coefficients. Different components of the active disturbance rejection controller use different factor coefficients in their fal functions, distinguished by subscripts. Taking the tracking differentiator as an example, the parameters are... α 1 and δ 1 ; The extended state observer model is as follows: In the formula, ω This is the actual velocity observed by ESO. Z 21 These are the observations output after ESO processing. e 2 is Z 21 and ω The error; Z 22 The total disturbance is the sum of all internal and external disturbances observed in the system. β 21 , β 22 These are the gain coefficients, Z· 22 It is the differential state of the observation error. b yes Compensation coefficient, u It is the output of the control system.
4. The active disturbance rejection controller design method according to claim 1, characterized in that: The method of combining Markov decision processes with active disturbance rejection controller models to build deep reinforcement learning models specifically includes: Using a motor control system based on feedforward field weakening control as the environment, and the motor's operating state as the state, rewards are set based on the steady-state performance of the speed and its anti-interference capability. Motor speed is set as the state for calculating the training effect of the deep reinforcement learning model. The motor control system based on feedforward field weakening control is set as the environment for interaction with the deep reinforcement learning model, obtaining the corresponding state after the neural network output. The output of the deep neural network is defined as the action, and the adjustment process of the action enables the deep reinforcement learning model to autonomously learn and explore strategies. Rewards are given by evaluating the new state given by the environment after taking an action. The initial reward function is set as follows: In the formula: e os It is the error between the observed and reference states of rotational speed. e l It is the speed error during sudden load changes. e ts It is the error between the observed torque and the reference torque; t ss and t sl These are the response time and the recovery time for suppressing disturbances, respectively, which can be calculated after the disturbance occurs; s 1-5 It is a per-unit coefficient because the dimensions of the optimization objectives are different; r 1-5 These are weighting coefficients that can be changed according to different application requirements; R Once the value converges to the minimum, the optimal neural network is obtained.
5. The design method of an active disturbance rejection controller according to claim 1, characterized in that: The training of the deep reinforcement learning agent involves the agent autonomously learning deep neural network optimization methods based on its state and rewards through continuous interaction with the environment. Specifically: The deep reinforcement learning agent is trained using the actor-commentator algorithm. The actor-commentator AC algorithm includes: an actor, which builds an actor policy network that outputs different actions under different conditions. The actor network can be adjusted and optimized by the deep reinforcement learning agent. The deep neural network of the active disturbance rejection controller is consistent with the actor network. The commentator, Critic, assigns the reward value obtained after the action. R Used as the basis for the reviewer's evaluation; the deep reinforcement learning agent makes the reviewer evaluate based on the action and the effect, and changes the action strategy of the agent according to the reviewer's evaluation. The agent is trained by continuously interacting and repeating the agent-reviewer strategy process until the reward value converges and the agent finds the optimal action strategy. After the actor takes an action, the intelligent system adds noise to the action to prevent the actor's behavior strategy from falling into a local optimum. In the formula: μ (s| θ μ ) is an action strategy network, θ μ It is the internet μ Hyperparameters; s t and a t They represent the current t The state and action at any given moment; the current action is selected by the agent based on the current state and policy; noise is added to the action. award r t and the next state s t+1 Data can be obtained after the action is completed. s t , a t , r t , s t+1 The data will be stored in a database, and some datasets randomly selected from the database will be used to train the method. The commenter network will evaluate the previous action based on the new state and reward generated.
6. The active disturbance rejection controller design method according to claim 1, characterized in that: The aforementioned dual-delay deep deterministic gradient descent strategy method, used for training the active disturbance rejection controller (ADRC) model to form the ADRC, specifically includes: Q ( s , a | θ Q It is used as an evaluation network for reviewers; through updates θ π and θ Q To optimize the method; Two commentator network parameter update systems are constructed. Each time the commentator network parameters are updated, parameters that can obtain smaller evaluation values are used to alleviate overestimation. Two gradient descent networks were constructed. The actual network was updated in real time, while the target network to be used was updated with a delay. This establishes a connection between the network and the actor. π (s| θ π ) and evaluators Q ( s , a | θ Q The two policy networks have identical structures, but their parameters are expressed as follows: π’ (s| θ π’ )and Q’ ( s , π’ (s| θ π’ )| θ Q’ The network of actors and evaluators is called the target network of actors and evaluators; The update formula is as follows: The target network and the update rate coefficient τ≤1 are used to slowly update the policy network parameters. The target calculated by two commentators Q Network evaluation value y i And the loss between the target network and the original network. L Through the formula below Obtain; With the smallest L To optimize the objective, Q Network parameters θ Q Optimize according to the formula; Descent gradient update of a randomly sampled dataset θ π To maximize the evaluation value of Q; Each corrective action will be punished or rewarded based on its effect. R a Through formula Adjustments were made to the formula. Calculate the final total assessment value; In the formula: α Obs , α a These are the corresponding weights for observation rewards and adjustment rewards.
7. A device, characterized in that, include: One or more processors; Memory, used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors perform an active disturbance rejection controller design method as described in any one of claims 1-6.
8. A computer-readable storage medium storing a computer program, characterized in that, When executed by the processor, the program implements a method for designing an active disturbance rejection controller as described in any one of claims 1-6.
Citation Information
Patent Citations
Driving joint dynamic error prediction and compensation system and method based on neural network
CN114310911A
Active-disturbance-rejection controller parameter optimization method based on deep reinforcement learning
CN115097736A