A Parameter Optimization Method for Active Disturbance Rejection Controller Based on Deep Reinforcement Learning

Through deep reinforcement learning, the parameters of self-immune controllers are optimized, which solves the problem of long-term parameter adjustment and achieves efficient self-immune controller performance improvement.

CN115097736BActive Publication Date: 2025-07-08SOUTHEAST UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210955313.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-10
Publication Date
2025-07-08
Estimated Expiration
2042-08-10

AI Technical Summary

Technical Problem

The self-immune controller has a large number of parameters and strong coupling, which makes the parameter adjustment time-consuming and low efficiency, and cannot exert the optimal performance of the controller.

Method used

The self-immune controller model is established based on deep reinforcement learning, combined with the Markov process to build a deep reinforcement learning model, and the Actor-Critic algorithm and deep deterministic strategy gradient method are used to optimize the self-immune controller parameters.

Benefits of technology

Improve parameter adjustment efficiency, save time, avoid local optimization, and give full play to the performance of self-immune controllers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115097736B_ABST
    Figure CN115097736B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for optimizing the parameters of an active disturbance rejection controller based on deep reinforcement learning, belonging to the field of mechatronics. Among them, the method includes: taking the parameters of the permanent magnet synchronous motor active disturbance rejection controller as the optimization objective; building a deep reinforcement learning model, taking the control system as the environment, taking the motor speed as the state, setting rewards based on the smoothness of the speed and the anti-interference ability, using the Actor-Critic algorithm to train the intelligent agent to select optimization actions according to the environment and state, improving the optimization actions according to the size of the rewards obtained after the actions, enabling the intelligent agent to autonomously learn the optimization of the active disturbance rejection parameters; designing a deep deterministic policy gradient method to make the parameter optimization process converge, completing the training of the parameter optimization model based on deep reinforcement learning, and obtaining the optimal parameters. By adopting the above scheme, the optimal parameters of the active disturbance rejection controller can be obtained with the minimum manual debugging cost, thereby solving the problems that the active disturbance rejection controller has many parameters, strong coupling, low sensitivity, and is difficult to debug to make it work in the optimal state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electromechanics, and in particular, to a method for optimizing the parameters of an active disturbance rejection controller based on deep reinforcement learning. Background Art

[0002] As a popular motor control technology, the active disturbance rejection controller has been widely studied and applied in various fields of automatic control and industrial production.

[0003] In the prior art, due to the large number of internal parameters, strong coupling, and low sensitivity of the active disturbance rejection controller, the parameters of the active disturbance rejection controller are generally adjusted manually and by experience, which is time-consuming and inefficient, and cannot exert the optimal performance of the controller. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention proposes a method for optimizing the parameters of an active disturbance rejection controller based on deep reinforcement learning. To achieve the self-adaptive adjustment of the parameters of the active disturbance rejection controller according to the models of the motor and the active disturbance rejection controller, and to achieve the optimal effect of the parameters of the active disturbance rejection controller, thereby solving the problems that the parameters of the active disturbance rejection controller are difficult to adjust, the traditional method is inefficient and cannot guarantee the optimal parameters.

[0005] The object of the present invention can be achieved by the following technical solutions: A method for optimizing the parameters of an active disturbance rejection controller based on deep reinforcement learning, comprising: establishing a speed-loop active disturbance rejection controller model and setting a parameter optimization target; building a deep reinforcement learning model for the parameters of the active disturbance rejection controller in combination with the Markov process; using the Actor-Critic (AC) algorithm to enable the intelligent agent to continuously interact with the environment and autonomously learn the method for optimizing the active disturbance rejection control parameters; designing a deep deterministic policy gradient method to make the parameter optimization process converge and complete the training of the parameter optimization model of the active disturbance rejection controller based on deep reinforcement learning. The main steps are as follows:

[0006] Step 1: Establish a speed-loop active disturbance rejection controller model; select the parameters and the number to be optimized according to the actual model and set the optimization target.

[0007] Specifically, the active disturbance rejection controller consists of three parts, namely the tracking differentiator, the extended state observer, and the nonlinear state error feedback. The Tracking Differentiator (TD) can track the input signal and arrange the transition state to enable fast and overshoot-free control. The Extended State Observer (ESO) is used to observe the output and its differential components of each order. At the same time, the internal disturbances and external variables of the system are regarded as the total disturbance, and the total disturbance is observed and estimated. The Nonlinear State Error Feedback (NLSEF) mainly controls the nonlinear combination of the reference input and the error signal of the extended state, and compensates for the observed total disturbance.

[0008] The controller of the speed loop is a first-order model. Its first-order tracking differentiator model is:

[0009]

[0010] In the formula: Z 11 represents the input value state processed by the tracking differentiator, represents the differential component of the input value state processed by the tracking differentiator. ω* represents the given speed, e1 is the error between the given value and the tracking value, r is the gain coefficient, and the fal function is a nonlinear function, and its expression is:

[0011]

[0012] In the formula: x is the function input, α is a constant representing the nonlinear factor, and δ is a constant affecting the filtering effect. The α and δ in the fal functions used by different modules of the active disturbance rejection controller are different, and are distinguished by subscripts. For example, in the tracking differentiator, the parameters are α1 and δ1.

[0013] The first-order extended state observer model is:

[0014]

[0015] In the formula: ω represents the actual speed collected, Z 21 represents the observed value of the system output by the extended state observer, and e2 is the error between the two; Z 22 represents the observed value of the total disturbance, β 21 , β 22 is the gain coefficient, represents the differential form of the error observed value, b represents the compensation coefficient, and u represents the output of the nonlinear state error feedback.

[0016] The first-order nonlinear state error feedback model is:

[0017]

[0018] Where: Z 11 represents the state of the input value after being processed by the tracking differentiator, and Z 21 represents the observed quantity of the system output by the extended state observer. e is the difference between Z 11 and Z 21 , and β3 is the gain coefficient.

[0019] Input the set speed value into the active disturbance rejection controller. After passing through the tracking differentiator, the state quantity of the given speed is obtained; input the actual speed value collected into the extended state observer to obtain the state quantity of the actual speed and the observed total error value; subtract the given speed state quantity from the actual speed state quantity, and the difference passes through the nonlinear state error feedback to obtain the initial output value u0. Add the compensation for the total disturbance observed by the extended state observer to obtain the final output value u, which is the initial given value of the torque required by the field-weakening system.

[0020] It can be seen that there are 12 parameters in the active disturbance rejection controller that need to be optimized and adjusted:

[0021]

[0022] Step 2: Build a deep reinforcement learning model for the parameters of the active disturbance rejection controller in combination with the Markov decision process. Take the current motor control system as the environment and the motor speed curve as the state, and set the reward based on the smoothness of the speed and the anti-interference ability.

[0023] Specifically, the motor speed value is set as: State, which is used to evaluate the parameter optimization effect; the motor control environment is set as: Environment, which is responsible for giving the real-time state after the parameters change; Action, learning to adjust and explore the 12 parameters; Reward, which is evaluated based on the new state given by the environment after the action. The reward function used is:

[0024] R = r1e os / s1 + r2t rs / s2 + r3t rl / s3 + r4|e l | / s4 (6)

[0025] Where: e os , e l , t rs and t rl are the optimization objectives. e os and e l are the speed errors during startup and sudden load transition respectively, and t rs and t rlThey are the start-up time and the time for the speed to return to normal after a sudden load, respectively. s1, s2, s3, and s4 are normalization coefficients because the dimensions between the optimization objectives are different; r1, r2, r3, and r4 are the weight coefficients of the four optimization objectives, which are changed according to the different requirements of the application environment. When the final evaluation value R is the smallest, the best optimization result will be obtained.

[0026] Step 3: Use the AC algorithm to train the agent to select different optimization actions according to the environment and state, and improve the optimization actions based on the size of the rewards obtained after the actions, so that the agent continuously interacts with the environment and autonomously learns the method of active disturbance rejection control parameter optimization.

[0027] Specifically, the classic Actor-Critic (AC) algorithm structure includes: an Actor that can learn and construct a policy network and select different actions according to the network in different states; a Critic that can evaluate the value of the actions of the optimized policy network. The agent evaluates according to the reward value and decides how the Actor should act based on the evaluation, that is, the direction and amplitude of the parameter adjustment. After this AC round, a new round of interaction and learning is carried out until the adjustment and optimization of the parameters converge. In addition, after the Actor takes an optimization action, the agent will add noise to this action, which can simulate the interference of the system and make the result more accurate.

[0028] The Actor structure network can be expressed as μ(s|θ μ ), θ μ are the internal parameters of the policy network μ. The current state and action are represented as s t and a t . respectively. The agent takes action a t according to μ(s|θ μ ) based on s t . The added noise can be expressed as:

[0029] a t = μ(s t |θ μ ) + Noise(7)

[0030] When an action is completed, the reward r t and the next state s t+1 are fed back, and the data (s t , a t , r t , s t+1 ) will be stored in the database. {(s i , a i , r i , s i+1 )|i = 1, 2, …, N} then contains some data sets randomly selected from the database for training. Subsequently, by Q(s, a|θQ ) The Critic evaluation network represented will evaluate based on the previous step's s and a.

[0031] Step 4: Design the deep deterministic policy gradient method to make the parameter optimization process converge and complete the training of the auto-disturbance rejection controller parameter optimization model based on deep reinforcement learning;

[0032] Specifically, the deterministic policy gradient method is used to make the critic converge and update the network parameters. Two gradient descent networks are constructed: the actual network θ that is updated in real time μ , and the target network θ that is updated with a delay and will be used finally Q . The optimization update of the algorithm is achieved by updating θ μ and θ Q . To give the algorithm sufficient learning time, μ(s|θ μ ) and Q(s,a|θ Q ) are not directly used. By using μ’(s|θ μ’ ) and Q’(s,μ’(s|θ μ’ )|θ Q’ ), the process of parameter update is split and amplified. The target network has the same structure as the actual network, and the target network slowly tracks the parameters of the actual network at the parameter update rate τ, τ ≤ 1:

[0033]

[0034] The evaluation value of the target network is y i , and the loss between the target network and the actual network is L, and the two can be calculated according to (9). Using the minimum L as the optimization goal can optimize the Q network parameter θ Q . Using the negative average value J of the network in (10) as the optimization goal to update θ μ can maximize the evaluation value of the actual network Critic.

[0035]

[0036] The action setting allows the algorithm to correct the ADRC parameters that need to be optimized and uses (11) to achieve the normalization, restoration, and correction of the parameters. (12) is used to keep the optimized parameters within the feasible range.

[0037]

[0038] In the formula: θ max , θ min θ i , They are the upper limit, lower limit, original value, and regularized value of the i-th generation parameters respectively. (13) is used to evaluate and process the error between the actual value and the given value after setting the optimization target. (14) is used to punish and reward parameter correction, and (15) is used as the final evaluation of the optimization.

[0039] R Obs = error evaluation (13)

[0040]

[0041] R = -(α Obs R Obs + α θ R θ )(15)

[0042] Where: α Obs , α θ are the corresponding weights of the observation reward and the parameter correction reward and punishment respectively.

[0043] Advantages of the present invention: Compared with the prior art, the present invention has the following remarkable advantages: By adaptively adjusting the parameters of the active disturbance rejection controller, the parameter adjustment efficiency can be improved, the parameter adjustment time can be saved, the parameters can be prevented from falling into local optimization, and the performance of the active disturbance rejection controller can be fully exerted. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The present invention will be further described below with reference to the drawings.

[0045] Figure 1 is a control system diagram of the parameter optimization of the speed loop active disturbance rejection controller for a permanent magnet motor based on deep reinforcement learning in an embodiment of the present invention;

[0046] Figure 2 is a schematic structural diagram of the active disturbance rejection controller in an embodiment of the present invention.

[0047] Figure 3 is a schematic structural diagram of the speed loop active disturbance rejection control for a permanent magnet motor in an embodiment of the present invention.

[0048] Figure 4 is a schematic structural diagram of the active disturbance rejection parameter optimization framework based on deep reinforcement learning in an embodiment of the present invention.

[0049] Figure 5 is a schematic structural diagram of the deep deterministic gradient policy and the AC algorithm in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0051] An anti-disturbance controller parameter optimization method based on deep reinforcement learning is provided in an embodiment of the present invention, including: establishing a speed-loop anti-disturbance controller model and setting a parameter optimization target; building a deep reinforcement learning model for the anti-disturbance controller parameters in combination with a Markov process; using the AC algorithm to enable the intelligent agent to continuously interact with the environment and autonomously learn the anti-disturbance control parameter optimization method; designing a deep deterministic policy gradient method to make the parameter optimization process converge and complete the training of the anti-disturbance controller parameter optimization model based on deep reinforcement learning. The specific optimization process and control structure can be referred to Figure 1 as shown. The main steps are as follows:

[0052] Step 1: Establish a speed-loop anti-disturbance controller model; select the parameters and the number to be optimized according to the actual model and set the optimization target.

[0053] Specifically, referring to Figure 2 , the anti-disturbance controller includes three parts, namely a tracking differentiator, an extended state observer, and a nonlinear state error feedback. The Tracking Differentiator (TD) can track the input signal and arrange the transition state to make the control fast and without overshoot. The Extended State Observer (ESO) is used to observe the output and its various order differential components, and at the same time regard the internal disturbance and external variables of the system as the total disturbance to observe and estimate the total disturbance. The Nonlinear State Error Feedback (NLSEF) mainly controls the nonlinear combination of the reference input and the error signal of the extended state, and at the same time compensates for the observed total disturbance.

[0054] The controller of the speed loop is a first-order model. Its first-order tracking differentiator model is:

[0055]

[0056] In the formula: Z 11 represents the input value state after being processed by the tracking differentiator, represents the differential component of the input value state after being processed by the tracking differentiator. ω* represents the given speed, r is the gain coefficient, and the fal function is a nonlinear function, and its expression is:

[0057]

[0058] The first-order extended state observer model is as follows:

[0059]

[0060] Where: ω represents the actual speed collected, Z 21 represents the observed quantity of the system output by the extended state observer, Z 22 represents the observed quantity of the total disturbance, β 21 , β 22 is the gain coefficient, Z 22 represents the differential form of the error observed quantity, and b represents the compensation coefficient.

[0061] The first-order non-linear state error feedback model is as follows:

[0062]

[0063] Where: u represents the output of the non-linear state error feedback.

[0064] Input the set speed value into the active disturbance rejection controller, and obtain the state quantity of the given speed through the tracking differentiator; input the actual speed value collected into the extended state observer to obtain the state quantity of the actual speed and the observed total error value; subtract the given speed state quantity from the actual speed state quantity, and the difference passes through the non-linear state error feedback to obtain the initial output value u0, and add the compensation for the total disturbance observed by the extended state observer to obtain the final output value u, which is the initial given value of the torque required by the field-weakening system.

[0065] At this time, the active disturbance rejection control system of the permanent magnet motor speed loop is built, and reference can be made to Figure 3 .

[0066] It can be seen that there are 12 parameters in the active disturbance rejection controller that need to be optimized and adjusted:

[0067]

[0068] Step 2: Build a deep reinforcement learning model for the parameters of the active disturbance rejection controller in combination with the Markov decision process. Take the current motor control system as the environment, take the motor speed curve as the state, and set the reward based on the smoothness of the speed and the anti-interference ability.

[0069] Specifically, the motor speed value is set as: State, which is used to evaluate the parameterization effect; the motor control environment is set as: Environment, which is responsible for giving the real-time state after the parameter changes. For the specific structure, reference can be made to Figure 3As shown; Action, learning to adjust and explore 12 parameters; Reward, evaluating according to the new state given by the environment after adopting the action. The adopted reward function is:

[0070] R = r1e os / s1 + r2t rs / s2 + r3t rl / s3 + r4|e l | / s4 (6)

[0071] In the formula: e os , e l , t rs , and t rl are the optimization objectives. e os and e l are the speed errors during startup and sudden load transition respectively. t rs and t rl are the startup time and the time for the speed to return to normal after sudden load respectively. s1, s2, s3, s4 are normalization coefficients because the dimensions of the optimization objectives are different. r1, r2, r3, r4 are the weight coefficients of the four optimization objectives and can be changed according to different requirements of the application environment. When the final evaluation value R is the smallest, the best optimization result will be obtained. For the optimization structure, refer to Figure 4 as shown.

[0072] Step 3: Use the AC algorithm to train the agent to select different optimization actions according to the environment and state, and improve the optimization actions according to the reward obtained after the action, so that the agent continuously interacts with the environment and autonomously learns the method of active disturbance rejection control parameter optimization.

[0073] Specifically, the classic Actor-Critic structure AC algorithm includes: Actor, which can learn and construct a policy network and select different actions according to the network in different states; Critic, which can evaluate the value of the actions of the optimized policy network. The agent evaluates according to the reward value and decides how the Actor should act according to the evaluation, that is, the adjustment direction and amplitude of the parameters. After this AC round, a new round of interaction and learning is carried out until the adjustment and optimization of the parameters converge. In addition, after the Actor takes the optimization behavior, the agent will add noise to this behavior, which can simulate the interference of the system and make the result more accurate.

[0074] Refer to Figure 5 as shown, the Actor structure network can be expressed as μ(s|θ μ ), θ μ are the internal parameters of the policy network μ. The current state and action are represented as s t and a t . respectively. The agent acts according to st of μ(s|θ μ ) to take action a t . The added noise can be expressed as:

[0075] a t = μ(s t |θ μ ) + Noise(7)

[0076] When an action is completed, the reward r t and the next state s t+1 are fed back, and the data (s t , a t , r t , s t+1 ) will be stored in the database. {(s i , a i , r i , s i+1 ) | i = 1, 2, …, N} then contains some data sets randomly selected from the database for training. Subsequently, the Critic evaluation network represented by Q(s, a|θ Q ) will evaluate based on s and a from the previous step.

[0077] Step 4: Design the deep deterministic policy gradient method to make the parameter optimization process converge and complete the training of the auto-disturbance rejection controller parameter optimization model based on deep reinforcement learning;

[0078] Specifically, the deterministic policy gradient method is used to make the critic converge and update the network parameters. As shown in Figure 5 , two gradient descent networks are constructed: the actual network θ μ that is updated in real time, and the target network θ Q that is updated with a delay and will be used at the end. The optimization update of the algorithm is achieved by updating θ μ and θ Q . To give the algorithm sufficient learning time, μ(s|θ μ ) and Q(s, a|θ Q ) are not directly used. By using μ’(s|θ μ’ ) and Q’(s, μ’(s|θ μ’ ) | θ Q’ ), the process of parameter update is split and amplified. The target network has the same structure as the actual network, and the target network slowly tracks the parameters of the actual network at a parameter update rate τ, τ ≤ 1:

[0079]

[0080] The evaluation value of the target network is y i, the loss between the target network and the actual network is L, and the two can be calculated according to (9). Using the minimum L as the optimization objective can optimize the Q-network parameters θ Q . Using the negative average value J of the network in (10) as the optimization objective to update θ μ , which can maximize the evaluation value of the actual network Critic.

[0081]

[0082] The action setting allows the algorithm to correct the ADRC parameters to be optimized and uses (11) to implement the normalization, restoration, and correction of the parameters. (12) is used to keep the optimized parameters within a feasible range.

[0083]

[0084] In the formula: θ max , θ min θ i , are respectively the upper limit, lower limit, original value, and regularized value of the parameters of the i-th generation. (13) is used to evaluate and process the error between the actual value and the given value after setting the optimization objective. (14) is used to punish and reward parameter correction, and (15) is used as the final evaluation of the optimization.

[0085] R Obs = error evaluation (13)

[0086]

[0087] R = -(α Obs R Obs + α θ R θ ) (15)

[0088] In the formula: α Obs , α θ are respectively the corresponding weights of the observation reward and the parameter correction reward and punishment.

[0089] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0090] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.

Claims

1. A parameter optimization method for an active disturbance rejection controller based on deep reinforcement learning, characterized in that, Including: Establish a speed-loop active disturbance rejection controller model; Select the active disturbance rejection controller parameters and their quantities to be optimized according to the speed-loop active disturbance rejection controller model, and set the optimization objectives; Build a deep reinforcement learning model for the active disturbance rejection controller parameters in combination with the Markov decision process, take the current motor control system as the environment, take the motor speed curve as the state, and set the rewards based on the smoothness of the speed and the anti-interference ability; Build a deep reinforcement learning model for the active disturbance rejection controller parameters in combination with the Markov process; use the Actor-Critic algorithm to enable the intelligent agent to continuously interact with the environment and autonomously learn the optimization method of the active disturbance rejection control parameters. Use the Actor-Critic algorithm to train the intelligent agent to select different optimization actions according to the environment and state, and improve the optimization actions based on the rewards obtained after the actions, so that the intelligent agent continuously interacts with the environment and autonomously learns the method of optimizing the active disturbance rejection control parameters; Design a deep deterministic policy gradient method to make the parameter optimization process converge and complete the training of the active disturbance rejection controller parameter optimization model based on deep reinforcement learning. The speed-loop active disturbance rejection controller model is a first-order model. The first-order speed-loop active disturbance rejection controller model includes a tracking differentiator, an extended state observer, and a non-linear error state feedback; Combine the Markov decision process with the active disturbance rejection control environment. The motor speed value is set as: the state, which is used to evaluate the parameter optimization effect; the motor control environment is set as: the environment, which is responsible for giving the real-time state after the parameters change; the action, which learns to adjust and explore the parameters; the reward, which is evaluated based on the new state given by the environment at this time after the action; the adopted reward function is: R = r1e os / s1 + r2t rs / s2 + r3t rl / s3 + r4|e l | / s4 where e os 、e l 、t rs and t rl are the optimization objectives, e os and e l are the speed errors during startup and sudden load transition respectively, t rs and t rl are the startup time and the time for the speed to return to normal after sudden load respectively, s1, s2, s3 and s4 are normalization coefficients because the dimensions of the optimization objectives are different; r1, r2, r3 and r4 are the weight coefficients of the four optimization objectives, which are changed according to the different requirements of the application environment. When the final evaluation value R is the smallest, the best optimization result will be obtained.

2. The parameter optimization method of the active disturbance rejection controller based on deep reinforcement learning according to claim 1, characterized in that Use the Actor-Critic algorithm to train the intelligent agent. The reward value is used as the evaluator, and the action is used as the actor. The intelligent agent evaluates according to the reward value and decides how the actor should act according to the evaluation; After this Actor-Critic round ends, a new round of interaction and learning is carried out until the adjustment and optimization of the parameters converge.

3. The method for optimizing the parameters of the active disturbance rejection controller based on deep reinforcement learning according to claim 1, characterized in that, Use the deep deterministic gradient policy to optimize and converge the Actor-Critic algorithm; calculate the direction of the gradient descent of each step of the action. To avoid falling into local optimization, two gradient descent networks are constructed to update the actual network in real time and delay the update of the target network to be used; after the intelligent agent finishes learning, use the target network to perform gradient descent on the Actor-Critic algorithm to make the parameter optimization process converge.

4. The self-disturbance rejection controller parameter optimization method based on deep reinforcement learning according to claim 1, characterized in that The tracking differentiator tracks the input signal and arranges the transition state to make the control fast and without overshoot.

5. The method for optimizing the parameters of the active disturbance rejection controller based on deep reinforcement learning according to claim 1, wherein The extended state observer is used to observe the output and its various order differential components. At the same time, the internal disturbance and external variables of the system are regarded as the total disturbance, and the total disturbance is observed and estimated.

6. The method for optimizing the parameters of the active disturbance rejection controller based on deep reinforcement learning according to claim 1, wherein The non-linear state error feedback mainly controls the non-linear combination of the reference input and the error signal of the extended state, and compensates for the observed total disturbance at the same time.

7. The self-disturbance rejection controller parameter optimization method based on deep reinforcement learning according to claim 1, wherein The Actor-Critic algorithm includes: the actor, which learns and constructs a policy network and selects different actions according to the network in different states; the evaluator, which evaluates the value of the actions of the optimized policy network.

8. The self-disturbance rejection controller parameter optimization method based on deep reinforcement learning according to claim 7, characterized in that After the actor takes an optimization action, the intelligent agent adds noise to this action, which can simulate the interference of the system and make the result more accurate.

Citation Information

Patent Citations

  • Flux weakening control strategy for permanent magnet motor

    CN112636652A