Permanent magnet synchronous motor active anti-interference optimization control method based on deep learning

By introducing a hybrid control strategy—Switching Active Disturbance Rejection Controller (SADRC) and Deep Reinforcement Learning (DRL)—to optimize parameters in permanent magnet synchronous motors, the problem of insufficient robustness of traditional controllers under complex disturbances and nonlinear parameters is solved, achieving better dynamic performance and anti-interference capability.

CN121333147APending Publication Date: 2026-01-13SUZHOU VOCATIONAL UNIVERSITY (SUZHOU OPEN UNIVERSITY) +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511481303.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing permanent magnet synchronous motor control systems suffer from high computational costs, complex parameter tuning, and insufficient robustness when dealing with complex disturbances and nonlinear parameters. Traditional active disturbance rejection controllers exhibit reduced local optima and efficiency in practical applications.

Method used

The Switch Active Disturbance Rejection Controller (SADRC) employs a hybrid control strategy, combines deep reinforcement learning (DRL) to optimize parameters, and performs automatic tuning using the CEC-DDPG algorithm to achieve seamless switching between linear and nonlinear active disturbance rejection control laws, thereby improving system performance.

Benefits of technology

It improves the dynamic performance and anti-interference capability of permanent magnet synchronous motors, reduces parameter dependence, and enhances the robustness and control accuracy of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121333147A_ABST
    Figure CN121333147A_ABST
Patent Text Reader

Abstract

The invention discloses a permanent magnet synchronous motor active disturbance rejection optimization control method based on deep learning, and the method comprises the following steps: S1, designing a hybrid power control strategy which is used for achieving the seamless switching between a linear active disturbance rejection control law and a nonlinear active disturbance rejection control law; s2, based on the proposed control strategy, designing a switching type active disturbance rejection controller SADRC for the permanent magnet synchronous motor; s3, aiming at a parameter setting problem of the SADRC, DRL is applied to automatic setting of parameter optimization of an ADRC, and a DRL model is established; and S4, in order to solve the parameter setting problem of the SADRC, providing a DRL (Depth Deterministic Policy Gradient), namely a CEC-DDPG (Classified Empirical Conduction) algorithm, based on a DDPG and CEC (Depth Deterministic Policy Gradient) algorithm. The SADRC provided by the invention has better dynamic performance, and for the parameters in the SADRC provided by the invention, the DRL is adopted to carry out parameter optimization, so that satisfactory dynamic performance is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of permanent magnet synchronous motor, and particularly relates to a permanent magnet synchronous motor active anti-disturbance optimization control method based on deep learning. BACKGROUND

[0002] Due to the advantages of high efficiency and fast dynamic response, permanent magnet synchronous motors (PMSM) are widely used in fields requiring excellent driving performance, such as robotics, aviation, electric vehicles, electric ship propulsion systems, and advanced CNC machine tool feed drive systems.

[0003] Field-oriented control (FOC) is a common control method for permanent magnet synchronous motors (PMSM) with good tracking performance. However, it is quite challenging to ensure dynamic performance throughout the speed range. To tackle this problem, the control system of permanent magnet synchronous motors (PMSM) has incorporated complex control algorithms such as sliding mode control (SMC), model predictive control (MPC), and active disturbance rejection control (ADRC) to achieve the desired control effect. Many complex working conditions, such as electric vehicle applications, require motor drive systems to have strong anti-disturbance capability.

[0004] The "disturbance estimation + feedforward compensation" architecture of the active disturbance rejection controller has an advantage in reducing disturbances, which is one of the commonly used anti-disturbance control methods for permanent magnet synchronous motors in electric vehicles. On the one hand, ADRC generates control quantities based on command tracking errors. On the other hand, it estimates disturbances through ESO to achieve feedforward compensation of control quantities, thereby reducing the impact of load disturbances on the system, ultimately making the system strong enough to handle a wide range of disturbances. However, ADRC requires high computational cost, complex parameter setting, and complex stability analysis. It has challenges in dealing with a large number of nonlinear parameters and the coupling relationship between them. Therefore, many scholars have studied the parameter tuning of the active disturbance rejection controller.

[0005] To solve the problems related to output signal jitter and the reduced anti-disturbance performance of the Fal function in active disturbance rejection control at the inflection point, an enhanced monkey algorithm is introduced for parameter adjustment. The active disturbance rejection control strategy based on the improved monkey algorithm reduces the dependence of permanent magnet synchronous motors on inherent parameters while enhancing the robustness of the system. However, this control strategy has limitations and can only be applied to specific situations. The existing technology also designs an active disturbance rejection controller based on an ant colony optimization (ACO) strategy to achieve disturbance compensation of the system. Based on the feedback signals of the system, the ant colony algorithm is set as a self-correcting algorithm for the parameters of the active disturbance rejection controller. Through the optimization mechanism and self-learning characteristics of the ant colony algorithm itself, the optimal parameters are obtained through iterative operations. However, the ant colony algorithm often falls into local optimality when dealing with complex practical problems.

[0006] Many scholars have studied parameter tuning and attempted to study the structure of active disturbance rejection controllers. The prior art proposes a two-degree-of-freedom fractional order proportional derivative (FOPD) controller with a general extended state observer (GESO) to achieve optimal set speed tracking and disturbance rejection performance for typical permanent magnet synchronous motor servo systems. In addition, an innovative linear active disturbance rejection control (LADRC) controller is introduced. The feedforward compensation element is integrated with the load torque observer, which helps to optimize the nonlinear function and improve the accessibility of parameter adjustment and calculation. However, the performance of the new nonlinear active disturbance rejection control (NADRC) controller is significantly reduced in the presence of a large amount of disturbance.

[0007] With the progress of contemporary computer technology, the integration of machine learning (ML) is becoming more and more common in different fields. Its ability to establish a direct relationship between observations and parameters has proven to be advantageous, bypassing the inherent limitations of heuristic algorithms. As a kind of machine learning, deep reinforcement learning (DRL) can independently learn and optimize functions without continuous supervision and develop reasonable control strategies to complete tasks, representing an ideal agent system. DRL has begun to be applied to some basic controller theories.

[0008] In the prior art, an enhanced speed and current control method is introduced, which uses the principle of DRL to enhance the robustness of the system. Since the proposed controller is independent of the parameters of the permanent magnet synchronous motor, the burden of the observer is heavy, and the precision of the control effect is not high. In order to solve this problem, a DRL-based permanent magnet synchronous motor vector control method is proposed. An advanced policy gradient algorithm is used to design a DRL vector controller to solve the inherent nonlinear complexity in system parameters. In addition, the prior art also proposes a DRL-based ADRC method, which is designed to address the constraints related to the inherent nonlinear error decay function in traditional active disturbance rejection methods. Although it performs commendable control effects under ideal conditions or theoretical simulations, the above applications have not yet been implemented to optimize the nonlinear parameters in ADRC. SUMMARY

[0009] To solve the above problems, the present application proposes a deep learning-based permanent magnet synchronous motor active disturbance optimization control method, which adopts a hybrid control strategy of a switching active disturbance rejection controller (SADRC) for permanent magnet synchronous motor drives. And use DRL algorithm to calibrate parameters to improve the overall system performance.

[0010] The specific scheme is as follows:

[0011] The deep learning-based permanent magnet synchronous motor active disturbance optimization control method comprises the following steps,

[0012] S1, design a hybrid control strategy for seamless switching between linear and nonlinear active disturbance rejection control law;

[0013] S2, based on the proposed control strategy, design a switching active disturbance rejection controller SADRC for permanent magnet synchronous motor;

[0014] S3, for the parameter setting problem of SADRC, deep reinforcement learning DRL is applied to the automatic setting of active disturbance rejection controller ADRC parameter optimization, and a DRL model is established;

[0015] S4, to solve the parameter setting problem of SADRC, a DRL based on deep deterministic policy gradient DDPG and classification experience transmission CEC policy gradient algorithm, namely CEC-DDPG algorithm, is proposed.

[0016] Further, in step S1, the switching criteria include the sum of the total disturbance estimates of LESO and nonlinear ESO under different weight coefficients, as shown in (19):

[0017] (19)

[0018] In the formula, ω * is the reference speed command; ω r is the feedback speed; e qh is the absolute value of the difference between the speed command and the feedback speed; z qh is defined as the absolute value of the disturbance combination considering the weight factor. z1 and z2 are the feedback speed and total disturbance observed by the nonlinear function, respectively, μ, β, and χ are weight coefficients, β max and χ max are the upper limits of the weight coefficients, s at (·) is a saturation function, and its expression is:

[0019] (20)

[0020] (21)

[0021] In the formula, e1 and e2 are the lower limit and upper limit of the speed error region observed during the transition process, respectively; D1 and D2 are the lower limit and upper limit of the disturbance region observed during the transition process, respectively.

[0022] According to (19), by setting different β and χ, the amplitude of β max and χ max can be changed; on this basis, the value of weight coefficient μ is determined according to β and χ;

[0023] (22)

[0024] According to (19), the significant difference between the target speed and the actual speed causes the parameter β to gradually increase; the adjustment of the χ variable depends on the estimated value derived from the total disturbance. On this basis, the weight coefficient μ combines the control rates originating from LADRC and NADRC, and obtains the control rate of the hybrid strategy:

[0025] (23)

[0026] In the formula, q is the lumped control quantity, 1 and 2 are the control quantities of LADRC and NADRC, respectively.

[0027] Further, in step S2, the SADRC control strategy designed for the permanent magnet synchronous motor combines the inherent advantages of LADRC and NADRC, and avoids their respective shortcomings through strategic conversion.

[0028] Further, in step S3, in the deep reinforcement learning DRL, the terms "environment", "state", "action" and "reward" constitute the basic components; R is a reward value function for calculating the reward value R(s, a) returned by the environment after the agent selects an action in the state; the agent updates the action according to the reward and the current state, and continuously interacts and evaluates the environment until R converges; DRL uses the current environment and input-output compatibility to determine the action, denoted as A, the state is represented by S, derived from the current output value, and then the environment calculates the relevant reward value, denoted as r.

[0029] Further, in step S4, a random conductor network is added to guide the network training of DRL on the basis of the original DRL, reducing the time required for network training; the permanent magnet synchronous motor uses its high-quality experience data to guide the current parameter selection of DRL; after the parameter selection is completed, the current running state of the permanent magnet synchronous motor is obtained, and the corresponding reward is obtained according to the current state; the CEC-DDPG algorithm adds a "talent pool" to store the latest high-quality sample data in the DRL network model for network training, improving the learning efficiency of the DRL network model;

[0030] The strategy π is a mapping from state to action, that is, what action to take in a certain state, denoted as:

[0031] (24)

[0032] In the Markov decision process, the state value function of the strategy p is denoted as v π (s):

[0033] (25)

[0034] The action value function of the policy is denoted as q π (s,a) represents

[0035] (26)

[0036] According to the Bellman equation, (32) and (33) can be rewritten as follows

[0037] (27)

[0038] (28)

[0039] According to the maximum cumulative reward and the Bellman equation, the optimal policy π* is found

[0040] (29)

[0041] In the CEC-DDPG algorithm, the policy network is used for policy update, where the deterministic policy is expressed by parameters, thereby generating a predetermined action value; the CEC-DDPG algorithm incrementally updates the target network at each time step and makes slight adjustments; the update method is as follows:

[0042] (30)

[0043] The default parameter τ update rate is 0.001;

[0044] The CEC-DDPG algorithm directly obtains the auxiliary policy through the participant network , and the action output by the conductor network is the guide policy ; the synthesis of the two strategies is used as the current output action; in order to add a certain degree of supervision and guidance to the parameter optimization, the difference between the guide policy and the current policy A is reduced to the target training network; the current policy is selected as defined below, and a guidance item is added to it through the guide policy :

[0045] (31)

[0046] N t represents random noise; the conductor network weight parameters are repeatedly updated:

[0047] (32)

[0048] The slope of the conductor network is represented as:

[0049] (33)

[0050] The difference between the reduced guide strategy Q and the current strategy A is taken as the target of training the network, and theta is updated μ In order to maximize the evaluation value of q; the objective function J is defined as the expectation of the discounted cumulative reward, denoted as

[0051] (34)

[0052] Optimizing the parameters of the active disturbance rejection system; (35) is used for normalizing the parameters, and (36) is used for making the optimized parameters feasible;

[0053] (35)

[0054] (36)

[0055] In the formula, theta max , theta min , and i , theta i are the upper limit, lower limit, original value and normalized value of the i-th parameter, respectively.

[0056] After setting the optimization target, (36) is used to evaluate and process the error between the actual value and the given value; formula (37) imposes a penalty and reward on the parameter adjustment, and formula (38) is used as the final evaluation index of the optimization process;

[0057] (37)

[0058] (38)

[0059] In the formula, alpha Obs and alpha θ are the corresponding weights of the observation reward and the parameter reward, respectively, and R Obs is the initial observation reward function value.

[0060] The beneficial effects of the present application are:

[0061] 1) A switch active disturbance rejection controller (SADRC) using a hybrid control strategy is proposed. Compared with the traditional method, the SADRC of the present application has better dynamic performance.

[0062] 2) For the parameters in the proposed SADRC, DRL is used for parameter optimization, and satisfactory dynamic performance is obtained. BRIEF DESCRIPTION OF DRAWINGS

[0063] Figure 1 is the ADRC control block diagram in the embodiment.

[0064] Figure 2is the SADRC control block diagram in the embodiment.

[0065] Figure 3 is the specific flow chart of DRL in the embodiment.

[0066] Figure 4 is the ADRC control system block diagram of permanent magnet synchronous motor based on DRL in the embodiment.

[0067] Figure 5 is the CEC-DDPG algorithm chart in the embodiment.

[0068] Figure 6 is the permanent magnet synchronous motor starting performance comparison chart in the embodiment, (a) speed comparison, (b) current comparison.

[0069] Figure 7 is the speed and phase current waveform chart of permanent magnet synchronous motor in the embodiment, (a) speed, (b) phase current waveform.

[0070] Figure 8 is the dynamic performance chart of permanent magnet synchronous motor under 5 N·m and 10 N·m step load at 2000 rpm in the embodiment, (a) speed comparison of three methods, (b) phase current comparison of three methods.

[0071] Figure 9 is the dynamic performance chart of permanent magnet synchronous motor under 5 N·m and 10 N·m step load at 1000 rpm in the embodiment, (a) speed comparison of three methods, (b) phase current comparison of three methods.

[0072] Figure 10 is the dynamic performance chart of permanent magnet synchronous motor under 5 N·m and 10 N·m step load at 200 rpm in the embodiment, (a) speed comparison of three methods, (b) phase current comparison of three methods. DETAILED DESCRIPTION

[0073] The application will be further illustrated below in combination with the drawings and specific embodiments, and it should be understood that the following specific embodiments are only used to illustrate the application and not to limit the scope of the application.

[0074] The application provides a deep learning-based permanent magnet synchronous motor active anti-disturbance optimization control method, including the following contents: 1, introduction of permanent magnet synchronous motor model and traditional active disturbance rejection controller ADRC, 2, introduction of SADRC design of the application, 3, experimental verification.

[0075] 1, introduction of permanent magnet synchronous motor model and traditional active disturbance rejection controller ADRC

[0076] A, modeling of permanent magnet synchronous motor

[0077] By coordinate transformation, the mathematical model of permanent magnet synchronous motor in ABC three-phase natural coordinate system can be simplified to d-q in rotating coordinate system.

[0078] The stator voltage equation can be expressed as:

[0079] (1)

[0080] The flux equation can be expressed as:

[0081] (2)

[0082] Where u d , u q, , i d , i q , R, ω e , ψ d , ψ q are d-axis voltage, q-axis voltage, d-axis current, q-axis current, stator resistance, rotor angular velocity, d-axis component of stator flux, q-axis component, respectively. d L q are d-axis and q-axis inductance, respectively. f y is the rotor permanent magnet flux linkage.

[0083] By substituting (2) into (1), the stator voltage equation can be rewritten as:

[0084] (3)

[0085] The electromagnetic torque equation can be expressed as:

[0086] (4)

[0087] Where P n is the number of poles of the motor.

[0088] The mechanical motion equation can be expressed as:

[0089] (5)

[0090] Where ω m is the mechanical angular velocity, J is the moment of inertia, T L is the load torque, and B is the viscous friction coefficient.

[0091] The second-order equation of velocity is:

[0092] (6)

[0093] Assume

[0094] (7)

[0095] (8)

[0096] In this case, the second-order equation of the speed can be expressed as

[0097] (9)

[0098] where, represents the overall disturbance experienced by the system. By effectively estimating and compensating for the entire disturbance, the anti-interference performance of the control system can be improved. According to the formula in (9), the expression of the permanent magnet synchronous motor state equation is as follows:

[0099] (10)

[0100] B. Basic principles of ADRC

[0101] 1) Tracking differentiation factor

[0102] A typical active disturbance rejection controller consists of three parts: a tracking differentiator (TD), a state error feedback (SEF) control law, and an extended state observer (ESO). The principle of the active disturbance rejection controller is introduced by taking a second-order system as an example.

[0103] (11)

[0104] where, y is the system output variable, x is the system control state variable, u is the system input variable, w is the external disturbance state variable, and b is the system control gain. The establishment of the second-order SADRC controller design is consistent with the formula stated in (11).

[0105] As an important part of the active disturbance rejection control theory, TD aims to create a transient profile for the input command signal to prevent the rate of change of the input signal from exceeding the tracking ability of the system, thereby causing excessive tracking error and overshoot. In obtaining the derivative of the input signal, TD adopts the idea of the fastest tracking, and by constructing the fastest tracking differential equation, it converts the process of obtaining the derivative of the input signal into the integral of the equation. Therefore, it overcomes the problem of noise amplification introduced by the traditional Euler method, and to some extent, it alleviates the contradiction between the "speed" and "smoothness" of the differential operation. In summary, TD can accurately obtain a smooth differential signal from the original input signal that changes rapidly and contains random noise. Figure 1 The control block diagram of ADRC is shown. The general form of TD is as follows:

[0106] (12)

[0107] where v is the original input, v1 and v2 are the tracking value of v and its differential, respectively, and the sign function denoted as fhan(*) acts as the optimal control synthesis function of TD, which is considered to be the fastest, and is defined as follows:

[0108] (13)

[0109] fhan(*) has two adjustable parameters r0 and h0. Among them, r0 is called the "speed factor", which determines the tracking speed of TD; h0 is called the "filter factor", which plays the role of noise suppression function, which can be adjusted independently to meet the system's requirements for tracking speed and noise level.

[0110] 2) Extended State Observer

[0111] ESO, as the core of ADRC, regards all unmodeled parts different from the standard series integral model and other unknown factors as a whole, collectively known as concentrated disturbance, and extends it to an independent state to establish a state observer for observation. In the case of the second-order system introduced in (10), the subsequent establishment involves the formula of the second-order extended state observer:

[0112] (14)

[0113] where z1 and z2 are used to track the state variables of the system, z3 is used to track the state variable of the total disturbance f(x, w), β1, β2, β3 are the gains of ESO, φ1(e), φ2(e), φ3(e) are the processing functions of observation errors. When φ1(e), φ2(e) and φ3(e) are selected as components of nonlinear functions, the corresponding ESO is called nonlinear ESO. The general representation of such nonlinear functions can be represented as:

[0114] (15)

[0115] where α and δ are undetermined parameters. When the system running time t < T or the extended state observer tracking deviation |e| > 1 or the total disturbance |z n+1 > 1 is one of the three, the ESO switches to linear ESO (LESO). When none of the above three conditions is met, the observer switches to nonlinear ESO.

[0116] 3) State Error Feedback Control Law

[0117] State Error Feedback (SEF) control law combines the tracking output of TD with the error derived from the estimated state variables of various orders by ESO. This merging, together with the estimated concentrated disturbance, produces the actual control input of the controller. The expression of SEF is as follows.

[0118] (16)

[0119] where u is the control amount of the controller, k1 and k2 are proportional gains. When the system running time t < T or the extended state observer tracking deviation |e| > 1 or |v i −z i |>1 is one of the three, the control law becomes linear, and the control amount can be expressed as:

[0120] (17)

[0121] When the above three conditions are not met, the nonlinear control law shown in the switching formula.

[0122] (18)

[0123] Second, the design of SADRC is introduced

[0124] S1, design a hybrid control strategy for seamless switching between linear and nonlinear active disturbance rejection control law;

[0125] Considering the transient effect of switching different control strategies on motor operation in actual system, it is necessary to design a hybrid switching strategy. In order to realize seamless switching between linear and nonlinear active disturbance rejection control law, the absolute values of command speed error and feedback speed error are considered. The seamless switching criteria include the sum of total disturbance estimates of LESO and nonlinear ESO under different weight coefficients, as shown in (19):

[0126] (19)

[0127] In the formula, ω * is the reference speed command; ω r is the feedback speed; e qh is the absolute value of the difference between the speed command and the feedback speed; z qh is defined as the absolute value of the estimated disturbance combination considering the weight factor. z1 and z2 are the feedback speed and total disturbance observed by the nonlinear function respectively, μ, β, and χ are weight coefficients, β max and χ max are the upper limits of the weight coefficients, s at (·) is a saturation function, and its expression is

[0128] (20)

[0129] (21)

[0130] According to (19), by setting different β and χ, β max and χ maxthe amplitude of the oscillations; on this basis, the value of the weight coefficient μ is determined as a function of β and χ;

[0131] (22)

[0132] According to (19), a significant difference between the target speed and the actual speed leads to an increase in the parameter β; the adjustment of the χ variable depends on the estimate derived from the total disturbance. On this basis, the weight coefficient μ combines the control rates originating from LADRC and NADRC, leading to the control rate of the hybrid strategy:

[0133] (23)

[0134] S2, based on the proposed control strategy, a switched active disturbance rejection controller SADRC is designed for permanent magnet synchronous motor;

[0135] The block diagram of the permanent magnet synchronous motor control system based on SADRC is shown in Figure 2 In this section, the typical ADRC is improved so that the ADRC controller can obtain optimal control even under large disturbances. The SADRC control strategy designed for the permanent magnet synchronous motor combines the inherent advantages of LADRC and NADRC, and avoids their respective shortcomings through strategic switching. This integration improves the anti-interference performance and accuracy of motor control. However, in order to apply SADRC in practice, a more convenient parameter selection strategy is still needed.

[0136] S3, for the parameter setting problem of SADRC, deep reinforcement learning DRL is applied to the automatic setting of active disturbance rejection controller ADRC parameter optimization, and a DRL model is established;

[0137] Deep reinforcement learning (DRL) combines the powerful perception ability inherent in deep learning with the strategic decision-making ability of reinforcement learning. This integration gives the system powerful deep learning perception ability and reinforcement learning proficiency in strategy selection and decision-making. In practical application scenarios, DRL uses the feature extraction ability of deep learning, followed by the skilled decision-making ability of reinforcement learning to formulate the control strategy of action output.

[0138] The present application uses DRL to optimize SADRC parameters. Figure 3 is the specific flowchart of DRL. In deep reinforcement learning DRL, the terms "environment", "state", "action" and "reward" constitute the basic components; R is a reward value function for calculating the reward value R(s,a) returned by the environment after the agent selects an action in the state; the agent updates the action according to the reward and the current state, and continuously interacts and evaluates the environment until R converges;

[0139] The RL-DRC Agent conceptualizes the entire motor control system as encompassing the environment, designating motor speed as a state variable, and measuring motor control efficiency as a reward signal. Within this framework, the DRL utilizes the current environment and input-output compatibility to determine the action, denoted as A, with the state represented by S, derived from the current output value. Subsequently, the environment calculates the relevant reward value, denoted as r. The evaluation of the selected action is performed by a critic, guiding subsequent updates from the agent. In response, the DRL-ADRCA Agent participates in the evaluation, optimization, and update process to determine subsequent actions. By executing the selected action, the agent interacts with the environment, obtaining rewards and subsequent states from the subsequent feedback loop. This iterative cycle continues until the reward value converges. Once training is complete, the agent autonomously determines the optimization and selection of actions based on the learned experience.

[0140] S4. To solve the parameter tuning problem of SADRC, a DRL based on Deep Deterministic Policy Gradient (DDPG) and Classification Experience Transmission (CEC) Policy Gradient Algorithm is proposed, namely the CEC-DDPG algorithm.

[0141] Traditional Deep Deterministic Policy Gradient (DDPG) algorithms in DRL suffer from low learning efficiency and long training times. This paper proposes Classification Experience-Guided DDPG (CEC-DDPG). A random conductor network is added to the original DRL to guide network training, reducing the training time. The permanent magnet synchronous motor (PMSM) utilizes its high-quality empirical data to guide parameter selection in the current DRL. After parameter selection, the current operating state of the PMSM is obtained, and a corresponding reward is awarded based on this state. The CEC-DDPG algorithm incorporates a "talent pool" into the DRL network model to store recent high-quality sample data for network training, improving the learning efficiency of the DRL network model. Figure 4 This is a block diagram of a DRL-based active disturbance rejection control system for permanent magnet synchronous motors.

[0142] Policy π is a mapping from state to action, that is, what action to take in a given state, expressed as:

[0143] (twenty four)

[0144] In Markov decision-making, the state-value function of policy p is represented by v. π (s) means:

[0145] (25)

[0146] The action-value function of the strategy is q π (s,a) represents

[0147] (26)

[0148] According to Bellman equation, (32) and (33) can be rewritten as follows

[0149] (27)

[0150] (28)

[0151] In the formula, π*, v* and q* are used to represent the optimal policy, the optimal state value letter and the expected return of decision-making in the Markov decision process according to the optimal policy respectively.

[0152] According to the maximum cumulative income and Bellman equation, the best strategy π* is found

[0153] (29)

[0154] In the CEC-DDPG algorithm, the policy network is used for policy update, in which the deterministic policy is expressed by parameters, so as to generate the predetermined action value; the CEC-DDPG algorithm updates the target network incrementally at each time step and makes slight adjustment; the update method is as follows:

[0155] (30)

[0156] The default parameter τupdate rate is 0.001;

[0157] The CEC-DDPG algorithm directly obtains the auxiliary strategy through the participant network , and the action output by the conductor network is the guide strategy ; the synthesized strategy of the two is taken as the current output action; in order to add a certain degree of supervision and guidance to the parameter optimization, the difference between the guide strategy and the current strategy A is reduced to the target training network; the current strategy is selected as defined as follows, and the guide item is added to it through the guide strategy :

[0158] (31)

[0159] N t represents random noise; the conductor network weight parameters are repeatedly updated:

[0160] (32)

[0161] The slope of the conductor network is represented as:

[0162] (33)

[0163] The difference between the reduced guidance policy Q and the current policy A is taken as the target of training the network, and θ is updated μ to maximize the evaluation value of q; the objective function J is defined as the expectation of the discounted cumulative reward, denoted as

[0164] (34)

[0165] Eleven parameters of the active disturbance rejection system are optimized; these parameters need to be corrected; (35) is used for normalizing the parameters, and (36) is used for making the optimized parameters feasible.

[0166] (35)

[0167] (36)

[0168] where θ max , θ min , and i , θ i are the upper limit, lower limit, original value, and normalized value of the i-th parameter, respectively;

[0169] After setting the optimization target, (36) is used to evaluate and process the error between the actual value and the given value; formula (37) imposes a penalty and reward on the parameter adjustment, and formula (38) is used as the final evaluation index of the optimization process;

[0170] (37)

[0171] (38)

[0172] Figure 5 A DDPG block diagram based on CEC is shown.

[0173] III. Experimental verification

[0174] In order to verify the effectiveness of the proposed scheme, the following experimental tests were carried out on the experimental platform. The experimental platform consists of a permanent magnet synchronous motor, a torque sensor, a magnetic powder brake, a current sensor, a low-voltage box, a drive board, and a dSPACE DS1007. The simulation model of the permanent magnet synchronous motor is built based on the motor parameters described in Table I. In view of the autonomy of daytime running light training, time and computing resources are required, and the training process is carried out in an offline environment. The experimental results consist of PI control with optimal speed loop (method 1), SADRC method (method 2), and CEC-SADRC method (method 3). These results are used as reference points for comparative analysis.

[0175] The specific parameters of the motor are shown in Table 1. All control processes are completed by dSPACE DS1007 PPC. The sampling frequency is 10 kHz. The proposed control scheme is implemented on the dSPACE 1007 test bench. Experimental measurement results are exported from the dSPACE platform to MATLAB and plotted.

[0176] Table 1 Permanent magnet synchronous motor system parameters

[0177]

[0178] A. Optimization settings of CEC-SADRC

[0179] The optimization objectives are established to improve the efficiency of the proposed method, including reducing steady-state error, enhancing anti-interference, and minimizing the error difference between the specified torque and the actual torque in the control system. The observed reward value is designed as

[0180] (39)

[0181] where e os is the error of the observed state and the reference state, e l is the speed error of sudden load transition, e ts is the error of the observed torque and the reference torque. T ss and t sl represent the speed increase time and recovery time of the disturbance load, respectively. These time parameters can be calculated after the disturbance starts. S 1-5 is the standardization coefficient between the optimization objectives due to different dimensions. R 1-5 is the weight factor, which can be adjusted according to different application requirements. The best result is obtained when the final evaluation value R is the lowest.

[0182] B. Start performance comparison experiment

[0183] The starting performance of three control methods for permanent magnet synchronous motors is compared. Figure 6 The experimental results of the starting speed waveform and phase current based on the CEC-SARC control method, the LADRC control method, and the SADRC control method are shown, where the reference speed is established at a rated speed of 2000 rpm.

[0184] Compared with the experimental results in Figure 6 , it can be seen that the LADRC has a longer stable time and a larger speed overshoot, the SDARC has a shorter stable time and a smaller overshoot, the CEC-ADRC has the shortest stable time and the smallest overshoot, and the CEC-SADRC has the best starting performance.

[0185] C. Steady-state performance analysis

[0186] The performance of the permanent magnet synchronous motor in a given state will be compared. The given speed is 2000 rpm and the load torque is set to 1 N-m. Figure 7 The speed and phase current waveforms of the permanent magnet synchronous motor are shown. The steady-state performance of the LADRC, SADRC, and CEC-SADRC control strategies is compared and evaluated. From Figure 7 It can be seen that the speed fluctuation of the LADRC control is the largest, while the speed fluctuation of the CEC-SADRC is smaller, with superior steady-state performance.

[0187] D. Dynamic characteristics

[0188] This section introduces the dynamic performance comparison experiment of the permanent magnet synchronous motor, specifically, the experimental comparison of the three control methods under sudden load. The purpose of the comparison experiment under various specified speeds is to verify the superiority and effectiveness of CEC-SADRC in the entire speed range. The experiment applies a sudden load of 5 N-m and 10 N-m at 0.4 seconds and 1.2 seconds, respectively. Figures 8-10 The corresponding speed and a-phase current graphs of these methods are shown. Table 2 lists the comparative analysis of the experimental results, which illustrates the dynamic performance properties of the permanent magnet synchronous motor.

[0189] The experimental results show that the CEC-SADRC control method has a smaller overshoot and shorter adjustment time in the speed range of 0-2000 rpm. The dynamic performance of the permanent magnet synchronous motor using CEC-SADRC and LADRC control is superior to that of the permanent magnet asynchronous motor using CEC-ADRC.

[0190] Table 2 Dynamic performance comparison of permanent magnet synchronous motor

[0191]

[0192] The technical means disclosed in the present application scheme are not limited to the technical means disclosed in the above-mentioned embodiments, and also include technical solutions composed of any combination of the above technical features. It should be noted that for ordinary skilled persons in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which are also considered within the protection scope of the present application.

Claims

1. A deep learning-based active disturbance rejection optimization control method for permanent magnet synchronous motor, comprising the following steps: S1, designing a hybrid control strategy for realizing seamless switching between linear and nonlinear active disturbance rejection control laws; S2, based on the proposed control strategy, designing a switching active disturbance rejection controller SADRC for the permanent magnet synchronous motor; S3, for the parameter setting problem of SADRC, applying deep reinforcement learning DRL to the automatic setting of ADRC parameter optimization, and establishing a DRL model; S4, to solve the parameter setting problem of SADRC, a DRL based on deep deterministic policy gradient DDPG and classification experience-based CEC policy gradient algorithm, namely CEC-DDPG algorithm, is proposed.

2. The deep learning-based permanent magnet synchronous motor active disturbance rejection optimization control method according to claim 1, characterized in that, In step S1, the switching criteria include the sum of the total disturbance estimates of LESO and nonlinear ESO under different weight coefficients, as shown in (19): (19) where ω * is the reference speed command; ω r is the feedback speed; e qh is the absolute value of the difference between the speed command and the feedback speed; z qh is defined as the absolute value of the disturbance combination estimated taking into account the weight factors; z1 and z2 are respectively the nonlinear function of the observed feedback speed and the total disturbance, μ, β, and χ are weight coefficients, β max and χ max are upper limits of the weight coefficients; s at (·) is a saturation function expressed as: (20) (21) In the formula, e1 and e2 are the lower and upper limits of the observed speed error region during the transition process; D1 and D2 are the lower and upper limits of the disturbance region observed during the transition process; According to (19), by setting different β and χ, the amplitude of β max and χ max can be changed; on this basis, the value of the weight coefficient μ is determined according to β and χ; (22) According to (19), the significant difference between the target speed and the actual speed causes the parameter β to gradually increase; the adjustment of the χ variable depends on the estimate derived from the total disturbance. On this basis, the weight coefficient μ combines the control rates from LADRC and NADRC to obtain the control rate of the hybrid strategy: (23) In the formula, q is a lumped control variable, 1 and 2 are control variables of LADRC and NADRC, respectively.

3. The deep learning-based permanent magnet synchronous motor active disturbance rejection optimization control method according to claim 2, characterized in that, In step S2, the SADRC control strategy designed for the permanent magnet synchronous motor combines the inherent advantages of LADRC and NADRC, and avoids their respective shortcomings through strategic conversion.

4. The deep learning-based permanent magnet synchronous motor active disturbance rejection optimization control method according to claim 3, characterized in that, In step S3, in deep reinforcement learning DRL, the terms "environment", "state", "action" and "reward" constitute the basic components; R is a reward value function for calculating the reward value R(s,a) returned by the environment after the agent selects an action in a state; the agent updates the action according to the reward and the current state, and continuously interacts and evaluates the environment until R converges; DRL uses the current environment and input-output compatibility to determine the action, denoted as A, the state is represented by S, derived from the current output value, then the environment calculates the relevant reward value, denoted as r.

5. The deep learning-based permanent magnet synchronous motor active disturbance rejection optimization control method according to claim 4, characterized in that, In step S4, a random conductor network is added to guide the network training of DRL based on the original DRL, reducing the time required for network training; The permanent magnet synchronous motor uses its high-quality experience data to guide the current parameter selection of DRL; after the parameter selection is completed, the current running state of the permanent magnet synchronous motor is obtained, and the corresponding reward is obtained according to the current state; the CEC-DDPG algorithm adds a "talent pool" to the DRL network model to store the latest high-quality sample data for network training, improving the learning efficiency of the DRL network model; The strategy π is a mapping from state to action, i.e. what action to take in a certain state, denoted as: (24) In a Markov decision process, the state-value function for a policy p is denoted by v π (s) and is defined as the expected return of a policy p starting from state s. (25) The action value function for a policy is denoted by q π (s, a) (26) According to the Bellman equation, (32) and (33) are rewritten as follows (27) (28) According to the maximum cumulative return and the Bellman equation, the optimal strategy π* is found as (29) In the CEC-DDPG algorithm, the policy network is used for policy updating, in which the deterministic policy is expressed by parameters to generate predetermined action values; the CEC-DDPG algorithm incrementally updates the target network at each time step and makes slight adjustments; the updating method is as follows: (30) Where the default parameter τupdate rate is 0.001; The CEC-DDPG algorithm directly obtains the auxiliary policy through the participant network , the conductor network outputs the action as the guide policy ; the synthesis policy of the two is taken as the current output action; in order to add a certain degree of supervision and guidance to parameter optimization, the difference between the guide policy and the current policy A is reduced as the target training network; the current policy is selected as defined below, and the guide policy is added to it as a guide item: (31) N t represents random noise; repeatedly updating the conductor network weight parameters gives: (32) The slope of the conductor network is represented as: (33) Minimize the difference between the reduced guidance policy Q and the current policy A as the target of training the network, and update θ μ Maximize the evaluation value of q with Δ; the objective function J is defined as the expectation of the discounted cumulative reward, denoted as (34) The parameters of the active disturbance rejection system are optimized; (35) is used for normalizing the parameters, and (36) is used for making the optimized parameters feasible; (35) (36) where θ max , θ min , and i , θ i are the upper limit, lower limit, original value, and normalized value of the i-th parameter, respectively; After setting the optimization target, (36) is used to evaluate and process the error between the actual value and the given value; formula (37) imposes penalties and rewards on parameter adjustment, and formula (38) serves as the final evaluation index of the optimization process; (37) (38) where α Obs and α θ are respective weights for observed reward and parametric reward, respectively, and R Obs is an initial observed reward function value.