Method for dynamic tuning of electric servo position feedback based on deep reinforcement learning in semi-closed loop scenario
Patent Information
- Application Number
- CN202311187531.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-14
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-09-14
AI Technical Summary
虽然因算法的低复杂性和稳定性在工业实践中获得了广泛的应用,但是当其面临一些复杂和特殊的驱动场景,PID三环控制器也面临诸多挑战和不足:
[0023]有益效果:本发明所述方法针对仅能测量永磁同步电机转动角度等电机反馈信号,而无法测得负载机构实际位置的半闭环控制情形,考虑到负载模型的高阶非线性特征,提出了在永磁同步电机FOC控制框架下,采用传统PID三环控制器作为基础,并使用双延迟深度确定策略梯度算法训练调优策略网络,使其观测永磁同步电机反馈量、输出位置环反馈位置调优值,以改进传统PID三环控制在面对高阶非线性负载模型时的控制精度和响应速度。
Smart Images

Figure CN117335700B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer control system technology, specifically relating to the position control algorithm of permanent magnet synchronous motors, and more particularly to an optimization method for position control of permanent magnet synchronous motors in a semi-closed-loop feedback scenario where the actual position of the load cannot be measured. Background Technology
[0002] Permanent magnet synchronous motors (PMSMs) are widely used in electric transportation, industrial robots, and aerospace due to their advantages such as compact structure, high efficiency and power density, and good speed regulation performance. The field-oriented control (FOC) strategy commonly used in PMSM position servo systems is often based on a PID three-loop controller to control the motor's rotation angle. While the low complexity and stability of its algorithm have led to its widespread application in industrial practice, the PID three-loop controller also faces several challenges and shortcomings when encountering complex and special driving scenarios. First, when a servo motor drives a high-order nonlinear load mechanism, the parameters of the PID controller are difficult to determine through models and specifications. Therefore, trial and error adjustments based on human experience are necessary, resulting in poor dynamic response performance. Second, considering the motor system's limitations on speed and current, and the inverter's output voltage threshold, the controller's performance is further degraded when it receives highly dynamic position, speed, and current commands. Most challenging is that the special equipment addressed in this invention, due to its unique working environment, cannot have external sensors installed or reliably measure the actual position of the load mechanism. The controller receives feedback only from the motor shaft angle detected by a magnetic encoder, resulting in a semi-closed-loop control system. Compared to closed-loop control that monitors the final actuator, the semi-closed-loop scenario causes a significant performance degradation in traditional PID controllers based on feedback errors. For example, this invention considers the motor shaft driving a lead screw via gears to feed or retract, thereby causing the swing mechanism connected to the triangular linkage structure to deflect at a certain angle. The swing angle of the swing mechanism cannot be measured by sensors, and the motor rotation angle and the swing angle of the swing mechanism are affected by the elastic motion of the lead screw, involving high-order nonlinear dynamic equations that are difficult to express simply with functions. Therefore, the control algorithm cannot obtain actual angle feedback to form closed-loop control.
[0003] Therefore, designing a control tuning algorithm that can run on embedded chips to optimize the tracking accuracy and response speed of the actual load angle to the command deflection angle under semi-closed-loop control is of great significance and presents considerable challenges. Summary of the Invention
[0004] Purpose of the invention: To address the shortcomings and problems of the aforementioned PID three-loop control method in the semi-closed-loop case of permanent magnet synchronous motors when facing high-order nonlinear loads, this invention provides a dynamic optimization method for electric servo position feedback based on deep reinforcement learning in a semi-closed-loop scenario.
[0005] Technical Solution: A dynamic optimization method for electric servo position feedback based on deep reinforcement learning in a semi-closed-loop scenario. The method is based on a high-order nonlinear load and a motor simulation model under the FOC control framework. On the basis of PID three-loop control, the state space and action settings of the reinforcement learning agent are determined. Then, with the goal of improving the tracking accuracy and response speed of the load swing angle to the command target angle under semi-closed-loop control, the system model is modeled using simulation software, and the optimization strategy network for motor position feedback is obtained using deep reinforcement learning method to output the optimal optimization value. The method includes the following steps: (1) Construct a permanent magnet synchronous motor operation model based on the FOC control framework. Under FOC control, the motor system will receive power from the controller. The voltage command values of the two shafts are used as inputs, and the input torque to the load mechanism is expressed as the rotation angle of the motor rotor. (2) Construct a mathematical model of the load and its transmission mechanism. The motor load model considers the elastic deformation of the transmission mechanism, the dynamic equation of the swing mechanism, and also the nonlinear factors including Coulomb friction and torque transmission in the triangular linkage mechanism. (3) For the PMSM and high-order nonlinear load models modeled above, a three-loop control scheme based on PID controller is constructed on the FOC control framework; (4) The policy network is trained by reinforcement learning algorithm. The system performance is improved by changing the input of the position loop PI controller, including the use of TD3 reinforcement learning algorithm with continuous state and action space for agent optimization.
[0006] Furthermore, in order to simulate the behavior of the permanent magnet synchronous motor operating model within the FOC control framework using simulation software, the control variables, response variables, and related constraints are described by the following set of variables and equations: (11) The coordinate system under relative rotation , Voltage of two axes and As a system input, the drive motor applies torque to the transmission mechanism, which manifests as the rotation angle of the motor rotor. At the same time, the load and transmission device apply a time-varying load torque to the motor. ; (12) The controller obtains the rotation angle of the motor rotor. and angular velocity As a calculation parameter, the current sensor can also detect the stator current. , , And use Clarke transform and Park transform to convert it to , Current values of both axes , ; (13) Based on the actual motor model, the motor model is determined to be a surface-mounted permanent magnet synchronous motor, satisfying:
[0007] Considering the rotor speed of the permanent magnet synchronous motor operating in steady state and stator current It should be kept within the threshold range, satisfying:
[0008]
[0009] in, This represents the maximum rotor speed. This represents the maximum value of the stator current.
[0010] (14) The FOC strategy of permanent magnet synchronous motor decomposes the phase variables of the motor into magnetic field components and torque components, and controls them independently. Precise control of the motor's magnetic field and torque requires obtaining the rotor position and speed information of the motor for the phase variables in the stationary coordinate system. The transformation calculation between relative rotation variables of coordinates, in In the coordinate system, the motor equations can be described as follows:
[0011] in: Stator resistance; , They are respectively axis, Shaft voltage; , They are respectively axis, shaft current; , They are respectively axis, Shaft inductance; Permanent magnet Axial flux; , These are the output torque and the load torque, respectively. The bearing viscosity coefficient; The total moment of inertia of the motor and load; , These are the rotor mechanical angular velocity and the rotor electromagnetic angular velocity, respectively. For the number of permanent magnets, satisfying This module responds to input. The voltage across both axes depends on the external load torque. Output motor angle and angular velocity And obtained by current detector Axis current.
[0012] Furthermore, to simulate the special scenario of semi-closed-loop control, a dynamic model of the transmission and load mechanism was built using Simulink, taking into account higher-order nonlinear factors in the model, including: (21) The driving torque of the motor shaft is transmitted through a transmission ratio of The two-stage gear transmission drives the lead screw to rotate, and the lead screw rotates at... The reduction ratio produces a corresponding process, where the rotation generates a force on the elastic lead screw as follows:
[0013] (22) Considering the mass as Elastic damping is The combined stiffness is The elastic motion of the lead screw under compression, its dynamic equation is:
[0014] in, This refers to the compression stroke of the lead screw; (23) The elastic deformation of the lead screw generates a force that pushes the lead screw side of the triangular linkage mechanism, causing the lever arm of the load side to generate a torque that causes the load to deflect around the axis. Among them, the extension / retraction amount of the lead screw With deflection angle The relationship can be approximated as follows:
[0015] in , These are the two adjacent sides of the swing angle in a triangular linkage mechanism. , long; , These are the pendulum angles when the load deflection angle is 0. and edge length; (24) Consider a oscillating moment of inertia as The oscillation damping is The positional resistance torque is Frictional torque is The swing mechanism receives a size of The torque, its dynamic equation is:
[0016] The frictional torque is modeled as Coulomb friction: .
[0017] Furthermore, to ensure the stability and robustness of the system operation, the method in step (3) employs a three-loop control based on a PI controller to perform basic control on the final position angle, including: (31) The current loop controlling the torque is constructed as a decoupled current controller, which decouples the originally coupled current loops. , The two-axis voltage term is decomposed into linear and nonlinear terms:
[0018] in, and It can be controlled by a linear PID current controller, without the nonlinear term. and It can be calculated from the rotor speed value of the encoder:
[0019] (32) Input error in PID controller and output control value The connection between the two: ; (33) Feedback motor speed and speed command The difference is used as the input to the speed PI controller, and the controller's output is used as... Reference value for shaft current; (34) Based on the speed controller, a position control loop is constructed, specifically a load-based triangular linkage structure, which directs the command deflection angle. Approximately converted to lead screw precession length and compare it with the rotor angle feedback signal. Converted precession length The error value is used as the input to the position loop PI controller, and the speed loop reference value is used as the output.
[0020] Furthermore, for basic PID three-loop control in a semi-closed-loop control scenario, the method uses the TD3 reinforcement learning algorithm, which has continuous state and action spaces, to optimize the agent in step (4), enabling it to predict the lead screw precession length. Approximate lead screw precession length The optimized value of the deviation between them is used to alleviate the problem of insufficient control accuracy caused by position feedback error, wherein: (41) The value observed by the agent is determined as the instruction deflection angle. Calculated lead screw precession length and rotor feedback angle Approximate lead screw precession length as well as Rate of change and speed feedback After quantizing, saturating, and low-pass filtering its value, the state space of the agent is... Determined as:
[0021] (42) The agent's output is continuously processed. As part of the PID position loop, the command deflection angle Conversion screw precession length Rotor feedback angle Approximate lead screw precession length The optimization value for the deviation between them; (43) To improve the overall performance of the controller, random step instructions, low-frequency sine instructions and high-frequency sine instructions are mixed for training. Each episode selects a task as an instruction with equal probability and continues for 10 seconds. (44) To reduce the angle of the swing mechanism From the perspective of instructions The error is calculated by taking the negative of the absolute value of the difference between the two values as the reward value for each time step.
[0022] (45) Considering the deployment requirements on embedded chips, the actor is set to a 2-layer hidden MLP; the critic network is set to a 4-layer hidden MLP; and the agent sampling time is set to 0.01s during training.
[0023] Beneficial Effects: The method described in this invention addresses the semi-closed-loop control scenario where only the motor feedback signals, such as the rotation angle of the permanent magnet synchronous motor, can be measured, but the actual position of the load mechanism cannot be determined. Considering the high-order nonlinear characteristics of the load model, this invention proposes a method within the FOC control framework of the permanent magnet synchronous motor. This method uses a traditional PID three-loop controller as the foundation and employs a dual-delay depth deterministic strategy gradient algorithm to train and optimize the strategy network. This network observes the feedback quantity of the permanent magnet synchronous motor and outputs the position loop feedback position optimization value, thereby improving the control accuracy and response speed of the traditional PID three-loop control when facing high-order nonlinear load models. Attached Figure Description
[0024] Figure 1 This is a structural diagram of the overall control flow model in this invention; Figure 2 This is a Simulink model of the permanent magnet synchronous motor in this invention; Figure 3 This is an equivalent triangle structure diagram formed between the lead screw edge, the swing mechanism, and the fulcrum in this invention; Figure 4 This is a Simulink model of the load mechanism in this invention; Figure 5 This is a structural diagram of the decoupled current controller based on PID control in this invention; Figure 6 This is a structural diagram of the speed loop and position loop based on the PID controller in this invention; Figure 7(a) and Figure 7(b) are the response curves of the q-axis current controlled by the current controller to the step current command under the conditions of not considering voltage limiting and using voltage limiting, respectively, in this invention. Figure 8(a) shows the error convergence phenomenon of the precession distance of the command angle and the approximate precession distance of the rotating shaft position in the position loop in the present invention. Figure 8(b) is a schematic diagram of the phenomenon that the actual angle and the command angle of the swing mechanism differ greatly in the error convergence phenomenon. Figure 9 This is a structural diagram of the intelligent agent being added to the system control as an offset of the position loop nonlinearity in this invention; Figures 10(a) and 10(b) are schematic diagrams of the oscillation of the agent's output value and the oscillation of the motor state quantity observed by the agent in this invention. Figure 11(a) and Figure 11(b) are the position loop closure curves of the PID scheme and the reinforcement learning scheme in this invention, respectively. Detailed Implementation
[0025] To illustrate the technical solutions disclosed in this invention in detail, the invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0026] First, the key problem addressed by the method described in this invention is how to improve the control algorithm to enhance the tracking accuracy of the actual load position to the command position when performing PID three-loop position control of a permanent magnet synchronous motor, in a semi-closed-loop control situation where only motor parameters such as the motor rotation angle can be measured, but the actual position of the high-order nonlinear load mechanism cannot be measured.
[0027] The main design idea of this invention is based on deep reinforcement learning. It uses a dual-delay deep deterministic policy gradient algorithm to optimize a policy network, enabling the network to use real-time observable or computable feedback values from the motor as state variables. From these, it calculates the closed-loop error of the position loop input error in traditional PID three-loop control or the error caused by threshold limitations in the control. The overall control flow provided by this method is as follows: Figure 1 As shown.
[0028] The construction and training process of a deep reinforcement learning-based dynamic tuning method for electric servo position feedback in a semi-closed loop scenario includes the following steps: Step 1: Constructing a mathematical model of the permanent magnet synchronous motor under the FOC framework In a permanent magnet synchronous motor, the rotor consists of a rotor core and permanent magnets arranged around it. Regardless of the rotor configuration, the magnetic flux (magnetic induction intensity) density distribution generated by the paired magnetic poles in the air gap is similar. In physical modeling, it is typically assumed that the magnetic flux density distribution generated by the permanent magnet poles mounted on the rotor surface or embedded in the core is sinusoidal. Therefore, the fundamental wave of the magnetic flux density curve is taken as the ideal magnetic flux density distribution. Furthermore, for this sinusoidal magnetic flux density signal, its coordinate axis is defined as the magnetic field angle. Define the number of permanent magnet pairs mounted on the rotor core as follows: Mechanical angle of rotor rotation The relationship with the corresponding magnetic field angle is:
[0029] Simultaneously, the polar axis of a permanent magnet pole (the extreme point of the sine wave) is defined as the d-axis. Located between the two magnetic poles, its magnetic field angle differs from the d-axis by a certain angle. Magnetic flux is The point is the q-axis.
[0030] The rotor configuration leads to the classification of permanent magnet synchronous motors into two categories: salient-pole motors and non-salient-pole motors. Salient-pole motors have internal magnets, while non-salient-pole motors have surface-mounted magnets. The reason for caring about the difference between salient-pole and non-salient-pole motors is that the permeability of permanent magnets is almost the same as that of free air, while the permeability of the iron core is much higher than that of air (ferromagnetism).
[0031] According to Abe's Law, the magnetic flux density at a point in space is proportional to the permeability at that point. Therefore, considering a constant magnetic field strength generated by a current-carrying solenoid: when the rotor rotates, the radial length of the magnetic field lines passing through the iron core is the same regardless of the direction of rotation of the surface-mounted motor, meaning the magnetic reluctance in the magnetic circuit is the same; however, for a salient-pole motor, the magnetic circuit has the least amount of iron core passing through it and the greatest magnetic reluctance when the rotor rotates to the d-axis; while it has the most iron core passing through it and the least magnetic reluctance when it rotates to the q-axis, exhibiting a non-uniform magnetic air gap. This phenomenon is called magnetic salient polarity.
[0032] The stator windings of an electric motor are essentially superimposed, energized solenoids with different directions of rotation, and the coils rotate at 120°. The displacement is distributed within the stator slots surrounding the stator core, and these are named the A, B, and C phase windings. In practical circuits, the tails of the A, B, and C phase windings are usually connected to form a delta connection circuit. In this case, we have:
[0033] For the three-phase motor windings mentioned above, a phase difference of 120° is applied respectively. The three-phase alternating current and the three-phase sinusoidal time-varying current can be combined in space to form a rotating magnetic field.
[0034]
[0035] For a rotating rotor, if the rotating magnetic field can maintain the same rotational speed as the rotor and the magnetic phase is constant, the interaction between their magnetic fields can generate a constant torque, namely magnetic torque. For motors with salient polarity, another type of torque, namely reluctance torque, will also be generated. These torques drive the rotor to rotate along with the load.
[0036] Using the direction of the square magnetic field generated by the three-phase windings abc described above as the coordinate axis direction, we define the abc reference coordinate system. In this coordinate system, the phase variable in the time domain can be expressed as... , , ,in, It can represent phase voltage, phase current, and magnetic flux linkage. Considering Faraday's law of electromagnetic induction and Ohm's law, the three-phase voltage can be expressed as:
[0037] Based on the above discussion of the rotor, for a surface-mounted rotor motor, the permeability of the self-inductance and mutual inductance of each stator coil remains unchanged. The relationship between the self-inductance and mutual inductance of any rotor configuration and the magnetic angle is as follows:
[0038]
[0039] For surface-mount motors, Therefore, the flux linkage of the three-phase windings abc is considered to be the sum of the self-inductance and mutual inductance flux linkages, plus the flux linkage leaked from the permanent magnet into the coil:
[0040] in The maximum flux linkage supplied by the N pole of a permanent magnet to a coil is given by substituting the above two equations:
[0041] The motor model described above in the stationary reference frame suffers from parameters that change over time, complicating the design of the control system. This control complexity caused by rotation can be addressed by projecting the phase variables of the model onto a two-term model in the rotating reference frame. The coordinate system has two orthogonal axes fixed on the rotor, namely the d-axis of the rotor's permanent magnet poles and the q-axis orthogonal to it.
[0042] To transform the motor model from a three-phase stationary abc reference coordinate system to a two-phase rotating coordinate system, we first need to know the angles of the dq axes relative to the stationary abc coordinate system. Then the three-phase variables are projected onto the d-axis to obtain... Projecting onto the q-axis yields Mathematically, the Park transformation can be used to solve for the transformation process, and its matrix form is as follows:
[0043] The coefficients This is to ensure that the magnitude of the transformation remains equal after the transformation, and because for the phase variable:
[0044] Therefore, the above transformation matrix, after adding this constraint, becomes invertible. The transformation from the dq reference coordinate system to the abc reference coordinate system can be achieved by the following formula:
[0045] in The component is 0. FOC control transforms the three-phase rotating magnetic field into a rotor magnetic field through changes in Park. When the relative rotation of the shaft changes during motor startup or sudden load conditions, the three-phase variables abc are unbalanced sinusoidal signals. The phase variable is generally a time-varying type; when the motor is running in a steady state, the rotating magnetic field created by the phase variable remains relatively stationary with respect to the rotor. The axis variables then become some DC signals. In this case, control... The variable of the axis is equivalent to two equivalent solenoids that are always aligned with or perpendicular to the magnetic axis, and the corresponding motor control algorithm becomes relatively simple.
[0046] Transforming the above equation using the Park variation, we obtain the voltage equation for the rotor reference frame as follows:
[0047] in and These are the stator voltages along the d-axis and q-axis, respectively. and These are the stator currents along the d-axis and q-axis, respectively. and Let be the stator flux linkages along the d-axis and q-axis, respectively, and their values be:
[0048] in and These are the d-axis and q-axis inductances, respectively. The d-axis magnetic flux is the magnetic flux directly opposite the magnetic pole.
[0049] In the dq coordinate system, the expressions for the self-inductance and mutual inductance coefficients of each phase, which originally changed with rotation, become constants. Combining the two equations yields the motor current-voltage equation:
[0050] in, , These are the system inputs (control variables), and their input determines the current and torque. In the rotor coordinate system, the instantaneous input power of the permanent magnet synchronous motor during operation is:
[0051] in, This is to compensate for the coefficients multiplied during the Park transformation. Note that... This term is the resistance voltage drop term and does not contribute to the final motor output power; The term represents the magnetic field voltage drop; this electrical power is stored in the magnetic field and therefore does not contribute to the final motor output power. Therefore, the actual value of electrical energy converted into mechanical power is:
[0052] According to the torque theorem:
[0053] The electromagnetic torque is:
[0054] When a permanent magnet synchronous motor is connected to a mechanical load, the dynamics of the motor's mechanical components are described by the following equation:
[0055] in For load torque, The viscosity coefficient of the motor bearing. This is the total moment of inertia of the motor and the load.
[0056] Ultimately, the permanent magnet synchronous motor can be built as a Simulink submodule for easy access, such as... Figure 2 As shown. The voltages of the d and q axes input to this module are determined according to the external load torque. It outputs the motor angle and angular velocity, and obtains the dq axis current through a current detector.
[0057] Based on the above analysis, for the actual motor model, this invention constructs a mathematical model of permanent magnet synchronous motor operation based on the FOC control framework and implements it in Simulink.
[0058] Under FOC control, the motor system will receive power from the controller. The voltage command values of the two shafts are used as inputs, and the input torque to the load mechanism is expressed as the rotation angle of the motor rotor.
[0059] The FOC (Field-Oriented Control) strategy for permanent magnet synchronous motors decomposes the motor's phase variables into magnetic field and torque components, and controls them independently to achieve precise control of the motor's magnetic field and torque. It requires acquiring the motor's rotor position and speed information for the phase variables in the stationary coordinate system. Calculation of transformations between relative rotational variables of coordinates. In the coordinate system, the motor equations can be described as follows:
[0060] in: Stator resistance; , They are respectively axis, Shaft voltage; , They are respectively axis, shaft current; , They are respectively axis, Shaft inductance; Permanent magnet Axial flux; , These are the output torque and the load torque, respectively. The bearing viscosity coefficient; The total moment of inertia of the motor and load; , These are the rotor mechanical angular velocity and the rotor electromagnetic angular velocity, respectively. For the number of permanent magnets, satisfying This module responds to input. The voltage across both axes depends on the external load torque. Output motor angle and angular velocity And obtained by current detector Axis current.
[0061] Based on actual needs, the motor type is determined to be a surface-mounted PMSM. Considering that permanent magnet synchronous motors operating in steady state must remain under certain operating limits, this invention requires that the rotor speed and stator current be kept within a threshold range, i.e., satisfying:
[0062]
[0063]
[0064] in, This represents the maximum rotor speed. This represents the maximum value of the stator current.
[0065] Step 2: Construct a mathematical model of the load and its transmission mechanism.
[0066] Based on practical requirements, a high-order load mechanism with nonlinear factors is determined. The mechanism transmits the motor shaft torque to a lead screw via gears. The lead screw's stroke, through a push-pull triangular linkage, forms a lever arm, which in turn drives the oscillating mechanism to deflect at an angle. In the modeling process, to accurately reflect the mechanical properties of the actual load, the elastic motion of the lead screw, the motion equations of the oscillating mechanism, and the nonlinear Coulomb friction during the motion are considered. The dynamic equations of this load can be described as follows:
[0067] in, For rotor mechanical angle, This is the gear reduction coefficient. This is the lead screw reduction coefficient. This refers to the amount of lead screw retraction / extension. For rotational stiffness, The force exerted by the torque on the lead screw; For the mass of the lead screw, For elastic damping, This refers to the compression stroke of the lead screw; This refers to the load torque on the motor shaft. For transmission efficiency; For the oscillating torque, For precession combined stiffness, The length of the lever arm; The swing angle of the swing mechanism. For the moment of inertia of oscillation, For oscillation damping, For positional resistance torque, The frictional torque is modeled as Coulomb friction, and its expression is as follows:
[0068] The above load model considers the elastic deformation of the lead screw. If we only consider the geometric relationship of the triangular linkage mechanism, the equivalent triangular structure formed between the lead screw side, the oscillating mechanism, and the fixed fulcrum can be represented by... Figure 3 This indicates that, without considering the elastic deformation of the lead screw, the lead screw's retraction / extension amount... With deflection angle The relationship can be approximated as follows:
[0069] in , These are the two adjacent sides of the swing angle in a triangular linkage mechanism. , long; , These are the pendulum angles when the load deflection angle is 0. and edge length.
[0070] The overall load is modeled as a high-order nonlinear system. There is a complex differential equation relationship between the load swing angle and the shaft angle of the motor observable. This makes it difficult for traditional error-based closed-loop controllers to achieve good accuracy and response speed performance in this scenario, leaving room for optimization of data-based artificial intelligence algorithms.
[0071] Build an equivalent load module in Matlab / Simulink to implement the above system of equations, such as... Figure 4 As shown. This module will adjust the motor rotor angle. As input, the output is fed back to the motor load torque. and control quantity .
[0072] Step 3: Constructing a basic PID three-loop controller and its solution
[0073] In this invention, the PID controller calculates the difference between the external input reference value of the controlled object and the feedback value of the measuring element, and then linearly superimposes the difference, the integral value over time, and the derivative value as the output control quantity for the next actuator. Input error. and output control value The relationship between the two is as follows:
[0074] in , , It is the proportional coefficient of the three phases, which can be adjusted manually or automatically through parameter tuning or other optimization algorithms.
[0075] Within the FOC framework, vector control is achieved through control... and The method of controlling torque by current. Therefore, vector control relates to the innermost control layer of the motor drive system, and subsequent speed and position control should be based on current control. For the surface-mounted permanent magnet synchronous motor considered in this invention:
[0076] Electromagnetic torque is A linear function, and It has no effect on torque, any The work done by shaft current results in wasted input power (mainly dissipated by resistance and magnetic field). Therefore, control... Control the torque while maintaining This achieves maximum torque-to-current ratio control (MTPA), meaning that the maximum torque output can be achieved under any stator current.
[0077] In the rotor reference coordinate system, the motor model is subjected to a speed-voltage term (i.e., and The cross-coupling effect of PI controllers is significant, especially at high speeds, where it dominates the voltage equation. This practically weakens the performance of PI controllers, thus requiring a decoupling circuit for the current control scheme in vector control. To achieve this... and Control linearization can be achieved by providing the d-axis voltage and q-axis voltage respectively from a combination of the following two signals:
[0078] in, and It can be controlled by linear PI current control, without nonlinear terms. and It can be calculated from the rotor speed value of the encoder:
[0079] Based on the above decoupled current controller, the feedback motor speed and speed command The difference is used as the input to the speed PI controller, and the controller's output is used as... Reference value for shaft current. Also note the system's speed limit; use the motor speed reference value. Limited to Within the specified range. Based on the speed controller, a position control loop is constructed. Considering the triangular linkage structure of the load, the deflection angle of the commanded oscillating mechanism is determined. The cosine formula described above can be used to approximately convert it into the lead screw precession length. and compare it with the rotor angle feedback signal. Converted precession length The error calculation is used as the input to the position loop PI controller, and the speed loop reference value is used as the output.
[0080] The Simulink modeling of the decoupled current control in the aforementioned PID three-loop controller is as follows: Figure 5 As shown, the modeling of the position loop and the velocity loop is as follows: Figure 6 As shown, the PMSM three-loop PI control model is a linear control that controls the angle of the actual swing mechanism. While ensuring system robustness, it can meet certain requirements in terms of accuracy and response speed. Furthermore, due to the characteristics of PI control, it can ensure that the motor current and speed will not exceed system limits. However, the above three-loop control flow ultimately performs poorly in terms of control performance, mainly due to the following issues: On the one hand, in actual systems, the equivalent voltage of the SVPWM output cannot exceed the inverter's limit (220V), therefore it is necessary to... Figure 5 The output value of the decoupled current controller shown is limited, and this limiting function disrupts the linearity of the linear term of the decoupled current controller. Figures 7(a) and 7(b) illustrate the control behavior of the controller before and after considering the threshold limit. The response curve of the shaft current to the reference value of the step current command shows that, under the same controller parameter configuration, without threshold limitation... The axis current can track the command size relatively quickly, but after adding a threshold limit, The shaft current curve exhibits high irregularity, large overshoot, and long convergence time. This nonlinearity reduces the response speed of the current loop to some extent and increases the risk of the actual current exceeding system limits, thereby reducing the performance and reliability of the controller.
[0081] On the other hand, the system cannot observe the actual angle of the swing mechanism; it can only transmit the commanded angle. Rotor position feedback Approximate conversion to lead screw process and However, without considering the complex differential equation relationship between the motor rotor angle and the swing mechanism angle, the system exhibits a position loop error that is much larger than the actual swing angle when facing low-frequency signals, as shown in Figures 8(a) and 8(b).
[0082] Step 4: Construct a position feedback dynamic optimization scheme based on deep reinforcement learning
[0083] Training was performed using TD3, and the agent's predicted actions were defined as the instruction deflection angle within the PID position loop. Conversion screw precession length Rotor feedback angle Approximate lead screw precession length The optimization value for the deviation between them. Simultaneously, the agent's state input is determined as the command deflection angle. Calculated lead screw precession length and rotor feedback angle Approximate lead screw precession length as well as Rate of change and speed feedback .
[0084] To better optimize the training algorithm, the input values are quantized and limited. Simultaneously, the two feedback quantities are... and Add a Butterworth low-pass filter.
[0085] To comprehensively improve controller performance, the tasks faced by the reinforcement learning controller should be as diverse as possible to include all possible instruction patterns, including random step instructions, low-frequency sine instructions, and high-frequency sine instructions.
[0086] To reduce the angle of the swing mechanism From the perspective of instructions The error is used as the reward value for each time step, with the absolute value of the difference between them being the opposite of the sum:
[0087] State Use optimization strategy function This enables the agent to output tuning values. Therefore, two [entities / facilities] need to be established. The network evaluates the output of the policy network and uses the evaluation value to perform gradient optimization on the actor network.
[0088] Meanwhile, the agent sampling time was set to 0.01s.
[0089] Specifically, the present invention includes the following considerations: 1) This invention installs the intelligent agent at the position loop calculation deviation point, adding it as a non-linear offset to the system control, such as... Figure 9 As shown. Note that for the amplitude limit at... Angle commands within the range are converted using Equation 7. And with approximate precession feedback The difference item after subtraction is generally in Within the range. Therefore, the predicted offset is limited to... This is quite reasonable. This range allows for adjustment of the position loop to obtain a better position deviation, while also preventing the motor from becoming completely uncontrolled by the PID controller and exhibiting unsafe behavior under the black-box control of the neural network due to a large control range.
[0090] 2) The state-space variables in reinforcement learning should be values that the controller can collect or compute. Specifically, because... shaft current , shaft current It is already managed by a decoupled current controller. Close to 0, while Because the aforementioned voltage threshold limitations can cause numerical oscillations, introducing them as state values for the agent makes it difficult to learn the control logic and may even lead to performance degradation. Therefore, they are not used as observation values for reinforcement learning agents. As for the instruction deflection angle... Calculated lead screw precession length and rotor feedback angle Approximate lead screw precession length Highly correlated with the final deviation prediction, it is necessary to incorporate them into the state variables. Simultaneously, the controller should consider the movement trends of the load and target commands, therefore... Rate of change and speed feedback As the state value of the intelligent controller.
[0091] 3) To better optimize the training algorithm, the input values are quantized so that the input values remain consistent across various operating states. This order of magnitude. Meanwhile, for positional instructions involving mutations (such as step jumps), the derivative value becomes a very large number, leading to ill-conditioned empirical samples; therefore, the value of its derivative needs to be limited to a certain range.
[0092] Due to the highly nonlinear nature of neural networks, at the beginning of training, the network output often reaches boundary values and oscillates between the upper and lower boundaries. In the motor model, oscillations in the speed command lead to oscillations in the current command, which in turn cause oscillations in the motor feedback value, further leading to oscillations in the reinforcement learning controller, ultimately placing the motor in a highly unstable state, as shown in Figures 10(a) and 10(b). Using such ill-conditioned data as the agent's experience values often fails to yield any learning. In many cases, even after numerous training rounds, regardless of whether the agent achieves the controller's effect, the final output still oscillates repeatedly, which is unacceptable in practical motor control and even risks damaging the motor. To mitigate this situation, the two feedback values of the motor can be adjusted... and Adding a Butterworth low-pass filter will Add passband The second-order low-pass filter will Add passband The first-order low-pass filter is used. After this operation, even if system oscillations occur at the beginning of training, the presence of the low-pass filter will filter the input state into a low-frequency signal, and gradually make the output a low-frequency signal during training, thus ensuring the stability of training to a certain extent.
[0093] 4) To comprehensively improve controller performance, the tasks faced by the reinforcement learning controller should be as diverse as possible to include all possible instruction modes. According to automatic control principles, common performance analysis methods for linear systems include: examining the response capability to step commands in the time domain or examining the open-loop amplitude-frequency characteristic and open-loop logarithmic phase-frequency characteristic in the frequency domain. Based on the final performance evaluation scheme, the training tasks are divided into the following three types: a. Random Step Command: Using the zero-position steady state (all derivatives are 0) as the initial state, a new random step target is generated every 2.5 seconds to ensure the step range. .
[0094] b. Low-frequency sine wave command: Zero-position steady state is used as the initial state to generate amplitude. angular frequency Low-frequency sine wave command .
[0095] c. High-frequency sine wave command: Zero-position steady state as the initial state, generating amplitude. angular frequency High-frequency sine wave command .
[0096] During training, each episode selects a task as an instruction with equal probability and continues for 10 seconds.
[0097] 5) TD3 predicts the cumulative reward value over multiple future steps using a Q-network, and the setting of the reward function directly affects the optimization of the algorithm. Considering that the objective of this invention is to optimize the angle of the swing mechanism... From the perspective of instructions The error can be minimized, so the absolute value of the difference between them can be used as the reward value for each time step. Such a reward is not a sparse reward, and the range of the reward can be stabilized by offset and scaling. Inside, it may be beneficial for training.
[0098]
[0099] 6) During training, for the input agent state values:
[0100] The goal is to train an optimal policy function. , making for Able to output optimal compensation value:
[0101] Therefore, it is necessary to establish The network evaluates the output of the policy network to optimize its output. This is because, as seen in both DQN and DDPG, [the following occurs]. In cases where the value is overestimated, leading to a final performance degradation, the TD3 algorithm randomly initializes... and two Network and a Policy network, in which , , These are the network parameters. A copy of each of the three networks is then created to obtain the desired network parameters. , , , as the target network for optimization.
[0102] For each step index of time T and agent time step Ts. Agent sampling of state 4.4 And a noisy one is generated using a policy network. Actions:
[0103] The next state resulting from this action Including the state of this step ,action And the reward value for this step calculated by the environment. As training samples Experience replay cache In the middle. When The number of samples in the batch reaches a mini-batch of size N, and each step starts from... The following optimizations were made to a batch of samples from the middle sampling process;
[0104]
[0105] Here we use and To mitigate the overestimation of Q-values, the smaller of the two Q-values is used. The Q-value at step i+1 is estimated using a fixed target network, multiplied by a discount factor, and then added to the environmental reward at step i. This gives the Q-values predicted by both Q-networks at step i. , The target value at that time, that is, for , Update its parameters as follows:
[0106] Because of one of them The target network has fixed parameters, ensuring stable training through parameter updates. These updates allow the two Q-networks to gradually approximate the true Q-function. After training the Q-evaluation network, the policy network needs to be updated to optimize action decisions. This is done every *d* times the Q-network is updated using a direct gradient method. , The gradient value is:
[0107] Simultaneously with a smooth ratio Update the parameters of the target network:
[0108]
[0109] 7) Considering the deployment requirements on embedded chips, the actor network should not be set too large. This invention sets it to a 2-layer MLP, while the critic network is set to a 4-layer MLP. Meanwhile, considering that the frequency of the sinusoidal command input to the controller can reach 40 rad / s, and the training speed should not be too slow, setting the agent sampling time to 0.01s is appropriate. Other training parameter settings are as follows:
[0110] To comprehensively evaluate the performance improvement of the reinforcement learning-based position loop tuning scheme compared to the traditional PID method, the following three metrics are used to assess the performance.
[0111] 1) Load location characteristics
[0112] Under maximum load conditions, the angle command sequence As command input, data collection Actual swing angle sequence of the oscillating mechanism within the time range .in:
[0113] Will As the x-axis, The position loop curve is plotted on the ordinate. The nominal position curve is the line connecting the midpoints of the position loop curves on the horizontal axis, and the nominal position baseline is a first-order linear fit of the nominal position curve. The tracking accuracy in the time domain is analyzed using loop width and zero-position deviation algorithms. The maximum swing angle in both positive and negative directions characterizes the tracking capability when approaching the command limit position; The maximum loop width represents the maximum value of the tracking error; Zero-position deviation measures the symmetry of the control algorithm in the positive and negative directions when facing a symmetrical signal such as a sinusoidal low-frequency signal.
[0114] Combining Figure 11(a), Figure 11(b) and the table below, the RL method outperforms the PID method in all tracking accuracy metrics. However, the linear PID method is better than the RL method in terms of symmetry of sinusoidal commands, but it is still within the required range.
[0115]
[0116] 2) Velocity Characteristics Experiment
[0117] Under maximum load conditions, the amplitude is The step instruction is used as system input. The sampling angle is located at The actual position sequence of the load over a time period within a range, expressed as the average angular velocity of the swing. The time-domain response speed of the algorithm was analyzed, and the experimental results are shown in the table below. It can be seen that the reinforcement learning scheme responds faster to step-change signals than the linear method like PID.
[0118]
[0119] 3) Frequency response experiment
[0120] Under maximum load conditions, the sine command will be used. As ,in , Six command cycles were simulated for each frequency and amplitude. The phase attenuation of the output swing angle compared to the input command was measured. and amplitude attenuation The response characteristics of the algorithm in the frequency domain are analyzed using the following formula. and In Orthogonal decomposition is performed on the sine and cosine bases of the same frequency to obtain... and sine and cosine components , and , Calculate the relative standard excitation amplitude. and phase angle .
[0121]
[0122]
[0123] The gain of the actual angle signal with respect to the command signal is calculated using the above formula. and phase lag The results are recorded below.
[0124] Experiments show that the controller using the RL method generally outperforms the PID method in terms of phase decay performance.
[0125]
[0126] The above three test methods are used to quantitatively measure the control performance of the original PID method and the RL method in various indicators. The comparison shows that the RL method is better than the original PID method in most indicators, and the system has achieved better tracking accuracy and response speed.
Claims
1. A dynamic optimization method for electric servo position feedback based on deep reinforcement learning in a semi-closed loop scenario, characterized in that: The method described above is based on a permanent magnet simulation model under a high-order nonlinear load and FOC control framework. On the basis of PID three-loop control, the state space and action settings of the reinforcement learning agent are determined. Then, with the goal of improving the tracking accuracy and response speed of the load swing angle to the command target angle under semi-closed-loop control, the system model is modeled using simulation software and the optimization strategy network of motor position feedback is obtained by using deep reinforcement learning method to output the optimal optimization value. The method includes the following steps: (1) Construct a permanent magnet synchronous motor operation model based on the FOC control framework. Under FOC control, the motor system will receive power from the controller. The voltage command values of the two shafts are used as inputs, and the input torque to the load mechanism is expressed as the rotation angle of the motor rotor. (2) Construct a mathematical model of the load and its transmission mechanism. The motor load model considers the elastic deformation of the transmission mechanism, the dynamic equation of the swing mechanism, and also the nonlinear factors including Coulomb friction and torque transmission in the triangular linkage mechanism. (3) Considering the PMSM and high-order nonlinear load model in the above modeling, a three-loop control scheme based on PID controller is constructed on the FOC control framework; (4) The policy network is trained by reinforcement learning algorithm, and the system performance is improved by changing the input of the position loop PID controller, including the use of TD3 reinforcement learning algorithm with continuous state and action space for agent optimization; For basic PID three-loop control in semi-closed-loop control scenarios, the TD3 reinforcement learning algorithm, which has continuous state and action spaces, is used to optimize the agent, enabling it to predict the lead screw precession length. Approximate lead screw precession length The optimized value of the deviation between them is used to alleviate the problem of insufficient control accuracy caused by position feedback error, wherein: (41) The value observed by the agent is determined as the instruction deflection angle. Calculated lead screw precession length and rotor feedback angle Approximate lead screw precession length as well as Rate of change and speed feedback After quantizing, saturating, and low-pass filtering its value, the state space of the agent is... Determined as: (42) The agent's output is continuously processed. As part of the PID position loop, the command deflection angle Conversion screw precession length Rotor feedback angle Approximate lead screw precession length The optimization value for the deviation between them; (43) To improve the overall performance of the controller, random step instructions, low-frequency sine instructions and high-frequency sine instructions are used for mixed training. Each episode selects a task as an instruction with equal probability and continues for 10 seconds. (44) In order to reduce the angle of the swing mechanism From the perspective of instructions The error is calculated by taking the negative of the absolute value of the difference between the two values as the reward value for each time step. (45) Considering the deployment requirements on embedded chips, the actor is set to a 2-layer hidden MLP; the critic network is set to a 4-layer hidden MLP; and the agent sampling time is set to 0.01s during training.
2. The method for dynamic optimization of electric servo position feedback based on deep reinforcement learning in a semi-closed-loop scenario according to claim 1, characterized in that: The behavior of the described permanent magnet synchronous motor operating model in the FOC control framework was simulated using simulation software, where the control and response variables, as well as the associated constraints, are described by the following set of variables and equations: (11) The coordinate system under relative rotation , Voltage of two axes and As a system input, the drive motor applies torque to the transmission mechanism, which manifests as the rotation angle of the motor rotor. At the same time, the load and transmission device apply a time-varying load torque to the motor. ; (12) The controller obtains the rotation angle of the motor rotor. and angular velocity As a calculation parameter, the current sensor simultaneously detects the stator current. , , And use Clarke transform and Park transform to convert it to , Current values of both axes , ; (13) Based on the actual motor model, the motor model is determined to be a surface-mounted permanent magnet synchronous motor, satisfying: Considering the rotor speed of the permanent magnet synchronous motor operating in steady state and stator current It should be kept within the threshold range, satisfying: in, This represents the maximum rotor speed. This represents the maximum value of the stator current. (14) The FOC strategy decomposes the phase variables of the permanent magnet synchronous motor into magnetic field components and torque components, and controls them independently. Precise control of the motor's magnetic field and torque requires obtaining the rotor position and speed information of the motor for the phase variables in the stationary coordinate system. The transformation calculation between relative rotation variables of coordinates, in In the coordinate system, the motor equations are described as follows: in: Stator resistance; , They are respectively axis, Shaft voltage; , They are respectively axis, shaft current; , They are respectively axis, Shaft inductance; Permanent magnet Axial flux; , These are the output torque and the load torque, respectively. The bearing viscosity coefficient; This is the total moment of inertia of the motor and the load. , These are the rotor mechanical angular velocity and the rotor electromagnetic angular velocity, respectively. For the number of permanent magnets, satisfying The permanent magnet synchronous motor operating model responds to the input. The voltage across both axes depends on the external load torque. Output motor angle and angular velocity And obtained by current detector Axis current.
3. The method for dynamic optimization of electric servo position feedback based on deep reinforcement learning in a semi-closed-loop scenario according to claim 2, characterized in that: For the specific scenario of semi-closed-loop control, a dynamic model of the transmission and load mechanism is built using Simulink, taking into account higher-order nonlinear factors in the model, including: (21) The driving torque of the motor shaft is transmitted through a transmission ratio of The two-stage gear transmission drives the lead screw to rotate, and the lead screw rotates at... The reduction ratio produces a corresponding process, in which rotation generates a force on the elastic screw. for: In the formula, Indicates the rotational stiffness; (22) Considering the mass as Elastic damping is The combined stiffness is The elastic motion of the lead screw under compression, its dynamic equation is: in, This refers to the compression stroke of the lead screw; (23) The elastic deformation of the lead screw generates a force that pushes the lead screw side of the triangular linkage mechanism, causing the lever arm of the load side to generate a torque that causes the load to deflect around the axis. Among them, the extension / retraction amount of the lead screw With deflection angle The relationship can be approximated as follows: in , These are the two adjacent sides of the swing angle in a triangular linkage mechanism. , long; , These are the pendulum angles when the load deflection angle is 0. and edge length; (24) Consider a oscillating moment of inertia as The oscillation damping is The positional resistance torque is The frictional torque is The swing mechanism receives a size of The torque, its dynamic equation is: The frictional torque is modeled as Coulomb friction: 。 4. The method for dynamic optimization of electric servo position feedback based on deep reinforcement learning in a semi-closed-loop scenario according to claim 2, characterized in that: For a semi-closed-loop electric servo system under the FOC control framework, a three-loop control based on a PID controller is used to perform basic control on the final position and angle, ensuring the stability and robustness of the system operation. Specifically: (31) The current loop controlling the torque is constructed as a decoupled current controller, which decouples the originally coupled current loops. , The two-axis voltage term is decomposed into linear and nonlinear terms: in, and It is controlled by a linear PID current controller, without nonlinear terms. and The rotor speed value is calculated from the encoder. ; (32) Input error in PID controller and output control value The connection between the two: ; (33) Feedback motor speed and speed command The difference is used as the input to the speed PID controller, and the controller's output is used as... Reference value for shaft current; (34) Based on the speed controller, a position control loop is constructed, specifically a load-based triangular linkage structure, which controls the command deflection angle. Approximately converted to lead screw precession length and compare it with the rotor angle feedback signal. Converted precession length The error value is used as the input to the position loop PID controller, and the speed loop reference value is used as the output.
Citation Information
Patent Citations
Superhigh precision servo driving system based on PID online calibration machine tool
CN105824290A
Machine learning apparatus and method and motor control apparatus
CN106815642A