Electromagnetic bearing rotor system unbalance control method based on DDPG algorithm
The DDPG algorithm optimizes the imbalance control of the electromagnetic bearing rotor system, and solves the problems of complex calculations and poor compensation in the prior art, and achieves fast and accurate imbalance compensation.
Patent Information
- Application Number
- CN202510655088.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-21
AI Technical Summary
In the active electromagnetic bearing rotor system, the existing control method has complex calculations and poor compensation effect.
The deep reinforcement learning algorithm DDPG is adopted to optimize the identification and compensation of imbalance coefficients by constructing compensation current expressions and adaptive identification equations, and combining multi-objective reward functions.
It realizes the effects of simple calculation, fast compensation speed and high compensation accuracy. It is suitable for low and high speed scenarios, and widens the scope of application of the algorithm.
Smart Images

Figure CN120217899A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of active electromagnetic bearing rotor system control, and specifically relates to an imbalance control method for an electromagnetic bearing rotor system based on the DDPG algorithm. Background Art
[0002] The active electromagnetic bearing (AMB) enables the rotor to levitate through controllable electromagnetic force. Due to its advantages such as long lifespan and suitability for high rotational speeds, it is widely used in fields such as high-speed motors, centrifugal compressors, and artificial heart pumps. In the electromagnetic bearing-rotor system, due to issues such as material and assembly processes, the center of mass of the rotor deviates from the geometric center. During the rotation of the rotor, the center-of-mass deviation will cause an unbalanced force, which acts on the rotor in the opposite direction, causing the rotor to generate unbalanced vibrations. The unbalanced force increases continuously with the increase in rotational speed, leading to intensified unbalanced vibrations, reducing the stability of rotor levitation, and even causing instability. Therefore, it is necessary to suppress the unbalanced vibrations of the rotor.
[0003] Currently, the control methods for imbalance compensation in active electromagnetic bearing rotor systems are mainly divided into two categories: one is the imbalance compensation method based on identifying the position of the unbalanced mass of the rotor, which calculates the unbalanced vibration compensation signal in real time according to the position of the unbalanced mass of the rotor and the rotational speed of the rotor. However, when the frequency and amplitude of the unbalanced disturbance increase, the convergence speed of the search algorithm will slow down, resulting in a worse compensation effect; the other is to use the adaptive control method, such as using neural networks, adaptive filters based on LMS, etc. However, the search algorithm is relatively complex, has high requirements for the performance of the controller, and the actual compensation effect is poor.
[0004] The present invention constructs a compensation current expression based on the external disturbances received by the rotor, extracts the imbalance coefficient therein, designs appropriate observation variables using the DDPG algorithm, and combines multiple indicators as the reward function to accurately identify the imbalance coefficient, thereby obtaining the optimal compensation current. Compared with the existing inventions, the compensation algorithm proposed by the present invention has the characteristics of simple calculation, good compensation effect within the full rotational speed range, and high compensation accuracy. Summary of the Invention
[0005] The purpose of the present invention is to provide an imbalance control method for an electromagnetic bearing rotor system based on the DDPG algorithm. By combining the imbalance coefficient reflected by the unbalanced force received by the rotor at both ends of the rotor with the deep reinforcement learning algorithm DDPG for identification, the calculation amount is reduced, the identification speed of the imbalance coefficient and the compensation efficiency of the imbalance compensation algorithm within the full rotational speed range are improved, and the rapid stability and high-efficiency compensation of the rotor system are achieved.
[0006] The technical solution adopted by the present invention is: an imbalance control method for an electromagnetic bearing rotor system based on the DDPG algorithm, including the following steps: S1: Establish the four-degree-of-freedom AMB-rigid rotor dynamic equation based on rotor dynamics, and achieve decoupling control of the active electromagnetic bearing rotor system through linear state feedback decoupling; S2: Construct a compensation current according to the unbalanced forces acting on the four degrees of freedom of the rotor, extract the unbalanced coefficients in the compensation current, and construct an adaptive identification equation; S3: Construct a DDPG agent, use the rotor position error, the integral of the rotor position error, and the derivative of the rotor position error as the observed variables of the system speed, and design a multi-objective reward function; S4: Use the DDPG agent to optimize the undetermined parameters in the adaptive identification equation, and dynamically adjust the compensation current in combination with the rotor speed to cancel the unbalanced force disturbance.
[0007] Furthermore, the specific expression of the four-degree-of-freedom AMB-rigid rotor dynamic equation is: ; ; where, J represents the transverse moment of inertia, J z represents the moment of inertia of the rotor about the z-axis, represents the rotor rotation speed, L bA represents the distance from the plane of the electromagnetic bearing at end A of the rotor to the plane of the rotor centroid, f xa represents the electromagnetic force of the electromagnetic bearing A on the rotor in the x direction, L bB represents the distance from the plane of the electromagnetic bearing at end B of the rotor to the plane of the rotor centroid, f xb represents the electromagnetic force of the electromagnetic bearing B on the rotor in the x direction, u z is the projection of the distance between the centroid of the unbalanced mass and the rotor centroid on O z , ε z is the projection of the distance between the unbalanced mass point and the centroid on the O z axis, m represents the rotor mass, represents the acceleration of the rotor in the x direction, represents the angular acceleration of the rotor about the x-axis, represents the angular velocity of the rotor about the x-axis, represents the angular acceleration of the rotor about the y-axis, represents the angular velocity of the rotor about the y-axis, f yb represents the electromagnetic force of the electromagnetic bearing B on the rotor in the y direction, f ya represents the electromagnetic force of the electromagnetic bearing A on the rotor in the y direction, represents the acceleration of the rotor in the y direction, represents the unbalanced mass, represents the rotor rotation speed, represents the rotor rotation angle, Denote ε z and O x the initial included angle with the O axis, Denote the rotor rotational acceleration, Fu is the generalized unbalance vector received by the electromagnetic bearing rotor system, Fx represents the unbalanced force received by the rotor in the x direction, and Fy represents the unbalanced force received by the rotor in the y direction, Denote the unbalanced force received by the rotor when rotating around the x axis, Denote the unbalanced force received by the rotor when rotating around the y axis.
[0008] Furthermore, the state equation and output equation of the active electromagnetic bearing rotor system are: ; wherein, denotes the first derivative of X with respect to time, X represents the n-dimensional state vector, A represents the n×n order state matrix, B represents the n×l order input matrix, U represents the l-dimensional control vector, Y represents the m-dimensional output vector, and C represents the m×n order output matrix; The active electromagnetic bearing rotor system becomes a decoupled controlled system after linear state feedback decoupling , and the decoupled controlled system is an integral type system. The expressions of the decoupling matrix R and decoupling matrix F of the decoupled controlled system are: ; ; ; wherein, W1 represents the input matrix after decoupling, W2 represents the output matrix after decoupling, W2 -1 represents the inverse matrix of the output matrix W2 after decoupling, T qb represents the lever arm matrix of the electromagnetic bearing-rotor system, K s represents the force-displacement stiffness coefficient matrix of the electromagnetic bearing rotor system, K i represents the force-current stiffness coefficient matrix of the electromagnetic bearing rotor system, M represents the generalized mass matrix of the rotor system, M -1 represents the inverse matrix of the generalized mass matrix M of the rotor system, L f represents the lever arm matrix, and G represents the gyro matrix.
[0009] Furthermore, the specific expression of the compensation current is: ; ; ; ; ; where, I u represents the compensation current, and I d represents the unbalance disturbance, represents the unbalance coefficient in the x-axis direction at the A end of the rotor, represents the unbalance coefficient in the y-axis direction at the A end of the rotor, represents the unbalance coefficient in the x-axis direction at the B end of the rotor, represents the unbalance coefficient in the y-axis direction at the B end of the rotor, and it is defined that and are the unbalance coefficients at the A and B ends of the rotor; The unbalance coefficient at the A end of the rotor and the unbalance coefficient at the B end of the rotor are recorded in complex form: ; ; where, represents the complex form of the unbalance coefficient at the A end of the rotor, represents the complex form of the unbalance coefficient at the B end of the rotor, and j represents the imaginary unit; The identified value of the unbalance coefficient at the A end of the rotor and the identified value of the unbalance coefficient at the B end of the rotor are: ; ; where, represents the identified value of the unbalance coefficient in the x-direction at the A end of the rotor, represents the identified value of the unbalance coefficient in the y-direction at the A end of the rotor, represents the identified value of the unbalance coefficient in the x-direction at the B end of the rotor, represents the identified value of the unbalance coefficient in the y-direction at the B end of the rotor; Then, the adaptive identification equation for the unbalance coefficient at the A end of the rotor is: ; The adaptive identification equation for the unbalance coefficient at the B end of the rotor is: ; where, represents the second derivative of the identified value of the unbalance coefficient in the x-direction at the A end of the rotor, represents the first derivative of the imbalance coefficient identification value in the x direction at the A end of the rotor, x bA represents the displacement in the x direction at the A end of the rotor, y bA represents the displacement in the y direction at the A end of the rotor, k1, k2, and k3 represent the undetermined parameters in the adaptive identification equation of the imbalance coefficient at the A end of the rotor, and k4, k5, and k6 are the undetermined parameters in the adaptive identification equation of the imbalance coefficient at the B end of the rotor represents the second derivative of the imbalance coefficient identification value in the y direction at the A end of the rotor represents the first derivative of the imbalance coefficient identification value in the y direction at the A end of the rotor represents the second derivative of the imbalance coefficient identification value in the x direction at the A end of the rotor represents the first derivative of the imbalance coefficient identification value in the x direction at the B end of the rotor, x bB represents the displacement in the x direction at the B end of the rotor, y bB represents the displacement in the y direction at the B end of the rotor represents the second derivative of the imbalance coefficient identification value in the y direction at the B end of the rotor represents the first derivative of the imbalance coefficient identification value in the y direction at the B end of the rotor
[0010] Furthermore, the DDPG agent includes an action network, an action target network, a critic network, a critic target network, and an experience pool. The training process of the DDPG agent is as follows S301: Initialize the weights of the action network of the active magnetic bearing rotor system imbalance compensation model based on the DDPG algorithm and the weights of the critic network ; ; ; S302: Create an experience pool for storing the experience data generated during the interaction between the DDPG agent and the external environment. Initialize the experience pool capacity to D, the number of training sampling samples to n, and the maximum number of training episodes to Episode S303: Define the state space of the active magnetic bearing rotor system as (s t , a t , r t , s t+1 ), where s t represents the state of the active magnetic bearing rotor system at time t, a t represents the action output by the DDPG agent at time t, r t represents the total reward value obtained by the DDPG agent at time t, and s t+1 represents the state of the active magnetic bearing rotor system at time t + 1 S304: According to the state s of the active electromagnetic bearing rotor system at time t t and the exploration noise N t Generate the action a output by the DDPG agent at time t through the action network t , and update the state s of the active electromagnetic bearing rotor system at time t + 1 t+1 and calculate the total reward value r obtained by the DDPG agent at time t t ; S305: The relevant data group (s t , a t , r t , s t+1 ) obtained after each exploration of the DDPG agent is stored in the experience pool. If the number of samples in the experience pool is greater than the training sampling number, then during training, multiple data groups (s t , a t , r t , s t+1 ) are randomly selected from the experience replay pool for training to update the target parameters, otherwise continue to execute step S304; S306: Increase the time step and conduct the next round of training until the training ends.
[0011] Furthermore, the expression of the observed variable in S3 is: ; where R obserbation represents the observed variable, e error represents the rotor position error, T represents the total time, t represents the current moment, and V rpm represents the system rotation speed; The expression of the multi-objective reward function is: ; ; ; ; where r error represents the accuracy reward, r response represents the response time reward, r overshoot represents the overshoot reward, a1 represents the error weight coefficient, a2 represents the response weight coefficient, a3 represents the overshoot weight coefficient, r set represents the value that the system is set to reach, r final represents the actual value after the system stabilizes, T response represents the time required for the system to stabilize, and r max represents the maximum response value during the operation of the system.
[0012] Furthermore, the steps of updating the target parameters in S305 are as follows: S3051: For the extracted small batch of samples, calculate the target Q value using the evaluation target network. The specific formula is: ; where γ is the discount factor used to correct the online network, a t+1 represents the action output by the DDPG agent at time t+1, is the target state-action value for executing action a t+1 in state s t+1 ; S3052: Update the parameters of the evaluation network of the DDPG agent by minimizing the loss function L. The minimization of the loss function is defined as the mean square error between the current Q value estimate and the calculated target Q value. The specific expression is: ; where, represents the expectation of the data group (s t , a t , r t , s t+1 ); S3053: Adopt a delayed update strategy for updating the action network parameters, and update the action network parameters by the gradient ascent method. The gradient formula is: ; where, represents the gradient of the Q value function with respect to action a t , is the gradient of the policy function with respect to μ i , represents the expected value of the current action, represents the policy at time t; S3054: Adopt a soft update strategy for the DDPG agent, and calculate the target value using the action target function and the evaluation target function. The update methods of the action target network and the evaluation target network are as follows: ; In the formula are the parameters of the action target network, are the parameters of the action network, are the parameters of the evaluation target network, are the parameters of the evaluation network, is the sliding average coefficient, and its value is less than 1.
[0013] The beneficial technical effects of the present invention are as follows: By combining the DDPG algorithm with the compensation algorithm for identifying the imbalance coefficient, the problems of low efficiency in identifying the imbalance coefficient and complex calculation are solved, achieving the effects of simple calculation, fast compensation speed, and high compensation accuracy; the present invention combines the deep reinforcement learning algorithm with the imbalance compensation algorithm, solving the problems of low efficiency in identifying traditional imbalance coefficients and complex parameter selection, and further improving the compensation efficiency of the algorithm while realizing the stable control of the imbalance compensation of the rotor system. Compared with the prior art, the compensation algorithm proposed by the present invention has the characteristics of simple calculation, small occupation of computing resources, high accuracy, and high efficiency in identifying the imbalance coefficient, and can be adapted to scenarios with low and high rotational speeds and high precision requirements for rotor operation, broadening the applicable scenarios of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0015] Figure 1 It is the control schematic diagram of the embodiment of the present invention; Figure 2 It is the transfer function block diagram of the linear state feedback decoupling method in the embodiment of the present invention; Figure 3 It is the DDPG agent training evaluation reward function diagram in the embodiment of the present invention; Figure 4 It is the comparison diagram of the maximum amplitude of the rotor before and after compensation under constant acceleration in the embodiment of the present invention; Figure 5 It is the comparison diagram of the axial displacement of the A end of the rotor before and after compensation at a constant rotational speed in the embodiment of the present invention; among them, (a) is the schematic diagram of the axial displacement of the A end of the rotor before compensation, and (b) is the schematic diagram of the axial displacement of the A end of the rotor after compensation; Figure 6 It is the identification diagram of the imbalance coefficient of the traditional algorithm and the embodiment of the present invention; Figure 7 It is the axial trajectory diagram of the A end and both ends of the B end of the rotor at a constant rotational speed of 10,000 rpm. Among them, (a) is the axial trajectory diagram of the A end of the rotor, and (b) is the axial trajectory diagram of both ends of the B end of the rotor. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] To more clearly understand the above objects, features, and advantages of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the present invention is not limited by the limitations of the specific embodiments disclosed below.
[0017] As Figure 1 and Figure 2 shown, an imbalance control method for an electromagnetic bearing rotor system based on the DDPG algorithm is proposed in an embodiment of the present invention, including the following steps: S1: Establish a four-degree-of-freedom AMB-rigid rotor dynamic equation according to rotor dynamics, and achieve decoupling control of the active electromagnetic bearing rotor system through linear state feedback decoupling.
[0018] The specific expression of the four-degree-of-freedom AMB-rigid rotor dynamic equation is: ; ; where J represents the transverse moment of inertia, J z represents the moment of inertia of the rotor about the z-axis, represents the rotor rotation speed, L bA represents the distance from the plane of the electromagnetic bearing at end A of the rotor to the plane of the rotor centroid, f xa represents the electromagnetic force of the electromagnetic bearing A on the rotor in the x direction, L bB represents the distance from the plane of the electromagnetic bearing at end B of the rotor to the plane of the rotor centroid, f xb represents the electromagnetic force of the electromagnetic bearing B on the rotor in the x direction, u z is the projection of the distance between the centroid of the unbalanced mass and the rotor centroid on O z , ε z is the projection of the distance between the unbalanced mass point and the centroid on the O z axis, m represents the rotor mass, represents the acceleration of the rotor in the x direction, represents the angular acceleration of the rotor about the x-axis, represents the angular velocity of the rotor about the x-axis, represents the angular acceleration of the rotor about the y-axis, represents the angular velocity of the rotor about the y-axis, f yb represents the electromagnetic force of the electromagnetic bearing B on the rotor in the y direction, f ya represents the electromagnetic force of the electromagnetic bearing A on the rotor in the y direction, represents the acceleration of the rotor in the y direction, represents the unbalanced mass, represents the rotational speed of the rotor, represents the rotational angle of the rotor, represents ε z and O x the initial included angle with the O axis, represents the rotational acceleration of the rotor, Fu is the generalized unbalance vector acting on the electromagnetic bearing rotor system, Fx represents the unbalanced force acting on the rotor in the x direction, and Fy represents the unbalanced force acting on the rotor in the y direction. represents the unbalanced force acting on the rotor during rotation about the x axis, represents the unbalanced force acting on the rotor during rotation about the y axis.
[0019] Rewriting the above equation into matrix form gives: ; where M represents the generalized mass matrix of the rotor system; represents the 4-degree-of-freedom displacement vector of the rotor's center of mass, where represents the angle of rotation about the y axis, x represents the displacement in the x direction, represents the angle of rotation about the x axis, and y represents the displacement in the y direction; G is the gyro matrix; F xy is the electromagnetic force vector, and L f is the electromagnetic force coefficient matrix.
[0020] ; ; ; ; where diag(•) represents constructing a diagonal matrix.
[0021] The state equation and output equation of the active electromagnetic bearing rotor system are: ; where X represents the n-dimensional state vector, represents the first derivative of the n-dimensional state vector X with respect to time, A represents the n×n state matrix, B represents the n×l input matrix, U represents the l-dimensional control vector, Y represents the m-dimensional output vector, and C represents the m×n output matrix.
[0022] The active electromagnetic bearing rotor system becomes a decoupled controlled system after linear state feedback decoupling , and the decoupled controlled system is an integral type system, and its closed-loop transfer function is expressed as: ; where s represents the Laplace complex variable, I represents the input control current matrix, R represents the input decoupling matrix, F represents the feedback decoupling matrix, and h i (s) represents the i-th non-associated integral subsystem, i = 1, 2, 3, …, m, and m represents the number of integral subsystems.
[0023] The relationship between the reference input V and the output Y can be expressed as: ; where a i represents the i-th order decoupling constant, v i represents the i-th input, represents the i-th output after decoupling.
[0024] From the state equation of the system it can be seen that the i-th output equation is: ; where y i represents the i-th output, and c i represents the i-th row vector of the output matrix C.
[0025] Differentiate both sides of the relationship equation between the reference input V and the output Y with respect to time, and then substitute the state equation and output equation of the active electromagnetic bearing rotor system to obtain: ; where represents the first derivative of the i-th output.
[0026] If c i B≠0, then U is obviously included in the above formula. Let a i = 1 and no longer differentiate the above formula. Otherwise, continue to differentiate both sides of the above formula with respect to time and substitute the state equation and output equation of the active electromagnetic bearing rotor system to obtain: ; where represents the second derivative of the i-th output.
[0027] If c i AB≠0, then U is obviously included in the above formula. Let a i = 1 and no longer differentiate the above formula. Otherwise, repeat the above steps until is satisfied. At this time, the a i -th derivative of y with respect to time is: i ; ; So far, a total of m decoupling constants α1, α2, …, α i , …, α m , …, can be obtained, and the specific expressions are as follows: ; where k represents the k-th decoupling order constant.
[0028] Substitute the expression of the a i -th derivative of y with respect to time into the relational expression between the reference input V and the output Y, and the a i -th derivative of the output Y can be obtained as: i ; ; When W2 is a non-singular matrix, the above equation can be solved inversely to obtain the state feedback control law: .
[0029] It can be seen from this that the expressions of the decoupling matrix R and the decoupling matrix F of the decoupled controlled system are: ; ; ; where W1 represents the decoupled input matrix, W2 represents the decoupled output matrix, W2 -1 represents the inverse matrix of W2, T qb represents the lever arm matrix of the electromagnetic bearing-rotor system, K s represents the force-displacement stiffness coefficient matrix of the electromagnetic bearing rotor system, K i represents the force-current stiffness coefficient matrix of the electromagnetic bearing rotor system, M represents the mass matrix of the system, M -1 represents the inverse matrix of M, L f represents the lever arm matrix of the system, and G represents the gyro matrix.
[0030] S2: Construct the compensation current according to the unbalanced forces acting on the four degrees of freedom of the rotor, extract the unbalance coefficients in the compensation current, and construct an adaptive identification equation.
[0031] To cancel the influence caused by the unbalanced disturbance I d , an additional compensation current can be input into the output current V u of the controller to cancel the influence of I d . The compensation current varies with the output current V uEntering the magnetic bearing coil, the generated electromagnetic force can exactly offset the unbalance force of the rotor. In the ideal case, the unbalance force is completely compensated, and the rotor will rotate around the geometric axis. At this time, the vibration displacement of the rotor is 0, achieving effective control of the unbalance vibration displacement. The output expression of the system is: ; where, represents the acceleration of the centroid coordinate of the magnetic bearing, represents the compensation current.
[0032] The specific expression of the compensation current is: ; ; ; ; ; where I u represents the compensation current, I d represents the unbalance perturbation, represents the unbalance coefficient in the x-axis direction at the A end of the rotor, represents the unbalance coefficient in the y-axis direction at the A end of the rotor, represents the unbalance coefficient in the x-axis direction at the B end of the rotor, represents the unbalance coefficient in the y-axis direction at the B end of the rotor, and it is defined that and are the unbalance coefficients at the A end and B end of the rotor.
[0033] The unbalance coefficient at the A end of the rotor and the unbalance coefficient at the B end of the rotor are recorded in complex form: ; ; where, represents the complex form of the unbalance coefficient at the A end of the rotor, represents the complex form of the unbalance coefficient at the B end of the rotor, and j represents the imaginary unit.
[0034] The identified values of the unbalance coefficient at the A end of the rotor and the identified values of the unbalance coefficient at the B end of the rotor are: ; ; Among them, represents the identification value of the unbalance coefficient in the x-direction at the A-end of the rotor, represents the identification value of the unbalance coefficient in the y-direction at the A-end of the rotor, represents the identification value of the unbalance coefficient in the x-direction at the B-end of the rotor, represents the identification value of the unbalance coefficient in the y-direction at the B-end of the rotor.
[0035] The vibration displacement of the rotor at the A-end of the electromagnetic bearing and the vibration displacement of the rotor at the B-end of the electromagnetic bearing are denoted as: ; ; Among them, Z gA represents the vibration displacement of the rotor at the A-end of the electromagnetic bearing, x bA represents the displacement in the x-direction at the A-end of the rotor, y bA represents the displacement in the y-direction at the A-end of the rotor, Z gB represents the vibration displacement of the rotor at the B-end of the electromagnetic bearing, x bB represents the displacement in the x-direction at the B-end of the rotor, y bB represents the displacement in the y-direction at the B-end of the rotor.
[0036] To realize the identification of the rotor unbalance coefficient, an adaptive identification equation can be constructed: ; ; Observing the above formula, when the rotor unbalance force is completely cancelled out, the rotor rotates around its geometric axis, and the vibration displacement of the rotor tends to 0. At this time, the identification value of the unbalance coefficient will tend to a constant solution, that is converges to the true value of the unbalance coefficient , the identification value of the unbalance coefficient will tend to a constant solution, that is converges to the true value of the unbalance coefficient , where k and k1 are undetermined coefficients.
[0037] Then the adaptive identification equation for the unbalance coefficient at the A-end of the rotor is: ; The adaptive identification equation for the unbalance coefficient at the B-end of the rotor is: ; Among them, represents the second derivative of the identification value of the unbalance coefficient in the x-direction at the A-end of the rotor, represents the first derivative of the identification value of the unbalance coefficient in the x-direction at the A-end of the rotor, x bA represents the displacement in the x-direction at the A-end of the rotor, ybA represents the displacement of the A - end of the rotor in the y - direction. k1, k2, and k3 represent the undetermined parameters in the self - adaptive identification equation of the imbalance coefficient at the A - end of the rotor, and k4, k5, and k6 are the undetermined parameters in the self - adaptive identification equation of the imbalance coefficient at the B - end of the rotor. represents the second - order derivative of the identified value of the imbalance coefficient in the y - direction at the A - end of the rotor. represents the first - order derivative of the identified value of the imbalance coefficient in the y - direction at the A - end of the rotor. represents the second - order derivative of the identified value of the imbalance coefficient in the x - direction at the A - end of the rotor. represents the first - order derivative of the identified value of the imbalance coefficient in the x - direction at the B - end of the rotor, x bB represents the displacement of the B - end of the rotor in the x - direction, y bB represents the displacement of the B - end of the rotor in the y - direction. represents the second - order derivative of the identified value of the imbalance coefficient in the y - direction at the B - end of the rotor. represents the first - order derivative of the identified value of the imbalance coefficient in the y - direction at the B - end of the rotor.
[0038] S3: Construct a DDPG agent, using the rotor position error, the integral of the rotor position error, and the differential of the rotor position error as the observed variables of the system speed, and design a multi - objective reward function.
[0039] The DDPG agent includes an action network, an action target network, a critic network, a critic target network, and an experience pool. The training process of the DDPG agent is as follows: S301: Initialize the weights of the action network of the model for the imbalance compensation of the active magnetic bearing rotor system based on the DDPG algorithm and the weights of the critic network ; ; S302: Create an experience pool for storing the experience data generated during the interaction between the DDPG agent and the external environment. Initialize the capacity of the experience pool as D, the number of training sampling samples as n, and the maximum number of training episodes as Episode. S303: Define the state space of the active magnetic bearing rotor system as (s t , a t , r t , s t+1 ), where s t represents the state of the active magnetic bearing rotor system at time t, a t represents the action output by the DDPG agent at time t, r t represents the total reward value obtained by the DDPG agent at time t, and s t+1 represents the state of the active magnetic bearing rotor system at time t + 1; S304: Based on the state s of the active electromagnetic bearing rotor system at time t t and the exploration noise N t generate the action a output by the DDPG agent at time t through the action network t , and update the state s of the active electromagnetic bearing rotor system at time t + 1 t+1 and calculate the total reward value r obtained by the DDPG agent at time t t ; S305: After each exploration of the DDPG agent, the relevant data groups (s t , a t , r t , s t+1 ) are stored in the experience pool. If the number of samples in the experience pool is greater than the training sampling number, then during training, randomly select multiple data groups (s t , a t , r t , s t+1 ) from the experience replay pool for training to update the target parameters, otherwise continue to execute step S304.
[0040] The steps to update the target parameters are as follows: S3051: For the extracted mini-batch of samples, calculate the target Q value using the evaluation target network. The specific formula is: ; where γ is the discount factor used to correct the online network, a t+1 represents the action output by the DDPG agent at time t + 1, is the target state-action value for executing action a t+1 in state s t+1 ; S3052: Update the evaluation network parameters of the DDPG agent by minimizing the loss function L; the minimized loss function is defined as the mean squared error between the current Q value estimate and the calculated target Q value. The specific expression is: ; where, represents the expectation of the data group (s t , a t , r t , s t+1 ); S3053: Adopt a delayed update strategy for updating the action network parameters, and update the action network parameters through the gradient ascent method. The gradient formula is: ; where, Denote the gradient of the Q-value function with respect to action a t ; is the gradient of the policy function with respect to μ i ; denotes the expected value of the current action, denotes the policy at time t; S3054: Adopt a soft update strategy for the DDPG agent, and use the action target function and the evaluation target function to calculate the target value. The update methods of the action target network and the evaluation target network are as follows: ; wherein are the parameters of the action target network, are the parameters of the action network, are the parameters of the evaluation target network, are the parameters of the evaluation network, is the sliding average coefficient, and its value is less than 1; S306: Increase the number of time steps and perform the next round of training until the training ends.
[0041] The expression of the observed variable is: ; wherein, R obserbation denotes the observed variable, e error denotes the rotor position error, T denotes the total time, t denotes the current time, V rpm denotes the system rotation speed.
[0042] The expression of the multi-objective reward function is: ; ; ; ; wherein, r error denotes the accuracy reward, r response denotes the response time reward, r overshoot denotes the overshoot reward, a1 denotes the error weight coefficient, a2 denotes the response weight coefficient, a3 denotes the overshoot weight coefficient, r set denotes the value that the system is set to reach, r final denotes the actual value after the system stabilizes, T response denotes the time required for the system to stabilize, r max denotes the maximum response value during the operation of the system.
[0043] S4: Use the DDPG agent to optimize the undetermined parameters in the adaptive identification equation, dynamically adjust the compensation current in combination with the rotor speed, and cancel the unbalanced force disturbance.
[0044] The following is a verification of the effect of the control method described in the embodiments of the present invention through simulation experiments: The settings of the hyperparameters of the DDPG algorithm in the embodiments of the present invention are shown in Table 1.
[0045] Table 1 Hyperparameter Settings of the Algorithm
[0046]
[0047] After setting all the parameters, with the initial rotor speed being 0 and the acceleration a = 20 (rad / min) / s, the rotor system is offline trained. After the training is completed, the unbalance compensation of the rotor system can be achieved.
[0048] In the constant-speed model of the flywheel rotor, simulations are carried out under the condition that the rotor speed ω = 10000 rpm, and the unbalance compensation algorithm based on the DDPG agent is started at the 1st second. According to Figure 4 the speed of unbalance coefficient identification at constant speed, it can be known that the identification speed of this algorithm is 1.35 s faster than that of the traditional algorithm.
[0049] According to Figure 5 it can be known that at the constant speed ω = 10000 rpm, the maximum amplitude of the rotor drops from 2.5×10 -5 m to 2×10 - 17 m. Compared with the situation before compensation, the vibration suppression effect of the embodiments of the present invention reduces the amplitude of the rotor by 99%.
[0050] In the active magnetic bearing-rigid rotor acceleration motion model, the rotor accelerates from a stationary state to 8000 rpm with an acceleration a = 20 (rad / min) / s; according to Figure 4 it can be known that the maximum amplitude of the rotor drops from 2.4×10 -5 m to 1×10 -6 m, a decrease of 95%, which reflects the vibration suppression effect of the magnetic bearing compensation control algorithm under the rotor acceleration operation state.
[0051] From Figure 6 it can be seen that the identification speed of the unbalance coefficient in the embodiments of the present invention is much faster than that of the traditional algorithm. The identification speed of the unbalance coefficient in the embodiments of the present invention is only 0.15 s, which can quickly identify the specific unbalance coefficient, so as to quickly compensate for the unbalanced force.
[0052] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for controlling the imbalance of an electromagnetic bearing rotor system based on a DDPG algorithm, characterized in that: The steps include: S1: Based on the rotor dynamics, the four-degree-of-freedom AMB-rigid rotor dynamics equation is established, and the decoupling control of the active electromagnetic bearing rotor system is realized through linear state feedback decoupling; S2: Construct the compensation current according to the unbalanced force on the four degrees of freedom of the rotor, extract the unbalanced coefficient in the compensation current, and construct the adaptive identification equation; S3: Construct a DDPG agent, take the rotor position error, rotor position error integral and rotor position error differential as the system speed observation variables, and design a multi-objective reward function; S4: Use the DDPG agent to optimize the unknown parameters in the adaptive identification equation, and dynamically adjust the compensation current based on the rotor speed to offset the unbalanced force disturbance.
2. The unbalance control method of an electromagnetic bearing rotor system based on the DDPG algorithm according to claim 1 is characterized in that: The specific expression of the four-degree-of-freedom AMB-rigid rotor dynamics equation is: ; ; Where J represents the lateral moment of inertia, J z represents the moment of inertia of the rotor around the z-axis, Indicates the rotor rotation speed, L bA Indicates the distance from the plane where the electromagnetic bearing at the rotor end A is located to the plane of the rotor center of mass, f xa represents the electromagnetic force of the electromagnetic bearing A on the rotor in the x direction, L bB Indicates the distance from the plane where the electromagnetic bearing at the B end of the rotor is located to the plane of the rotor's center of mass, f xb represents the electromagnetic force of the electromagnetic bearing B on the rotor in the x direction, u z is the distance between the center of mass of the unbalanced mass and the center of mass of the rotor at O z The projection on z is the distance between the unbalanced mass point and the center of mass at O z The projection on the axis, m represents the rotor mass, represents the acceleration of the rotor in the x direction, represents the angular acceleration of the rotor around the x-axis, represents the angular velocity of the rotor around the x-axis, represents the angular acceleration of the rotor around the y-axis, represents the angular velocity of the rotor around the y-axis, f yb represents the electromagnetic force of the electromagnetic bearing B on the rotor in the y direction, f ya represents the electromagnetic force of the electromagnetic bearing A on the rotor in the y direction, represents the acceleration of the rotor in the y direction, represents the unbalanced mass, represents the rotor rotation speed, represents the rotor rotation angle, Represents ε z With O x The initial angle of the axis, represents the rotor rotation acceleration, Fu is the generalized unbalance vector of the electromagnetic bearing rotor system, Fx represents the unbalance force on the rotor in the x direction, and Fy represents the unbalance force on the rotor in the y direction. It represents the unbalanced force on the rotor when rotating around the x-axis. It represents the unbalanced force on the rotor when it rotates around the y-axis.
3. The unbalance control method of an electromagnetic bearing rotor system based on the DDPG algorithm according to claim 2 is characterized in that: The active electromagnetic bearing rotor system The state equation and output equation are: ; in, represents the first-order derivative of X with respect to time, X represents the n-dimensional state vector, A represents the n×n-order state matrix, B represents the n×l-order input matrix, U represents the l-dimensional control vector, Y represents the m-dimensional output vector, and C represents the m×n-order output matrix; The active electromagnetic bearing rotor system After linear state feedback decoupling, it becomes a decoupled controlled system , decouple the controlled system For an integral system, decouple the controlled system The expressions of the decoupling matrix R and the decoupling matrix F are: ; ; ; Among them, W1 represents the input matrix after decoupling, W2 represents the output matrix after decoupling, and W2 -1 represents the inverse matrix of the decoupled output matrix W2, T qb represents the arm matrix of the electromagnetic bearing-rotor system, K s The force-displacement stiffness coefficient matrix of the electromagnetic bearing rotor system, K i represents the force-current stiffness coefficient matrix of the electromagnetic bearing rotor system, M represents the generalized mass matrix of the rotor system, and M -1 represents the inverse matrix of the generalized mass matrix M of the rotor system, L f represents the arm matrix, and G represents the gyro matrix.
4. The unbalance control method of an electromagnetic bearing rotor system based on the DDPG algorithm according to claim 3 is characterized in that: The specific expression of the compensation current is: ; ; ; ; ; Among them, I u Represents the compensation current, I d represents an unbalanced disturbance, It represents the unbalance coefficient of the rotor end A in the x-axis direction. It represents the unbalance coefficient of the rotor end A in the y-axis direction. It represents the unbalance coefficient of the rotor B end in the x-axis direction, It represents the unbalance coefficient of the rotor B end in the y-axis direction and defines and is the unbalance coefficient of the rotor ends A and B; The unbalance coefficient of the rotor end A is and the rotor B end unbalance coefficient In plural form: ; ; in, Indicates the unbalance coefficient of the rotor A end plural form of Indicates the unbalance coefficient of the rotor B end The plural form of , j represents the imaginary unit; Note the unbalance coefficient of the rotor A end Identification value and the rotor B end unbalance coefficient Identification value for: ; ; in, Indicates the unbalance coefficient identification value of the rotor end A in the x direction, Indicates the unbalance coefficient identification value of the rotor end A in the y direction, Indicates the unbalance coefficient identification value of the rotor B end in the x direction, Indicates the unbalance coefficient identification value of the rotor B end in the y direction; Then the adaptive identification equation of the unbalance coefficient at the rotor A end is: ; The adaptive identification equation of the unbalance coefficient at the B end of the rotor is: ; in, Represents the second-order derivative of the unbalance coefficient identification value in the x direction at the rotor end A, It represents the first-order derivative of the unbalance coefficient identification value in the x direction at the rotor end A, x bA represents the displacement of the rotor end A in the x direction, y bA represents the displacement of rotor end A in the y direction, k1, k2 and k3 represent the undetermined parameters in the adaptive identification equation of the unbalance coefficient of rotor end A, k4, k5 and k6 represent the undetermined parameters in the adaptive identification equation of the unbalance coefficient of rotor end B, The second derivative of the unbalance coefficient identification value in the y direction at the rotor end A is: Represents the first derivative of the unbalance coefficient identification value in the y direction at the rotor end A; Represents the second-order derivative of the unbalance coefficient identification value in the x direction at the rotor end A, It represents the first-order derivative of the unbalance coefficient identification value in the x direction at the B end of the rotor, x bB represents the displacement of the rotor end B in the x direction, y bB represents the displacement of the rotor end B in the y direction, Represents the second-order derivative of the unbalance coefficient identification value in the y direction at the B end of the rotor, Represents the first-order derivative of the unbalance coefficient identification value in the y direction at the B end of the rotor.
5. The unbalance control method of an electromagnetic bearing rotor system based on a DDPG algorithm according to claim 4, characterized in that: The DDPG agent includes an action network, an action target network, an evaluation network, an evaluation target network and an experience pool. The training process of the DDPG agent is as follows: S301: Initialize the action network of the model of unbalance compensation of active electromagnetic bearing rotor system based on DDPG algorithm Weight and evaluation network Weight ; S302: Create an experience pool to store the experience data generated during the interaction between the DDPG agent and the external environment. Initialize the experience pool capacity to D, the number of training samples to n, and the maximum number of training rounds to Episode; S303: Define the spatial state of the active electromagnetic bearing rotor system as (s t , a t , r t ,s t+1 ), where s t represents the state of the active electromagnetic bearing rotor system at time t, a t represents the action output by the DDPG agent at time t, r t represents the total reward value obtained by the DDPG agent at time t, s t+1 represents the state of the active electromagnetic bearing rotor system at time t+1; S304: Based on the state s of the active electromagnetic bearing rotor system at time t t and exploration noise N t Generate the action a output by the DDPG agent at time t through the action network t , and update the state s of the active electromagnetic bearing rotor system at time t+1 t+1 And calculate the total reward value r obtained by the DDPG agent at time t t ; S305: The relevant data group (s) obtained after each exploration by the DDPG agent t , a t , r t ,s t+1 ) are stored in the experience pool. If the number of samples in the experience pool is greater than the number of training samples, multiple data sets (s t , a t , r t ,s t+1 ) to perform training and update the target parameters, otherwise continue to execute step S304; S306: Increase the number of time steps and perform the next round of training until the training is completed.
6. The unbalance control method of an electromagnetic bearing rotor system based on the DDPG algorithm according to claim 5, characterized in that: The expression of the observed variable in S3 is: ; Among them, R obserbation represents the observed variable, e error represents the rotor position error, T represents the total time, t represents the current moment, V rpm Indicates the system speed; The expression of the multi-objective reward function is: ; ; ; ; Among them, r error represents the accuracy reward, r response represents the response time reward, r overshoot represents the overshoot reward, a1 represents the error weight coefficient, a2 represents the response weight coefficient, a3 represents the overshoot weight coefficient, r set Indicates the value that the system setting needs to reach, r final Indicates the actual value after the system stabilizes, T response Indicates the time it takes for the system to stabilize, r max Indicates the maximum response value of the system during operation.
7. The unbalance control method of an electromagnetic bearing rotor system based on the DDPG algorithm according to claim 6 is characterized in that: The steps of updating the target parameters in S305 are as follows: S3051: For the extracted small batch samples, the target Q value is calculated using the evaluation target network. The specific formula is: ; Among them, γ is the discount factor used to correct the online network, a t+1 represents the action output by the DDPG agent at time t+1, For state s t+1 Next, perform action a t+1 The target state of the action is the value of the action; S3052: Update the evaluation network parameters of the DDPG agent by minimizing the loss function L; the minimization loss function is defined as the mean square error between the current Q value estimate and the calculated target Q value, and the specific expression is: ; in, Represents the data set (s t , a t , r t ,s t+1 ) expectations; S3053: A delayed update strategy is adopted for updating the action network parameters. The action network parameters are updated by the gradient ascent method. The gradient formula for: ; in, represents the Q value function for action a t The gradient of is the policy function with respect to μ i The gradient of Indicates the expected value of the current action. represents the strategy at time t; S3054: A soft update strategy is adopted for the DDPG agent, and the action objective function and the evaluation objective function are used to calculate the target value. The update method of the action target network and the evaluation target network is as follows: ; In the formula is the action target network parameter, are the action network parameters, To evaluate the target network parameters, To evaluate the network parameters, is the sliding average coefficient, and its value is less than 1.
Citation Information
Patent Citations
Unbalanced compensation algorithm of polygonal iterative search with variable step size for rotor imbalance coefficient
CN107133387A
Unbalance compensation method for rotor unbalance coefficient variable-angle iterative search
CN111177943A
Unbalanced intelligent fault quantitative diagnosis method based on reward optimization deep reinforcement learning
CN116561517A
Power load model parameter identification method based on improved deep reinforcement learning
CN119129367A
Adaptive compensation of sensor run-out and mass unbalance in magnetic bearing systems without changing rotor speed
WO2002016792A1
Cited By
Electromagnetic bearing-flexible rotor system unbalance control method
CN120469198A