Unbalance Control Method of Electromagnetic Bearing Rotor System Based on DDPG Algorithm

Through the imbalance control method based on the DDPG algorithm, the imbalance force of the electromagnetic bearing rotor system is quickly identified and compensated, which solves the problems of complex calculations and poor compensation effects in the prior art, and realizes efficient stable control of the rotor system.

CN120217899BActive Publication Date: 2025-08-29NANCHANG HANGKONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510655088.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-29
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

In the vibration suppression caused by unbalanced forces, the existing active electromagnetic bearing rotor system has problems such as complex calculations and poor compensation effect. Especially when the reaction of unbalanced forces at high speeds, the rotor suspension stability is reduced, and even instability is caused.

Method used

The imbalance control method based on the DDPG algorithm is adopted, and the compensation current expression is constructed, the imbalance coefficient is extracted, and the observation variables and multi-objective reward functions are designed in combination with the deep reinforcement learning algorithm, and the adaptive identification equation is optimized to achieve rapid identification and compensation of imbalance forces.

Benefits of technology

It improves the identification speed and compensation efficiency of the imbalance coefficient, realizes stable control of the rotor system within the full speed range, has high compensation accuracy, simple calculation, and less resource utilization, and is suitable for low-speed scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217899B_ABST
    Figure CN120217899B_ABST
Patent Text Reader

Abstract

The present application relates to an imbalance control method for an electromagnetic bearing rotor system based on the DDPG algorithm, comprising the following steps: establishing a rotor dynamics equation based on rotor dynamics and performing decoupling control; constructing a compensation current, extracting the imbalance coefficient from the compensation current and constructing an adaptive identification equation; constructing a DDPG agent, integrating position error, system response, and control output, and designing a multi-objective reward function through a normalized weighted optimization strategy; using the DDPG agent to optimize the imbalance coefficient identification parameters, dynamically adjusting the compensation current in combination with the rotor speed to offset the unbalanced force disturbance. The present invention identifies the imbalance coefficient reflected at both ends of the rotor by the unbalanced force exerted on the rotor in combination with the deep reinforcement learning algorithm DDPG, thereby improving the identification of the imbalance coefficient and the efficiency of imbalance compensation, and achieving rapid, stable, and efficient compensation of the rotor system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of active electromagnetic bearing rotor system control, and in particular to an imbalance control method for an electromagnetic bearing rotor system based on a DDPG algorithm. Background Art

[0002] Active electromagnetic bearings (AMBs) use controllable electromagnetic forces to suspend the rotor. Due to their long lifespan and suitability for high speeds, they are widely used in high-speed motors, centrifugal compressors, artificial heart pumps, and other fields. In the electromagnetic bearing-rotor system, issues such as material quality and assembly process can cause the rotor's center of mass to shift from its geometric center. During rotor rotation, this center of mass shift causes an unbalanced force, which reacts on the rotor, causing it to vibrate unbalancedly. This unbalanced force increases with increasing speed, exacerbating the unbalanced vibration and reducing the stability of the rotor suspension, potentially leading to instability. Therefore, it is necessary to suppress the rotor's unbalanced vibration.

[0003] At present, the control methods for imbalance compensation of active electromagnetic bearing rotor systems are mainly divided into two categories: one is the imbalance compensation method based on the identification of the rotor unbalance mass position, and the imbalance vibration compensation signal is calculated in real time according to the rotor unbalance mass position and the rotor speed. However, when the frequency and amplitude of the unbalance disturbance increase, the convergence speed of the search algorithm will slow down, resulting in a worse compensation effect; the other is to use adaptive control methods, such as neural networks, LMS-based adaptive filters, etc., but the search algorithm is relatively complex, and the performance requirements of the controller are high, and the actual compensation effect is poor.

[0004] This invention constructs a compensation current expression based on the external disturbances experienced by the rotor, extracts the imbalance coefficient, and uses the DDPG algorithm to design appropriate observation variables. This algorithm, combined with multiple indicators as a reward function, accurately identifies the imbalance coefficient and thus derives the optimal compensation current. Compared to existing inventions, the proposed compensation algorithm boasts simple calculations, excellent compensation across all speeds, and high compensation accuracy. Summary of the Invention

[0005] The purpose of the present invention is to provide an imbalance control method for an electromagnetic bearing rotor system based on the DDPG algorithm. The imbalance coefficient reflected at both ends of the rotor by the unbalanced force exerted on the rotor is identified in combination with the deep reinforcement learning algorithm DDPG, thereby reducing the amount of calculation, improving the identification speed of the imbalance coefficient and the compensation efficiency of the imbalance compensation algorithm within the full speed range, and realizing rapid, stable and efficient compensation of the rotor system.

[0006] The technical solution adopted by the present invention is: a method for controlling the imbalance of an electromagnetic bearing rotor system based on the DDPG algorithm, comprising the following steps:

[0007] S1: Based on rotor dynamics, the four-degree-of-freedom AMB-rigid rotor dynamics equation is established, and decoupling control of the active electromagnetic bearing rotor system is achieved through linear state feedback decoupling;

[0008] S2: Construct a compensation current based on the unbalanced forces acting on the four degrees of freedom of the rotor, extract the unbalance coefficient in the compensation current, and construct an adaptive identification equation;

[0009] S3: Construct a DDPG agent, use the rotor position error, rotor position error integral, and rotor position error differential as the system speed observation variables, and design a multi-objective reward function;

[0010] S4: Use the DDPG agent to optimize the undetermined parameters in the adaptive identification equation and dynamically adjust the compensation current based on the rotor speed to offset the unbalanced force disturbance.

[0011] Furthermore, the specific expression of the four-degree-of-freedom AMB-rigid rotor dynamics equation is:

[0012] ;

[0013] ;

[0014] Where, J represents the lateral moment of inertia, J z represents the moment of inertia of the rotor around the z-axis, Indicates the rotor rotation speed, L bA Indicates the distance from the plane where the electromagnetic bearing at end A of the rotor is located to the plane of the rotor's center of mass, f xa represents the electromagnetic force of the electromagnetic bearing A on the rotor in the x direction, L bB Indicates the distance from the plane where the electromagnetic bearing at the B end of the rotor is located to the plane of the rotor center of mass, f xb represents the electromagnetic force of the electromagnetic bearing B on the rotor in the x direction, u z The distance between the center of mass of the unbalanced mass and the center of mass of the rotor is z The projection on z The distance between the unbalanced mass point and the center of mass is z The projection on the axis, m represents the rotor mass, represents the acceleration of the rotor in the x direction, represents the angular acceleration of the rotor around the x-axis, represents the angular velocity of the rotor around the x-axis, represents the angular acceleration of the rotor around the y-axis, represents the angular velocity of the rotor around the y-axis, f yb represents the electromagnetic force of the electromagnetic bearing B on the rotor in the y direction, f ya represents the electromagnetic force of the electromagnetic bearing A on the rotor in the y direction, represents the acceleration of the rotor in the y direction, represents the unbalanced mass, represents the rotor rotation speed, represents the rotor rotation angle, Represents ε z With O x The initial angle of the axis, represents the rotor rotation acceleration, Fu is the generalized unbalance vector of the electromagnetic bearing rotor system, Fx represents the unbalance force on the rotor in the x direction, and Fy represents the unbalance force on the rotor in the y direction. It represents the unbalanced force on the rotor when it rotates around the x-axis. It represents the unbalanced force on the rotor when it rotates around the y-axis.

[0015] Furthermore, the active electromagnetic bearing rotor system The state equation and output equation are:

[0016] ;

[0017] in, represents the first-order derivative of X with respect to time, X represents the n-dimensional state vector, A represents the n×n-order state matrix, B represents the n×l-order input matrix, U represents the l-dimensional control vector, Y represents the m-dimensional output vector, and C represents the m×n-order output matrix;

[0018] The active electromagnetic bearing rotor system After linear state feedback decoupling, it becomes a decoupled controlled system , decoupled controlled systems For an integral system, decouple the controlled system The expressions of the decoupling matrix R and the decoupling matrix F are:

[0019] ;

[0020] ;

[0021] ;

[0022] Among them, W1 represents the input matrix after decoupling, W2 represents the output matrix after decoupling, and W2 -1 represents the inverse matrix of the decoupled output matrix W2, T qb represents the arm matrix of the electromagnetic bearing-rotor system, K s The force-displacement stiffness coefficient matrix of the electromagnetic bearing rotor system, K i represents the force-current stiffness coefficient matrix of the electromagnetic bearing rotor system, M represents the generalized mass matrix of the rotor system, M -1represents the inverse matrix of the generalized mass matrix M of the rotor system, L f represents the arm matrix, and G represents the gyro matrix.

[0023] Furthermore, the specific expression of the compensation current is:

[0024] ;

[0025] ;

[0026] ;

[0027] ;

[0028] ;

[0029] Among them, I u Indicates the compensation current, I d represents the unbalanced disturbance, Indicates the unbalance coefficient of the rotor A end in the x-axis direction, Indicates the unbalance coefficient of the rotor A end in the y-axis direction, Indicates the unbalance coefficient of the rotor B end in the x-axis direction, It represents the unbalance coefficient in the y-axis direction of the rotor B end and defines and is the unbalance coefficient between the rotor ends A and B;

[0030] The unbalance coefficient of the rotor A end and the rotor B-end unbalance coefficient In plural form:

[0031] ;

[0032] ;

[0033] in, Indicates the unbalance coefficient of the rotor A end plural form of Indicates the unbalance coefficient of the rotor B end The plural form of , j represents the imaginary unit;

[0034] Note the unbalance coefficient of rotor A end Identification value of and the rotor B-end unbalance coefficient Identification value of for:

[0035] ;

[0036] ;

[0037] in, Indicates the unbalance coefficient identification value of the rotor A end in the x direction, Indicates the unbalance coefficient identification value of the rotor A end in the y direction, Indicates the unbalance coefficient identification value of the rotor B end in the x direction, Indicates the unbalance coefficient identification value in the y direction of the rotor end B;

[0038] Then the adaptive identification equation of the unbalance coefficient at end A of the rotor is:

[0039] ;

[0040] The adaptive identification equation of the unbalance coefficient at the B end of the rotor is:

[0041] ;

[0042] in, Represents the second-order derivative of the unbalance coefficient identification value in the x direction at the rotor end A, Represents the first-order derivative of the unbalance coefficient identification value in the x direction at the rotor end A, x bA represents the displacement of the rotor A end in the x direction, y bA represents the displacement of rotor A end in the y direction, k1, k2 and k3 represent the undetermined parameters in the adaptive identification equation of the unbalance coefficient of rotor A end, k4, k5 and k6 represent the undetermined parameters in the adaptive identification equation of the unbalance coefficient of rotor B end, The second derivative of the unbalance coefficient identification value in the y direction at the rotor end A is represented by: Represents the first derivative of the unbalance coefficient identification value in the y direction at the rotor end A; Represents the second-order derivative of the unbalance coefficient identification value in the x direction at the rotor end A, Represents the first-order derivative of the unbalance coefficient identification value in the x direction at the B end of the rotor, x bB represents the displacement in the x direction of the rotor B end, y bB represents the displacement of the rotor end B in the y direction, Represents the second-order derivative of the unbalance coefficient identification value in the y direction at the B end of the rotor, Represents the first-order derivative of the unbalance coefficient identification value in the y direction at the B end of the rotor.

[0043] Furthermore, the DDPG agent includes an action network, an action target network, an evaluation network, an evaluation target network and an experience pool. The training process of the DDPG agent is as follows:

[0044] S301: Initialize the action network of the model for imbalance compensation of active electromagnetic bearing rotor system based on DDPG algorithm Weight and evaluation network Weight ;

[0045] S302: Create an experience pool to store the experience data generated during the interaction between the DDPG agent and the external environment. Initialize the experience pool capacity to D, the number of training samples to n, and the maximum number of training rounds to Episodes;

[0046] S303: Define the spatial state of the active electromagnetic bearing rotor system as (s t , a t , r t , s t+1 ), where s t represents the state of the active electromagnetic bearing rotor system at time t, a t represents the action output by the DDPG agent at time t, r t represents the total reward value obtained by the DDPG agent at time t, s t+1 represents the state of the active electromagnetic bearing rotor system at time t+1;

[0047] S304: Based on the state s of the active electromagnetic bearing rotor system at time t t and exploration noise N t Generate the action a output by the DDPG agent at time t through the action network t , and update the state s of the active electromagnetic bearing rotor system at time t+1 t+1 And calculate the total reward value r obtained by the DDPG agent at time t t ;

[0048] S305: The relevant data group (s) obtained after each exploration by the DDPG agent t , a t , r t , s t+1 ) are stored in the experience pool. If the number of samples in the experience pool is greater than the number of training samples, multiple data sets (s t , a t , r t , s t+1 ) to conduct training and update the target parameters, otherwise continue to execute step S304;

[0049] S306: Increase the number of time steps and perform the next round of training until the training is completed.

[0050] Furthermore, the expression of the observed variable in S3 is:

[0051] ;

[0052] Among them, R obserbation represents the observed variable, e error represents the rotor position error, T represents the total time, t represents the current moment, V rpm Indicates the system speed;

[0053] The expression of the multi-objective reward function is:

[0054] ;

[0055] ;

[0056] ;

[0057] ;

[0058] Among them, r error represents the accuracy reward, r response represents the response time reward, r overshoot represents the overshoot reward, a1 represents the error weight coefficient, a2 represents the response weight coefficient, a3 represents the overshoot weight coefficient, r set Indicates the value that the system setting needs to reach, r final Indicates the actual value after the system stabilizes, T response Indicates the time it takes for the system to stabilize, r max Indicates the maximum response value of the system during operation.

[0059] Furthermore, the step of updating the target parameters in S305 is as follows:

[0060] S3051: For the extracted small batch samples, the target Q value is calculated using the evaluation target network. The specific formula is:

[0061] ;

[0062] Among them, γ is the discount factor used to correct the online network, a t+1 represents the action output by the DDPG agent at time t+1, For state s t+1 Next, perform action a t+1 The target state - action value;

[0063] S3052: Update the evaluation network parameters of the DDPG agent by minimizing the loss function L; the minimization loss function is defined as the mean square error between the current Q value estimate and the calculated target Q value, specifically expressed as:

[0064] ;

[0065] in, Indicates the data set (s t , a t , r t , s t+1 ) expectations;

[0066] S3053: A delayed update strategy is adopted for the update of action network parameters. The action network parameters are updated by the gradient ascent method. The gradient formula for:

[0067] ;

[0068] in, Represents the Q-value function for action a t The gradient, is the policy function with respect to μ i The gradient, Indicates the expected value of the current action. represents the strategy at time t;

[0069] S3054: A soft update strategy is adopted for the DDPG agent. The action objective function and the evaluation objective function are used to calculate the target value. The action objective network and the evaluation objective network are updated as follows:

[0070] ;

[0071] In the formula are the action target network parameters, are the action network parameters, To evaluate the target network parameters, To evaluate the network parameters, is the sliding average coefficient, and its value is less than 1.

[0072] The beneficial technical effects of the present invention are as follows: by combining the DDPG algorithm with an imbalance coefficient identification compensation algorithm, the problems of low imbalance coefficient identification efficiency and complex calculations are solved, achieving the effects of simple calculation, fast compensation speed, and high compensation accuracy; the present invention combines a deep reinforcement learning algorithm with an imbalance compensation algorithm, solving the problems of low traditional imbalance coefficient identification efficiency and complex parameter selection, and further improving the algorithm's compensation efficiency while achieving stable control of rotor system imbalance compensation. Compared with the existing technology, the compensation algorithm proposed in the present invention has the characteristics of simple calculation, small computing resource usage, high accuracy, and high imbalance coefficient identification efficiency. It can be adapted to low and high speed scenarios and scenarios with high rotor operation accuracy, broadening the algorithm's applicable scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0074] Figure 1 This is a control principle diagram of an embodiment of the present invention;

[0075] Figure 2 This is a transfer function block diagram of the linear state feedback decoupling method in an embodiment of the present invention;

[0076] Figure 3 This is a diagram of the DDPG agent training evaluation reward function in an embodiment of the present invention;

[0077] Figure 4 This is a comparison diagram of the maximum rotor amplitude before and after compensation under constant acceleration according to an embodiment of the present invention;

[0078] Figure 5 The figures are comparative diagrams of the rotor A-end axial displacement before and after compensation at a constant speed according to an embodiment of the present invention; wherein (a) is a schematic diagram of the rotor A-end axial displacement before compensation, and (b) is a schematic diagram of the rotor A-end axial displacement after compensation;

[0079] Figure 6 This is an imbalance coefficient identification diagram of a traditional algorithm and an embodiment of the present invention;

[0080] Figure 7 The following are the axis trajectory diagrams of rotor end A and rotor B at a constant speed of 10000 rpm, where (a) is the axis trajectory diagram of rotor end A and (b) is the axis trajectory diagram of rotor B. DETAILED DESCRIPTION

[0081] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0082] like Figure 1 and Figure 2 As shown, an embodiment of the present invention proposes an imbalance control method for an electromagnetic bearing rotor system based on a DDPG algorithm, comprising the following steps:

[0083] S1: Based on the rotor dynamics, the four-degree-of-freedom AMB-rigid rotor dynamics equation is established, and the decoupling control of the active electromagnetic bearing rotor system is achieved through linear state feedback decoupling.

[0084] The specific expression of the four-degree-of-freedom AMB-rigid rotor dynamics equation is:

[0085] ;

[0086] ;

[0087] Where, J represents the lateral moment of inertia, J z represents the moment of inertia of the rotor around the z-axis, Indicates the rotor rotation speed, L bA Indicates the distance from the plane where the electromagnetic bearing at end A of the rotor is located to the plane of the rotor's center of mass, f xa represents the electromagnetic force of the electromagnetic bearing A on the rotor in the x direction, L bB Indicates the distance from the plane where the electromagnetic bearing at the B end of the rotor is located to the plane of the rotor center of mass, f xb represents the electromagnetic force of the electromagnetic bearing B on the rotor in the x direction, u z The distance between the center of mass of the unbalanced mass and the center of mass of the rotor is z The projection on z The distance between the unbalanced mass point and the center of mass is z The projection on the axis, m represents the rotor mass, represents the acceleration of the rotor in the x direction, represents the angular acceleration of the rotor around the x-axis, represents the angular velocity of the rotor around the x-axis, represents the angular acceleration of the rotor around the y-axis, represents the angular velocity of the rotor around the y-axis, f yb represents the electromagnetic force of the electromagnetic bearing B on the rotor in the y direction, f ya represents the electromagnetic force of the electromagnetic bearing A on the rotor in the y direction, represents the acceleration of the rotor in the y direction, represents the unbalanced mass, represents the rotor rotation speed, represents the rotor rotation angle, Represents ε z With O x The initial angle of the axis, represents the rotor rotation acceleration, Fu is the generalized unbalance vector of the electromagnetic bearing rotor system, Fx represents the unbalance force on the rotor in the x direction, and Fy represents the unbalance force on the rotor in the y direction. It represents the unbalanced force on the rotor when it rotates around the x-axis. It represents the unbalanced force on the rotor when it rotates around the y-axis.

[0088] Rewriting the above formula into matrix form yields:

[0089] ;

[0090] Where M represents the generalized mass matrix of the rotor system; represents the 4-DOF displacement vector of the rotor mass center, where represents the angle of rotation around the y-axis, x represents the displacement in the x-direction, represents the angle of rotation around the x-axis, y represents the displacement in the y-direction; G is the gyroscope matrix; F xy is the electromagnetic force vector, L f is the electromagnetic force coefficient matrix.

[0091] ;

[0092] ;

[0093] ;

[0094] ;

[0095] Where diag(•) indicates constructing a diagonal matrix.

[0096] The active electromagnetic bearing rotor system The state equation and output equation are:

[0097] ;

[0098] Where X represents the n-dimensional state vector, represents the first-order derivative of the n-dimensional state vector X with respect to time, A represents the n×n state matrix, B represents the n×l input matrix, U represents the l-dimensional control vector, Y represents the m-dimensional output vector, and C represents the m×n output matrix.

[0099] The active electromagnetic bearing rotor system After linear state feedback decoupling, it becomes a decoupled controlled system , decoupled controlled systems It is an integral system, and its closed-loop transfer function Expressed as:

[0100] ;

[0101] Where s represents the Laplace complex variable, I represents the input control current matrix, R represents the input decoupling matrix, F represents the feedback decoupling matrix, and h i (s) represents the i-th integral subsystem without any association, i=1, 2, 3,…, m, m represents the number of integral subsystems.

[0102] The relationship between the reference input V and the output Y can be expressed as:

[0103] ;

[0104] Among them, a i represents the i-th order decoupling constant, v i represents the i-th input, represents the i-th output after decoupling.

[0105] By system From the state equation, we can see that the i-th output equation is:

[0106] ;

[0107] Among them, y i represents the i-th output, c i Represents the i-th row vector of the output matrix C.

[0108] Refer to the relationship between input V and output Y and take the derivative of both sides with respect to time, then the active electromagnetic bearing rotor system Substituting the state equation and output equation into the equation, we can get:

[0109] ;

[0110] in, Represents the first derivative of the i-th output.

[0111] If c i B≠0, then the above formula obviously contains U, let a i =1, and no longer derive the above formula. Otherwise, continue to derive both sides of the above formula with respect to time and convert the active electromagnetic bearing rotor system Substituting the state equation and output equation into the equation, we can get:

[0112] ;

[0113] in, Represents the second derivative of the i-th output.

[0114] If c i AB≠0, then the above formula obviously contains U, let a i =1, and no longer derive the above formula, otherwise, repeat the above steps until it satisfies At this time, y i a i The time derivative of order is:

[0115] ;

[0116] So far, we can get the m-order decoupling constants α1, α2, ..., α i ,…,α m , the specific expression is:

[0117] ;

[0118] Where k represents the kth decoupling order constant.

[0119] y i a i Substituting the expression of the derivative of the order with respect to time into the relationship expression between the reference input V and the output Y, we can get the output Y a i The derivative is:

[0120] ;

[0121] When W2 is a non-singular matrix, the above equation can be inversely solved to obtain the state feedback control law:

[0122] .

[0123] It can be seen that the decoupled controlled system The expressions of the decoupling matrix R and the decoupling matrix F are:

[0124] ;

[0125] ;

[0126] ;

[0127] Among them, W1 represents the input matrix after decoupling, W2 represents the output matrix after decoupling, and W2 -1 represents the inverse matrix of W2, T qb represents the arm matrix of the electromagnetic bearing-rotor system, K s The force-displacement stiffness coefficient matrix of the electromagnetic bearing rotor system, K i represents the force-current stiffness coefficient matrix of the electromagnetic bearing rotor system, M represents the mass matrix of the system, and M -1 represents the inverse matrix of M, L f represents the system's force arm matrix, and G represents the gyro matrix.

[0128] S2: Construct a compensation current based on the unbalanced forces acting on the four degrees of freedom of the rotor, extract the unbalanced coefficient in the compensation current, and construct an adaptive identification equation.

[0129] In order to offset the unbalanced disturbance I d The effect can be on the controller output current V u Input additional compensation current to offset Id The compensation current increases with the controller output current V u Entering the magnetic bearing coil, the electromagnetic force generated can just offset the unbalanced force of the rotor. Ideally, the unbalanced force is completely compensated and the rotor will rotate around the geometric axis. At this time, the vibration displacement of the rotor is 0, achieving effective control of the unbalanced vibration displacement. The output expression of the system is:

[0130] ;

[0131] in, represents the acceleration of the center of mass coordinate of the electromagnetic bearing, Indicates the compensation current.

[0132] The specific expression of the compensation current is:

[0133] ;

[0134] ;

[0135] ;

[0136] ;

[0137] ;

[0138] Among them, I u Indicates the compensation current, I d represents the unbalanced disturbance, Indicates the unbalance coefficient of the rotor A end in the x-axis direction, Indicates the unbalance coefficient of the rotor A end in the y-axis direction, Indicates the unbalance coefficient of the rotor B end in the x-axis direction, It represents the unbalance coefficient in the y-axis direction of the rotor B end and defines and is the unbalance coefficient between the A and B ends of the rotor.

[0139] The unbalance coefficient of the rotor A end and the rotor B-end unbalance coefficient In plural form:

[0140] ;

[0141] ;

[0142] in, Indicates the unbalance coefficient of the rotor A end plural form of Indicates the unbalance coefficient of the rotor B end The plural form of , where j represents the imaginary unit.

[0143] Note the unbalance coefficient of rotor A end Identification value of and the rotor B-end unbalance coefficient Identification value of for:

[0144] ;

[0145] ;

[0146] in, Indicates the unbalance coefficient identification value of the rotor A end in the x direction, Indicates the unbalance coefficient identification value of the rotor A end in the y direction, Indicates the unbalance coefficient identification value of the rotor B end in the x direction, Indicates the unbalance coefficient identification value in the y direction at the B end of the rotor.

[0147] The rotor vibration displacement at the A end of the electromagnetic bearing and the rotor vibration displacement at the B end of the electromagnetic bearing are recorded as:

[0148] ;

[0149] ;

[0150] Among them, Z gA represents the rotor vibration displacement at end A of the electromagnetic bearing, x bA represents the displacement of the rotor A end in the x direction, y bA represents the displacement of the rotor end A in the y direction, Z gB represents the rotor vibration displacement at the B end of the electromagnetic bearing, x bB represents the displacement in the x direction of the rotor B end, y bB Represents the displacement of the rotor end B in the y direction.

[0151] In order to identify the rotor unbalance coefficient, an adaptive identification equation can be constructed:

[0152] ;

[0153] ;

[0154] Observe the above formula, when the rotor unbalance force is completely offset, the rotor rotates around its geometric axis and the vibration displacement of the rotor tends to 0. The identification value of the unbalance coefficient at this time is will tend to a constant solution, i.e. Converges to the true value of the imbalance coefficient , the identification value of the unbalance coefficient will tend to a constant solution, i.e. Converges to the true value of the imbalance coefficient , where k and k1 are unknown coefficients.

[0155] Then the adaptive identification equation of the unbalance coefficient at end A of the rotor is:

[0156] ;

[0157] The adaptive identification equation of the unbalance coefficient at the B end of the rotor is:

[0158] ;

[0159] in, Represents the second-order derivative of the unbalance coefficient identification value in the x direction at the rotor end A, Represents the first-order derivative of the unbalance coefficient identification value in the x direction at the rotor end A, x bA represents the displacement of the rotor A end in the x direction, y bA represents the displacement of rotor A end in the y direction, k1, k2 and k3 represent the undetermined parameters in the adaptive identification equation of the unbalance coefficient of rotor A end, k4, k5 and k6 represent the undetermined parameters in the adaptive identification equation of the unbalance coefficient of rotor B end, The second derivative of the unbalance coefficient identification value in the y direction at the rotor end A is represented by: Represents the first derivative of the unbalance coefficient identification value in the y direction at the rotor end A; Represents the second-order derivative of the unbalance coefficient identification value in the x direction at the rotor end A, Represents the first-order derivative of the unbalance coefficient identification value in the x direction at the B end of the rotor, x bB represents the displacement in the x direction of the rotor B end, y bB represents the displacement of the rotor end B in the y direction, Represents the second-order derivative of the unbalance coefficient identification value in the y direction at the B end of the rotor, Represents the first-order derivative of the unbalance coefficient identification value in the y direction at the B end of the rotor.

[0160] S3: Construct a DDPG agent, use the rotor position error, rotor position error integral, and rotor position error differential as the system speed observation variables, and design a multi-objective reward function.

[0161] The DDPG agent includes an action network, an action target network, an evaluation network, an evaluation target network and an experience pool. The training process of the DDPG agent is as follows:

[0162] S301: Initialize the action network of the model for imbalance compensation of active electromagnetic bearing rotor system based on DDPG algorithm Weight and evaluation network Weight ;

[0163] S302: Create an experience pool to store the experience data generated during the interaction between the DDPG agent and the external environment. Initialize the experience pool capacity to D, the number of training samples to n, and the maximum number of training rounds to Episodes;

[0164] S303: Define the spatial state of the active electromagnetic bearing rotor system as (s t , a t , r t , s t+1 ), where s t represents the state of the active electromagnetic bearing rotor system at time t, a t represents the action output by the DDPG agent at time t, r t represents the total reward value obtained by the DDPG agent at time t, s t+1 represents the state of the active electromagnetic bearing rotor system at time t+1;

[0165] S304: Based on the state s of the active electromagnetic bearing rotor system at time t t and exploration noise N t Generate the action a output by the DDPG agent at time t through the action network t , and update the state s of the active electromagnetic bearing rotor system at time t+1 t+1 And calculate the total reward value r obtained by the DDPG agent at time t t ;

[0166] S305: The relevant data group (s) obtained after each exploration by the DDPG agent t , a t , r t , s t+1 ) are stored in the experience pool. If the number of samples in the experience pool is greater than the number of training samples, multiple data sets (s t , a t , r t , s t+1 ) to conduct training and update the target parameters, otherwise continue to execute step S304.

[0167] The steps to update the target parameters are as follows:

[0168] S3051: For the extracted small batch samples, the target Q value is calculated using the evaluation target network. The specific formula is:

[0169] ;

[0170] Among them, γ is the discount factor used to correct the online network, a t+1 represents the action output by the DDPG agent at time t+1, For state s t+1 Next, perform action a t+1 The target state - action value;

[0171] S3052: Update the evaluation network parameters of the DDPG agent by minimizing the loss function L; the minimization loss function is defined as the mean square error between the current Q value estimate and the calculated target Q value, specifically expressed as:

[0172] ;

[0173] in, Indicates the data set (s t , a t , r t , s t+1 ) expectations;

[0174] S3053: A delayed update strategy is adopted for the update of action network parameters. The action network parameters are updated by the gradient ascent method. The gradient formula for:

[0175] ;

[0176] in, Represents the Q-value function for action a t The gradient, is the policy function with respect to μ i The gradient, Indicates the expected value of the current action. represents the strategy at time t;

[0177] S3054: A soft update strategy is adopted for the DDPG agent. The action objective function and the evaluation objective function are used to calculate the target value. The action objective network and the evaluation objective network are updated as follows:

[0178] ;

[0179] In the formula are the action target network parameters, are the action network parameters, To evaluate the target network parameters, To evaluate the network parameters, is the sliding average coefficient, and its value is less than 1;

[0180] S306: Increase the number of time steps and perform the next round of training until the training is completed.

[0181] The expression of the observed variable is:

[0182] ;

[0183] Among them, R obserbation represents the observed variable, e error represents the rotor position error, T represents the total time, t represents the current moment, V rpm Indicates the system speed.

[0184] The expression of the multi-objective reward function is:

[0185] ;

[0186] ;

[0187] ;

[0188] ;

[0189] Among them, r error represents the accuracy reward, r response represents the response time reward, r overshoot represents the overshoot reward, a1 represents the error weight coefficient, a2 represents the response weight coefficient, a3 represents the overshoot weight coefficient, r set Indicates the value that the system setting needs to reach, r final Indicates the actual value after the system stabilizes, T response Indicates the time it takes for the system to stabilize, r max Indicates the maximum response value of the system during operation.

[0190] S4: Use the DDPG agent to optimize the undetermined parameters in the adaptive identification equation and dynamically adjust the compensation current based on the rotor speed to offset the unbalanced force disturbance.

[0191] The following simulation experiments are used to verify the effect of the control method described in the embodiment of the present invention:

[0192] The settings of the DDPG algorithm hyperparameters in the embodiment of the present invention are shown in Table 1.

[0193] Table 1 Algorithm hyperparameter settings

[0194]

[0195] After setting all parameters, the rotor system is trained offline with an initial rotor speed of 0 and an acceleration of a=20 (rad / min) / s. After the training is completed, the imbalance compensation of the rotor system can be achieved.

[0196] In the flywheel rotor constant speed model, the simulation is carried out under the condition of rotor speed ω=10000rpm, and the imbalance compensation algorithm based on DDPG agent is started at the first second. Figure 4 From the speed of unbalance coefficient identification at constant speed, it can be seen that the identification speed of this algorithm is 1.35s faster than that of the traditional algorithm.

[0197] according to Figure 5 It can be seen that when the constant speed ω=10000rpm, the maximum amplitude of the rotor is 2.5×10 -5 m is reduced to 2×10 - 17 m. Compared with the vibration suppression effect before compensation, the vibration amplitude of the rotor is reduced by 99%.

[0198] In the active electromagnetic bearing-rigid rotor acceleration motion model, the rotor is accelerated from a stationary state to 8000 rpm at an acceleration of a=20 (rad / min) / s; according to Figure 4 It can be seen that the maximum amplitude of the rotor is 2.4×10 -5 m is reduced to 1×10 -6 m, is reduced by 95%, which reflects the vibration suppression effect of the magnetic bearing compensation control algorithm under the rotor acceleration operation state.

[0199] from Figure 6 It can be seen that the unbalance coefficient identification speed of the embodiment of the present invention is much faster than that of the traditional algorithm. The identification speed of the unbalance coefficient of the embodiment of the present invention is only 0.15s, which can quickly identify the specific unbalance coefficient, thereby realizing rapid compensation of the unbalanced force.

[0200] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for controlling the imbalance of an electromagnetic bearing rotor system based on the DDPG algorithm The method for controlling the imbalance of an electromagnetic bearing rotor system based on the DDPG algorithm is characterized in that: The steps include: S1: Based on rotor dynamics, the four-degree-of-freedom AMB-rigid rotor dynamics equation is established, and decoupling control of the active electromagnetic bearing rotor system is achieved through linear state feedback decoupling; S2: Construct a compensation current based on the unbalanced forces acting on the four degrees of freedom of the rotor, extract the unbalance coefficient in the compensation current, and construct an adaptive identification equation; S3: Construct a DDPG agent, including an action network, an action target network, an evaluation network, an evaluation target network, and an experience pool; use the rotor position error, the rotor position error integral, and the rotor position error differential as observation variables for the system speed; design a multi-objective reward function to guide the DDPG agent to complete training and output the optimal action until training is completed. The training process of the DDPG agent is as follows: S301: Initialize the action network of the model for imbalance compensation of active electromagnetic bearing rotor system based on DDPG algorithm Weight and evaluation network Weight ; S302: Create an experience pool to store the experience data generated during the interaction between the DDPG agent and the external environment. Initialize the experience pool capacity to D, the number of training samples to n, and the maximum number of training rounds to Episodes; S303: Define the spatial state of the active electromagnetic bearing rotor system as ( s t ,a t ,r t ,s t+1 ), where s t represents the state of the active electromagnetic bearing rotor system at time t, a t represents the action output by the DDPG agent at time t, r t represents the total reward value obtained by the DDPG agent at time t, s t+1 represents the state of the active electromagnetic bearing rotor system at time t+1; S304: Based on the state s of the active electromagnetic bearing rotor system at time t t and exploration noise N t Generate the action a output by the DDPG agent at time t through the action network t , and update the state s of the active electromagnetic bearing rotor system at time t+1 t+1 And calculate the total reward value r obtained by the DDPG agent at time t t ; S305: The relevant data group (s) obtained after each exploration by the DDPG agent t , a t , r t , s t+1 ) are stored in the experience pool. If the number of samples in the experience pool is greater than the number of training samples, multiple data sets (s t , a t , r t , s t+1 ) to conduct training and update the target parameters, otherwise continue to execute step S304; S306: Increase the number of time steps and perform the next round of training until the training is completed; S4: Use the DDPG agent to optimize the undetermined parameters in the adaptive identification equation and dynamically adjust the compensation current based on the rotor speed to offset the unbalanced force disturbance.

2. The imbalance control method of an electromagnetic bearing rotor system based on the DDPG algorithm according to claim 1, characterized in that: The specific expression of the four-degree-of-freedom AMB-rigid rotor dynamics equation is: ; ; Where, J represents the lateral moment of inertia, J z represents the moment of inertia of the rotor around the z-axis, Indicates the rotor rotation speed, L bA Indicates the distance from the plane where the electromagnetic bearing at end A of the rotor is located to the plane of the rotor's center of mass, f xa Indicates that the electromagnetic bearing A is The electromagnetic force on the rotor in the direction, L bB Indicates the distance from the plane where the electromagnetic bearing at the B end of the rotor is located to the plane of the rotor center of mass, f xb Indicates that the electromagnetic bearing B is The electromagnetic force on the rotor in the direction, u z The distance between the center of mass of the unbalanced mass and the center of mass of the rotor is z The projection on z The distance between the unbalanced mass point and the center of mass is z The projection on the axis, m represents the rotor mass, Indicates that the rotor is The acceleration in the direction, represents the angular acceleration of the rotor around the x-axis, represents the angular velocity of the rotor around the x-axis, represents the angular acceleration of the rotor around the y-axis, represents the angular velocity of the rotor around the y-axis, f yb Indicates that the electromagnetic bearing B is The electromagnetic force on the rotor in the direction, f ya Indicates that the electromagnetic bearing A is The electromagnetic force on the rotor in the direction, Indicates that the rotor is The acceleration in the direction, represents the unbalanced mass, represents the rotor rotation speed, represents the rotor rotation angle, Represents ε z With O x The initial angle of the axis, Indicates the rotor rotation acceleration, Fu is the generalized unbalance vector of the electromagnetic bearing rotor system, and Fx represents the rotor The unbalanced force in the direction of the rotor is Fy. The unbalanced force in the direction It represents the unbalanced force on the rotor when it rotates around the x-axis. It represents the unbalanced force on the rotor when it rotates around the y-axis.

3. The imbalance control method of an electromagnetic bearing rotor system based on the DDPG algorithm according to claim 2, characterized in that: The active electromagnetic bearing rotor system The state equation and output equation are: ; in, represents the first-order derivative of X with respect to time, X represents the n-dimensional state vector, A represents the n×n-order state matrix, B represents the n×l-order input matrix, U represents the l-dimensional control vector, Y represents the m-dimensional output vector, and C represents the m×n-order output matrix; The active electromagnetic bearing rotor system After linear state feedback decoupling, it becomes a decoupled controlled system , decoupled controlled systems For an integral system, decouple the controlled system The expressions of the decoupling matrix R and the decoupling matrix F are: ; ; ; Among them, W1 represents the input matrix after decoupling, W2 represents the output matrix after decoupling, and W2 -1 represents the inverse matrix of the decoupled output matrix W2, T qb represents the arm matrix of the electromagnetic bearing-rotor system, K s The force-displacement stiffness coefficient matrix of the electromagnetic bearing rotor system, K i represents the force-current stiffness coefficient matrix of the electromagnetic bearing rotor system, M represents the generalized mass matrix of the rotor system, M -1 represents the inverse matrix of the generalized mass matrix M of the rotor system, L f represents the arm matrix, and G represents the gyro matrix.

4. The imbalance control method of an electromagnetic bearing rotor system based on the DDPG algorithm according to claim 3 is characterized in that: The specific expression of the compensation current is: ; ; ; ; ; Among them, I u Indicates the compensation current, I d represents the unbalanced disturbance, Indicates the unbalance coefficient of the rotor A end in the x-axis direction, Indicates the unbalance coefficient of the rotor A end in the y-axis direction, Indicates the unbalance coefficient of the rotor B end in the x-axis direction, It represents the unbalance coefficient in the y-axis direction of the rotor B end and defines and is the unbalance coefficient between the rotor ends A and B; The unbalance coefficient of the rotor A end and the rotor B-end unbalance coefficient In plural form: ; ; in, Indicates the unbalance coefficient of the rotor A end plural form of Indicates the unbalance coefficient of the rotor B end The plural form of , j represents the imaginary unit; Note the unbalance coefficient of rotor A end Identification value of and the rotor B-end unbalance coefficient Identification value of for: ; ; in, Indicates the unbalance coefficient identification value of the rotor A end in the x direction, Indicates the unbalance coefficient identification value of the rotor A end in the y direction, Indicates the unbalance coefficient identification value of the rotor B end in the x direction, Indicates the unbalance coefficient identification value in the y direction of the rotor end B; Then the adaptive identification equation of the unbalance coefficient at end A of the rotor is: ; The adaptive identification equation of the unbalance coefficient at the B end of the rotor is: ; in, Represents the second-order derivative of the unbalance coefficient identification value in the x direction at the rotor end A, Represents the first-order derivative of the unbalance coefficient identification value in the x direction at the rotor end A, x bA represents the displacement of the rotor A end in the x direction, y bA represents the displacement of rotor A end in the y direction, k1, k2 and k3 represent the undetermined parameters in the adaptive identification equation of the unbalance coefficient of rotor A end, k4, k5 and k6 represent the undetermined parameters in the adaptive identification equation of the unbalance coefficient of rotor B end, The second derivative of the unbalance coefficient identification value in the y direction at the rotor end A is represented by: Represents the first derivative of the unbalance coefficient identification value in the y direction at the rotor end A; Represents the second-order derivative of the unbalance coefficient identification value in the x direction at the rotor end A, Represents the first-order derivative of the unbalance coefficient identification value in the x direction at the B end of the rotor, x bB represents the displacement in the x direction of the rotor B end, y bB represents the displacement of the rotor end B in the y direction, Represents the second-order derivative of the unbalance coefficient identification value in the y direction at the B end of the rotor, Represents the first-order derivative of the unbalance coefficient identification value in the y direction at the B end of the rotor.

5. The imbalance control method of an electromagnetic bearing rotor system based on the DDPG algorithm according to claim 1, characterized in that: The expression of the observed variable in S3 is: ; Among them, R obserbation represents the observed variable, e error represents the rotor position error, T represents the total time, t represents the current moment, V rpm Indicates the system speed; The expression of the multi-objective reward function is: ; ; ; ; Among them, r error represents the accuracy reward, r response represents the response time reward, r overshoot represents the overshoot reward, a1 represents the error weight coefficient, a2 represents the response weight coefficient, a3 represents the overshoot weight coefficient, r set Indicates the value that the system setting needs to reach, r final Indicates the actual value after the system stabilizes, T response Indicates the time it takes for the system to stabilize, r max Indicates the maximum response value of the system during operation.

6. The imbalance control method of an electromagnetic bearing rotor system based on the DDPG algorithm according to claim 1, characterized in that: The steps for updating the target parameters during the DDPG agent training process in S3 are as follows: S3051: For the extracted small batch samples, the target Q value is calculated using the evaluation target network. The specific formula is: ; Among them, γ is the discount factor used to correct the online network, a t+1 represents the action output by the DDPG agent at time t+1, For state s t+1 Next, perform action a t+1 The target state - action value; S3052: Update the evaluation network parameters of the DDPG agent by minimizing the loss function L; the minimization loss function is defined as the mean square error between the current Q value estimate and the calculated target Q value, specifically expressed as: ; in, Indicates the data set (s t , a t , r t , s t+1 ) expectations; S3053: A delayed update strategy is adopted for the update of action network parameters. The action network parameters are updated by the gradient ascent method. The gradient formula for: ; in, Represents the Q-value function for action a t The gradient, is the policy function with respect to μ i The gradient, Indicates the expected value of the current action. represents the strategy at time t; S3054: A soft update strategy is adopted for the DDPG agent. The action objective function and the evaluation objective function are used to calculate the target value. The action objective network and the evaluation objective network are updated as follows: ; In the formula are the action target network parameters, are the action network parameters, To evaluate the target network parameters, To evaluate the network parameters, is the sliding average coefficient, and its value is less than 1.

Citation Information

Patent Citations

  • Unbalanced compensation algorithm of polygonal iterative search with variable step size for rotor imbalance coefficient

    CN107133387A

  • Unbalance compensation method for rotor unbalance coefficient variable-angle iterative search

    CN111177943A