Robust control method of electromagnetic stress-driven fast tool servo system

CN122652952APending Publication Date: 2026-08-28NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510235448.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0003]在实际加工过程中,FTS的工作精度要求较高,这对控制器的带宽和稳定性提出了较高的要求,如何在保证系统鲁棒性的同时,确保高精度运动跟踪仍是一个技术挑战

Benefits of technology

[0049] The beneficial effects of the present invention are: (1) The nonlinear control strategy designed in this invention with the extended state observer as its core effectively estimates and compensates for disturbances, uncertainties and unmodeled dynamics in the system by extending internal and external disturbances into a new state variable. The nonlinear controller designed in combination with the backstepping control method can effectively suppress high-frequency disturbances and system nonlinearity problems, ensuring high-precision motion trajectory tracking of the system; (2) The present invention meets the superior performance requirements of the control system. The near-end policy optimization (PPO) algorithm is used to optimize the controller and observer parameters. Based on the theory, a control system model based on deep reinforcement learning is constructed. The optimized control system can exhibit higher dynamic response capability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122652952A_ABST
    Figure CN122652952A_ABST
Patent Text Reader

Abstract

The application discloses a robust control method of an electromagnetic stress driving fast tool servo system, and comprises the following steps: S1, constructing a mathematical model of the electromagnetic stress driving fast tool servo system; S2, based on the mathematical model, estimating system states by using a model-based extended state observer to obtain expected estimation values; S3, obtaining a nonlinear control law based on the expected estimation values; and S4, setting a reinforcement learning parameter learner, learning a current action by using the learner, and outputting the learned current action to a controller. The application can effectively suppress high-frequency disturbances and nonlinear problems, guarantee high-precision motion trajectory tracking of the system, meet the performance requirements of the control system, and show higher dynamic response capability and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of servo control technology and relates to a robust control method for an electromagnetic stress-driven high-speed tool servo system. Background Technology

[0002] Complex optical surfaces have been widely used in various fields due to their numerous superior properties. However, the increasing surface complexity of these components poses greater challenges to their manufacturing technology. Fast Tool Servo (FTS) technology offers significant advantages in precision, complex, and high-speed machining. Electromagnetic stress-driven FTS, with its advantages of fast response speed and simple structure, is gradually becoming an effective means of ultra-precision manufacturing of optical microstructure components.

[0003] In actual processing, FTS requires high working accuracy, which places high demands on the bandwidth and stability of the controller. Ensuring high-precision motion tracking while maintaining system robustness remains a technical challenge. Summary of the Invention

[0004] The purpose of this invention is to provide a robust control method for an electromagnetic stress-driven rapid tool servo system, design a robust control system based on an extended state observer, and improve controller performance by utilizing the parameter optimization capabilities provided by reinforcement learning algorithms.

[0005] The technical solution adopted in this invention is: a robust control method for an electromagnetic stress-driven rapid tool servo system, comprising the following steps:

[0006] S1. Construct a mathematical model for an electromagnetic stress-driven rapid tool servo system.

[0007] S2. Based on the mathematical model, the system state is estimated using a model-based extended state observer to obtain the desired estimate.

[0008] S3. Obtain the nonlinear control law based on the expected value;

[0009] S4. Set reinforcement learning parameters for the learner, use the learner to learn the current action, and output the learned current action to the controller.

[0010] Furthermore, S1 specifically refers to:

[0011] Electromagnetic stress-driven FTS systems satisfy the following:

[0012]

[0013] in, To incorporate the total disturbance, including both internal and external disturbances, y, Here, denoted as the system's current displacement, current velocity, and current acceleration, respectively; d represents the external disturbance; t represents the current time; b represents the input matrix; and u represents the controller output value.

[0014] Choose state variable x1 = y, Equation (1) is transformed into a state-space equation to obtain the mathematical model of the electromagnetic stress-driven rapid tool servo system:

[0015]

[0016] Where f(x1,x2,x3) is the rapid tool servo system.

[0017] Furthermore, S2 specifically refers to:

[0018] For the third-order controlled object shown in equation (2), the total disturbance σ of internal and external disturbances is treated as a new unknown state variable.

[0019]

[0020] That is, a new state is extended from the original system, and the original system then becomes a linear system:

[0021]

[0022] Based on equation (4), the model-based extended state observer is established as follows:

[0023]

[0024] Where A, B, and C are the system matrix, input matrix, and output matrix, respectively, and L = [β1β2β3β4]. T Let σ be the gain matrix that includes the observation gain, where σ is the sum of the internal and external perturbations;

[0025] Expanding equation (5) above, we obtain the expected estimate:

[0026]

[0027] Among them, e ob For observation error, σ represents the observed displacement, observed velocity, observed acceleration, and total system disturbance, respectively. Let β1, β2, β3, and β4 be their corresponding derivatives, and let β1, β2, β3, and β4 be the first, second, third, and fourth order gains of the observer.

[0028] Substituting into equation (5), we get e ob The expression is as follows

[0029]

[0030] Choose an appropriate gain matrix L such that the eigenvalues ​​of the error dynamic matrix A-LC are all located in the left half of the complex plane, thus ensuring that the estimated state of the state expansion observer asymptotically converges to the actual state of the system.

[0031] Furthermore, S3 specifically refers to:

[0032] The nonlinear control law is as follows:

[0033]

[0034] Where u is the controller output value, and b is the input matrix. For the desired acceleration, This is an estimate of the expected displacement error. The derivative of the estimated value of the desired displacement error. Let be the second derivative of the estimated value of the desired displacement error. This is an estimate of the expected speed error. The derivative of the estimated value of the desired speed error. Let ρ be the estimated value of the desired acceleration error, ρ be the upper bound function of σ, γ be a constant with a value range of 0.1≤γ≤0.5, and k1, k2, k3 be the first, second, and third feedback gains, respectively.

[0035] Furthermore, S4 specifically refers to:

[0036] The reinforcement learning parameter learner includes an agent, which is trained using the PPO algorithm. The training process of the agent is as follows:

[0037] The agent obtains the current state information required for training. Based on the current state information required for training, the action network outputs the current action, outputs the reward function based on the current action, trains the action network based on the reward function, and repeats the above training process until the training objective is achieved, and then obtains the reinforcement learning parameter learner that has been trained.

[0038] The current state information s is:

[0039]

[0040] in,

[0041] x d (t), x d (t+1), e(t) and e(t+1) represent the target displacement, target velocity, target acceleration, target jerk, target displacement, target velocity, target acceleration, target jerk, current tracking error, and next tracking error, respectively.

[0042] The current action a is:

[0043] a=[k1,k2,k3,ω0] (5)

[0044] Where ω0 is the observer bandwidth;

[0045] The reward function is a mapping between s, a and the reward value of the current action. The reward function includes a regular reward r1 and a penalty r2. All rewards are negative. The regular reward ensures the stability of training. When the controller diverges, the output penalty r2 = -10.

[0046]

[0047] Where e represents the system tracking error;

[0048] The action network has three intermediate layers: the first layer is a feature extraction layer with 128 nodes and the activation function is ReLU; the second layer is a non-linear activation layer with 128 nodes; and the third layer is a fully connected layer with 128 nodes and the activation function is tanh.

[0049] The beneficial effects of the present invention are: (1) The nonlinear control strategy designed in this invention with the extended state observer as its core effectively estimates and compensates for disturbances, uncertainties and unmodeled dynamics in the system by extending internal and external disturbances into a new state variable. The nonlinear controller designed in combination with the backstepping control method can effectively suppress high-frequency disturbances and system nonlinearity problems, ensuring high-precision motion trajectory tracking of the system; (2) The present invention meets the superior performance requirements of the control system. The near-end policy optimization (PPO) algorithm is used to optimize the controller and observer parameters. Based on the theory, a control system model based on deep reinforcement learning is constructed. The optimized control system can exhibit higher dynamic response capability and robustness.

[0050] In addition to the objectives, features, and advantages described above, the present invention has other objectives, features, and advantages. The invention will now be described in further detail with reference to the figures. Attached Figure Description

[0051] Figure 1 This is a flowchart of the present invention;

[0052] Figure 2The graph shows the observation curves of the extended state observer of the present invention at the initial position x = 0, the system motion frequency 100 Hz, and the bandwidth ω0 = 3000.

[0053] Figure 3 The nonlinear control strategy of this invention uses control parameter k1 = 5 × 10 7 When k2 = 3000 and k3 = 40000, and the observation bandwidth ω0 = 3000, the simulated trajectory tracking curve is shown.

[0054] Figure 4 This is a graph showing the reward function curve for the training task of the reinforcement learning parameter learner of the present invention.

[0055] Figure 5 The nonlinear control strategy of this invention is shown in the trajectory tracking curve under actual cutting conditions, with an ideal signal of a sine curve of amplitude 4μm and frequency 10Hz. Detailed Implementation

[0056] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0057] A robust control method for an electromagnetic stress-driven rapid tool servo system includes the following steps:

[0058] S1. Construct a mathematical model for an electromagnetic stress-driven rapid tool servo system.

[0059] S2. Based on the mathematical model, the system state is estimated using a model-based extended state observer to obtain the desired estimate.

[0060] S3. Obtain the nonlinear control law based on the expected value;

[0061] S4. Set reinforcement learning parameters for the learner, use the learner to learn the current action, and output the learned current action to the controller.

[0062] S1 specifically refers to:

[0063] Consider an electromagnetic stress-driven FTS system that satisfies the following conditions:

[0064]

[0065] in, To incorporate the total disturbance from both internal and external disturbances, y, Let x1 be the current displacement, current velocity, and current acceleration of the system, d be the external disturbance, t be the current time, b be the input matrix, and u be the controller output value. Select the state variable x1 = y. Then (1) can be transformed into a state-space equation to obtain the mathematical model of the electromagnetic stress-driven rapid tool servo system:

[0066]

[0067] S2 specifically refers to:

[0068] For the third-order controlled object shown in equation (2), the sum of the internal and external disturbances of the process, σ, is treated as a new unknown state variable.

[0069]

[0070] That is, a new state is extended from the original system, and the original system then becomes a linear system.

[0071]

[0072] Based on equation (4), a model-based extended state observer is established as follows:

[0073]

[0074] Where A, B, and C are the system matrix, input matrix, and output matrix, respectively, and L = [β1β2β3β4]. T Let σ be the gain matrix that includes the observation gain σ, where σ is the sum of the internal and external perturbations.

[0075] Expanding the above equation, we get...

[0076]

[0077] e o b represents the observation error. σ represents the observed displacement, observed velocity, observed acceleration, and system disturbance, respectively. The system disturbance includes external and internal disturbances. Let β1, β2, β3, and β4 be their corresponding derivatives, and let β1, β2, β3, and β4 be the first, second, third, and fourth order gains of the observer.

[0078] Substituting into equation (5), we get e ob The expression is as follows

[0079]

[0080] By simply choosing a suitable gain matrix L such that the eigenvalues ​​of the error dynamic matrix A-LC are all located in the left half of the complex plane, it can be guaranteed that the estimated state of the state expansion observer asymptotically converges to the actual state of the system.

[0081] S3 specifically refers to:

[0082] Let the control objective of the system be y→x d Substituting equation (6) into the system's error equation, we obtain the following:

[0083]

[0084] Among them, e, For the tracking error and its derivative, x d , Let e ​​be the desired signal and its derivative. ob , Let represent the observation error and its derivative.

[0085] The controller is designed as follows:

[0086] Get Lyapunov function

[0087]

[0088] make but

[0089]

[0090] Right now Zhengding Negative constant, tracking error e→0

[0091] Therefore, but

[0092]

[0093] x 2d For observations The target signal, and take Let be the estimated value of the desired speed error, then

[0094]

[0095] Get Lyapunov function

[0096]

[0097] make but

[0098]

[0099] Right now Zhengding Negative constant, target error

[0100] make but

[0101]

[0102] x 3d For observations The target signal, and take Let ε be the estimated value of the desired acceleration error, then

[0103]

[0104] Get Lyapunov function

[0105]

[0106] make but

[0107]

[0108] in, This is a high-frequency robust controller. ρ is an upper bound function of σ, and γ is used to suppress chattering of the controller and limit the range of steady-state error, typically 0.1 ≤ γ ≤ 0.5.

[0109] Therefore Zhengding Negative constant, target error

[0110] At this point, e→0. The entire system is stable.

[0111] Substituting equation (18) into (20), we obtain the expression for the control law as follows:

[0112]

[0113] Where u is the controller output value, and b is the input matrix. For the desired acceleration, This is an estimate of the expected displacement error. The derivative of the estimated value of the desired displacement error. Let be the second derivative of the estimated value of the desired displacement error. This is an estimate of the expected speed error. The derivative of the estimated value of the desired speed error. Let ρ be the estimated value of the desired acceleration error, ρ be the upper bound function of σ, γ be a constant, generally in the range of 0.1≤γ≤0.5, and k1, k2, k3 be the first, second, and third feedback gains, respectively.

[0114] S4 specifically involves: setting reinforcement learning parameters for the learner, using the learner to learn the current action, and outputting the learned current action to the controller.

[0115] In reinforcement learning, the state space is a complete observation of the environment, meaning the current state *s* contains complete information about the environment. The state space has 10 state features, namely...

[0116]

[0117] The action space describes the set of all possible actions an agent can take in each state. In this invention, the action space is a 4-dimensional vector:

[0118] a=[k1,k2,k3,ω0] (23)

[0119] Where ω0 is the observer bandwidth, and this four-dimensional vector is restricted to a certain range to avoid the agent from overexploring.

[0120] The reward function in this invention can be viewed as a mapping between s,a and the reward value of the current action. The training objective of this algorithm is for the agent to output a set of parameters that enable FTS to achieve trajectory tracking while minimizing the tracking error, and simultaneously avoiding controller divergence caused by excessive output parameter deviation. To achieve this objective, the reward function includes a regular reward r1 and a penalty r2. All rewards are negative to prevent the agent from getting trapped in local optima. The regular reward ensures the stability of training, and when the controller diverges, the output penalty r2 = -10.

[0121]

[0122] Where e represents the displacement tracking error.

[0123] The action network has three intermediate layers. The first layer is a feature extraction layer with 128 nodes and the ReLU activation function. The second layer is a non-linear activation layer with 128 nodes. The third layer is a fully connected layer with 128 nodes and the tanh activation function.

[0124] To facilitate practical application by engineers in engineering scenarios, the control parameter adjustment rules and the impact of corresponding parameter changes on control system performance involved in this invention can be simply summarized as follows:

[0125] 1. Based on nonlinear controller parameter tuning:

[0126] For the design parameters k1, k2, and k3 of a nonlinear controller, in an engineering sense, it is sufficient to ensure that each value is positive. When specific adjustments are required, a set of control parameters can be tuned using the concept of pole placement, and this set of control parameters can meet the trajectory tracking requirements.

[0127] 2. Deep reinforcement learner parameter tuning:

[0128] When the selection of controller parameters in nonlinear controller tuning relies on experience and cannot further improve controller performance, a reinforcement learner can be used for parameter tuning. After training, the reinforcement learner will directly provide the optimal control parameters under the current conditions.

[0129] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A robust control method for an electromagnetic stress-driven rapid tool servo system, characterized in that, Includes the following steps: S1. Construct a mathematical model for an electromagnetic stress-driven rapid tool servo system. S2. Based on the mathematical model, the system state is estimated using a model-based extended state observer to obtain the desired estimate. S3. Obtain the nonlinear control law based on the expected value; S4. Set reinforcement learning parameters for the learner, use the learner to learn the current action, and output the learned current action to the controller.

2. The robust control method for the electromagnetic stress-driven rapid tool servo system according to claim 1, characterized in that, S1 specifically refers to: Electromagnetic stress-driven FTS systems satisfy the following: in, To encompass the total disturbance, including both internal and external disturbances. Here, denoted as the system's current displacement, current velocity, and current acceleration, respectively; d represents the external disturbance; t represents the current time; b represents the input matrix; and u represents the controller output value. Choose the state variable x1 = y, Equation (1) is transformed into a state-space equation to obtain the mathematical model of the electromagnetic stress-driven rapid tool servo system: Where f(x1,x2,x3) is the rapid tool servo system.

3. The robust control method for the electromagnetic stress-driven rapid tool servo system according to claim 2, characterized in that, S2 specifically refers to: For the third-order controlled object shown in equation (2), the total disturbance σ of internal and external disturbances is treated as a new unknown state variable. That is, a new state is extended from the original system, and the original system then becomes a linear system: Based on equation (4), the model-based extended state observer is established as follows: Where A, B, and C are the system matrix, input matrix, and output matrix, respectively, and L = [β1β2β3β4]. T Let σ be the gain matrix that includes the observation gain, where σ is the sum of the internal and external perturbations; Expanding equation (5) above, we obtain the expected estimate: Among them, e ob For observation error, These are the observed displacement, observed velocity, observed acceleration, and total system disturbance, respectively. Let β1, β2, β3, and β4 be their corresponding derivatives, and let β1, β2, β3, and β4 be the first, second, third, and fourth order gains of the observer. Substituting into equation (5), we get e ob The expression is as follows Choose an appropriate gain matrix L such that the eigenvalues ​​of the error dynamic matrix A-LC are all located in the left half of the complex plane, thus ensuring that the estimated state of the state expansion observer asymptotically converges to the actual state of the system.

4. The robust control method for the electromagnetic stress-driven rapid tool servo system according to claim 3, characterized in that, S3 specifically refers to: The nonlinear control law is as follows: Where u is the controller output value, and b is the input matrix. For the desired acceleration, This is an estimate of the expected displacement error. The derivative of the estimated value of the desired displacement error. Let be the second derivative of the estimated value of the desired displacement error. This is an estimate of the expected speed error. The derivative of the estimated value of the desired speed error. Let ρ be the estimated value of the desired acceleration error, ρ be the upper bound function of σ, γ be a constant with a value range of 0.1≤γ≤0.5, and k1, k2, k3 be the first, second, and third feedback gains, respectively.

5. The robust control method for the electromagnetic stress-driven rapid tool servo system according to claim 4, characterized in that, S4 specifically refers to: The reinforcement learning parameter learner includes an agent, which is trained using the PPO algorithm. The training process of the agent is as follows: The agent obtains the current state information required for training. Based on the current state information required for training, the action network outputs the current action, outputs the reward function based on the current action, trains the action network based on the reward function, and repeats the above training process until the training objective is achieved, and then obtains the reinforcement learning parameter learner that has been trained. The current state information s is: in, These represent the target displacement, target velocity, target acceleration, target jerk at the current moment, target displacement, target velocity, target acceleration, target jerk at the next moment, target jerk at the next moment, tracking error at the current moment, and tracking error at the next moment, respectively. The current action a is: a=[k1,k2,k3,ω0] (5) Where ω0 is the observer bandwidth; The reward function is a mapping between s, a and the reward value of the current action. The reward function includes a regular reward r1 and a penalty r2. All rewards are negative. The regular reward ensures the stability of training. When the controller diverges, the output penalty r2 = -10. Where e represents the system tracking error; The action network has three intermediate layers: the first layer is a feature extraction layer with 128 nodes and the activation function is ReLU; the second layer is a non-linear activation layer with 128 nodes; and the third layer is a fully connected layer with 128 nodes and the activation function is tanh.