Adaptive feedforward disturbance compensation control method and system based on reinforcement learning
By adopting an adaptive feedforward disturbance compensation control method based on reinforcement learning, and combining feedforward and feedback control, the control strategy is adjusted in real time, which solves the vibration control problem caused by load centroid changes and achieves accurate compensation and stability improvement in high dynamic scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI
- Filing Date
- 2026-05-19
- Publication Date
- 2026-07-21
Smart Images

Figure CN122219057B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mechanical control technology, and in particular relates to an adaptive feedforward disturbance compensation control method and system based on reinforcement learning. Background Technology
[0002] With the rapid advancement of technology, especially the rapid development of high-precision fields such as semiconductor manufacturing, precision lithography, and nanometer measurement, the requirements for vibration control in equipment are becoming increasingly stringent. Active vibration isolation technology has become one of the key technologies to ensure the stable operation of these devices. Active vibration isolators effectively reduce the impact of vibration on equipment performance by sensing external and internal vibration disturbances in real time and generating appropriate compensation forces based on the disturbance signals.
[0003] However, in practical applications, active vibration isolation technology still faces many challenges, especially in handling vibrations caused by load motion, velocity changes, or changes in the load's center of mass position. In these scenarios, the sources and characteristics of vibrations are often complex, particularly the dynamic nonlinear effects caused by changes in the load's center of mass, which further exacerbates the difficulty of vibration control. Most existing control methods are based on fixed load parameters (such as mass, center of mass position, and moment of inertia), neglecting the dynamic impact of changes in the load's center of mass position on the system. Therefore, they cannot effectively address the nonlinear disturbances caused by real-time changes in the center of mass position during high-speed load movement. In such cases, traditional vibration control methods often fail to achieve ideal compensation effects, especially in applications with rapid load changes and high vibration frequencies, where the vibration suppression effect is significantly reduced.
[0004] In existing technologies, feedforward control is a widely used vibration control strategy. Feedforward control methods measure known input or disturbance signals and calculate the control input in advance based on these signals to suppress the impact of disturbances. This method is suitable for dealing with periodic or predictable disturbances. However, in practical applications, because changes in the load's center of mass can cause nonlinear changes in disturbances, traditional feedforward control often cannot effectively compensate for dynamic disturbances caused by changes in the load's center of mass. For example, feedforward algorithms can reduce the settling time caused by platform movement, but this method relies on a fixed model and has poor dynamic adaptability to changes in the load's center of mass, making it difficult to meet the demands of high-dynamic, high-precision applications. While feedforward control can reduce disturbances caused by platform motion, its compensation effect is limited under high load and rapid movement conditions because it relies heavily on gas spring pressure feedback.
[0005] In addition, nonlinear control methods such as sliding mode control and iterative learning control have been proposed and applied to vibration control. Sliding mode control can effectively handle uncertainties and nonlinear disturbances in the system, but its dynamic response capability often fails to meet expectations when faced with rapidly changing disturbances. Iterative learning control technology performs well in handling periodic disturbances and repetitive disturbances, but its compensation for nonlinear disturbances caused by changes in the load's center of mass is still limited, and it cannot achieve real-time accurate compensation. Especially when facing complex nonlinear disturbances such as changes in friction and inertia, existing control methods often suffer from insufficient modeling accuracy and weak dynamic compensation capability, resulting in the performance and robustness of the control system failing to meet the requirements of high-precision applications. Summary of the Invention
[0006] In view of this, the present invention aims to provide an adaptive feedforward disturbance compensation control method and system based on reinforcement learning. By combining reinforcement learning control technology, the control strategy can be adjusted in real time when the load centroid changes, thereby significantly improving the system's accuracy and stability. Furthermore, the present invention combines feedforward and feedback control mechanisms, effectively addressing multi-source disturbances, improving system robustness, and ensuring excellent performance in high-precision, high-dynamic environments. It is particularly suitable for fields with extremely high vibration control requirements, such as semiconductor lithography machines and precision measurement equipment.
[0007] To achieve the above objectives, the technical solution created by this invention is implemented as follows:
[0008] An adaptive feedforward disturbance compensation control method based on reinforcement learning includes:
[0009] S1: Place the load to be vibrated on the vibration isolation platform;
[0010] S2: Control the vibration isolation platform to isolate the moving load, and at the same time collect the driving force of the current vibration isolation platform on the load, the position of the center of mass of the load on the vibration isolation platform, and the acceleration of the vibration isolation platform;
[0011] S3: The three types of information collected in step S2 constitute the state of the vibration isolation platform; the current state is input into the trained reinforcement learning model to obtain the compensation force for feedforward vibration isolation of the load;
[0012] S4: Based on the vibration error of the vibration isolation platform, obtain the feedback signal for feedback control of the vibration isolation platform; use the compensation force of the feedforward vibration isolation obtained in step S3 and the feedback signal to perform composite control of the vibration isolation platform.
[0013] Furthermore, during the training process of the reinforcement learning model in step S3: a reward function is constructed using the acceleration of the vibration isolation platform, and the evaluation value of the reinforcement learning model is updated using the Bellman equation; a greedy strategy is adopted to select the compensation force corresponding to the maximum evaluation value, and the compensation force at this time is output.
[0014] Furthermore, the reward function is:
[0015] R(t) = -|a(t+1)| 2 ;
[0016] Where R(t) represents the reward function in the t-th training step, and a(t) represents the acceleration of the vibration isolation platform in the t-th training step.
[0017] Furthermore, the Bellman equation is as follows:
[0018] ;
[0019] Where Q(S(t),A(t)) represents the evaluation value corresponding to the state S(t) and the compensating force A(t) at the t-th training step, α represents the learning rate, and γ represents the discount factor. This represents the maximum evaluation value of the state during the (t+1)th training step.
[0020] Furthermore, in step S4: based on the vibration error of the vibration isolation platform, a PI controller outputs a feedback signal; the feedback signal is added to the compensation force to obtain the vibration isolation force for composite control of the vibration isolation platform.
[0021] Furthermore, step S4 also includes bandpass filtering the feedback signal output by the PI controller before adding it to the compensation force.
[0022] An adaptive feedforward disturbance compensation control system based on reinforcement learning includes: a vibration isolation platform for placing a load; a position sensor for sensing the position of the load on the vibration isolation platform; multiple vibration isolation components, each including a voice coil motor and an acceleration sensor group, the acceleration sensor group for sensing the horizontal and vertical acceleration of the vibration isolation platform, and the voice coil motor for providing horizontal and vertical vibration isolation forces to the vibration isolation platform; and a control center for acquiring the position information output by the position sensor, the acceleration information output by the acceleration sensor group, and the vibration isolation force output by the voice coil motor, and using the adaptive feedforward disturbance compensation control method based on reinforcement learning provided by this invention to output the vibration isolation force to the voice coil motor.
[0023] Furthermore, the vibration isolation force output by the multiple vibration isolation components and the vibration isolation force experienced at the center point of the vibration isolation platform satisfy the following relationship:
[0024] F F =R F ×F M ;
[0025] Among them, F F R represents the vibration isolation force output by multiple vibration isolation components. FLet F represent the allocation matrix. M This indicates the vibration isolation force acting on the center point of the vibration isolation platform.
[0026] Furthermore, the accelerations collected by multiple vibration isolation components satisfy the following relationship with the acceleration at the center point of the vibration isolation platform:
[0027] A F =R A ×A M ;
[0028] Among them, A F R represents the acceleration collected by multiple vibration isolation components. A A represents the transformation matrix between the accelerations collected by multiple vibration isolation components and the acceleration at the center point of the vibration isolation platform. M This indicates the acceleration at the center point of the vibration isolation platform.
[0029] Compared with the prior art, the present invention can achieve the following beneficial effects:
[0030] This invention presents an adaptive feedforward disturbance compensation control method and system based on reinforcement learning. Traditional vibration control methods rely on complex physical models, such as load centroid, mass, and moment of inertia, requiring precise debugging and updates of these models during system deployment. This invention employs a reinforcement learning algorithm, transforming control from a reliance on fixed physical models to an online learning process based on real-time data. By collecting data such as platform position, driving force, and acceleration, the system can automatically optimize feedforward compensation without pre-establishing complex physical models, reducing the complexity of system configuration and debugging, and improving system flexibility and adaptability. More importantly, this invention incorporates the load position into the reinforcement learning state, enabling the reinforcement learning model to better adapt to changes in the load centroid, thereby avoiding limited compensation for nonlinear disturbances caused by load centroid changes and achieving real-time, accurate compensation.
[0031] This invention addresses the challenge of traditional vibration isolation technologies failing to compensate for real-time changes in the center of mass caused by load motion, which often relies on fixed physical models. It proposes an innovative composite control method combining "data-driven feedforward + traditional feedback." The core of this method lies in incorporating the real-time center of mass position of the load, along with the platform's driving force and acceleration, into the state input of a reinforcement learning model. This allows the model to directly perceive and learn the complex, nonlinear disturbance dynamics caused by changes in the center of mass, thereby generating an adaptive feedforward compensation force online. Combined with feedback control, this method outputs the optimal vibration isolation force. This approach eliminates the reliance on precise prior models, achieving real-time and accurate compensation for time-varying disturbance sources, significantly improving vibration isolation accuracy and system robustness in high-dynamic scenarios. Attached Figure Description
[0032] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments and descriptions of the invention are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:
[0033] Figure 1 A flowchart illustrating the adaptive feedforward disturbance compensation control method based on reinforcement learning as described in an embodiment of the present invention;
[0034] Figure 2 A flowchart illustrating the adaptive feedforward disturbance compensation control method based on reinforcement learning as described in an embodiment of the present invention;
[0035] Figure 3 A schematic diagram of the overall structure of the adaptive feedforward disturbance compensation control system based on reinforcement learning as described in the embodiment of the present invention;
[0036] Figure 4 A schematic diagram of the structure of the vibration isolation component described in the embodiment of the present invention;
[0037] Figure 5 A schematic diagram illustrating the direction of the vibration isolation force output by each voice coil motor in the embodiments of the present invention;
[0038] Figure 6 This is a schematic diagram showing the direction of acceleration sensed by each of the acceleration sensors described in the embodiments of the present invention.
[0039] Explanation of reference numerals in the attached figures:
[0040] 1. Vibration isolation platform; 2. Vibration isolation components; 3. Load; 4. Motion platform; 5. Upper support plate; 6. Lower support plate; 7. Voice coil motor; 8. Accelerometer; 9. Support spring. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not constitute a limitation thereof.
[0042] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.
[0043] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0044] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0045] The invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0046] like Figure 1 and Figure 2 As shown in the embodiment of the present invention, the adaptive feedforward disturbance compensation control method based on reinforcement learning includes:
[0047] S1: Place the load to be vibrated on the vibration isolation platform.
[0048] S2: Control the vibration isolation platform to isolate the moving load, and at the same time collect the driving force of the current vibration isolation platform on the load, the position of the center of mass of the load on the vibration isolation platform, and the acceleration of the vibration isolation platform.
[0049] In this embodiment of the invention, the acceleration of the vibration isolation platform includes six dimensions: taking the center of mass of the load at the initial equilibrium position of the vibration isolation platform as the origin, the direction perpendicular to the upper surface of the vibration isolation platform as the Z direction, any direction on the plane of the vibration isolation platform as the X direction, and the direction perpendicular to the X direction as the Y direction, a right-handed Cartesian coordinate system is constructed. The acceleration of the vibration isolation platform can then be expressed as a(t) = [a x+ (t),a y+ (t),a z+ (t),a x- (t),a y- (t),a z- [(t)], where a(t) represents the acceleration of the vibration isolation platform at time t, a x+ (t), a y+ (t) and a z+(t) represent the accelerations of the vibration isolation platform in the positive X, Y, and Z directions at time t, respectively. x- (t), a y- (t) and a z- (t) represents the acceleration of the vibration isolation platform in the negative X, negative Y, and negative Z directions at time t, respectively.
[0050] S3: The three types of information collected in step S2 constitute the state of the vibration isolation platform; the current state is input into the trained reinforcement learning model to obtain the compensation force for feedforward vibration isolation of the load. In this embodiment of the invention, step S3 can be expressed by the following formula:
[0051] A(t) = π(S(t));
[0052] Where A(t) represents the compensation force of the feedforward vibration isolation on the load at time t, S(t) represents the state of the vibration isolation platform at time t, and π represents the reinforcement learning model.
[0053] In this embodiment of the invention, the state can be represented as S(t) = [F motion (t),P motion [(t),a(t)],F motion (t) represents the driving force P exerted by the vibration isolation platform on the load at time t. motion (t) represents the position of the load on the vibration isolation platform. Furthermore, in this embodiment of the invention, the compensation force for feedforward vibration isolation of the load is a six-dimensional force vector, specifically including a three-dimensional linear force and a three-dimensional torque. The compensation force can be specifically expressed as A(t) = [F...]. x (t),F y (t),F z (t),M x (t),M y (t),M z [(t)],F x (t), F y (t) and F z (t) represent the linear forces acting on the load at time t from the X, Y, and Z directions, respectively, and M. x (t), M y (t) and M z (t) represents the torques in three directions that the load experiences at time t.
[0054] In some embodiments, the training process of the reinforcement learning model includes: constructing a reward function based on the acceleration of the vibration isolation platform, updating the evaluation value of the reinforcement learning model using the Bellman equation, selecting the compensation force corresponding to the maximum evaluation value using a greedy strategy, and outputting the compensation force at this time.
[0055] In this embodiment of the invention, a deep reinforcement learning model is constructed using a deep Q-network algorithm. The input is the current state S(t), and the output is the compensation force A(t). The deep Q-network is used to map the state S(t) to the compensation force A(t). The reward function needs to ensure the achievement of the control objective, i.e., minimizing the acceleration a(t). In this way, the smaller the platform acceleration, the larger the reward, and the reinforcement learning model will tend to select a better compensation force.
[0056] The deep Q-network needs sufficient capacity (i.e., enough parameters) to learn the complex nonlinear mapping from high-dimensional states to six-dimensional compensating forces. The input / output layer dimensions are defined by the system physics. The number of hidden layers ranges from [3,5]. In this embodiment, three hidden layers are specifically set, with 128, 256, and 512 neurons in each layer, respectively. A ReLU activation function is set after each hidden layer.
[0057] The reward function is as follows:
[0058] R(t) = -|a(t+1)| 2 ;
[0059] Where R(t) represents the reward function in the t-th training step, and a(t) represents the acceleration of the vibration isolation platform in the t-th training step;
[0060] The Bellman equation is as follows:
[0061] ;
[0062] Where Q(S(t),A(t)) represents the evaluation value corresponding to the state S(t) and the compensating force A(t) at the t-th training step, α represents the learning rate, and γ represents the discount factor. This represents the maximum evaluation value of the state at training step t+1. Through multiple iterations using the Bellman equation, the system learns how to select the optimal compensation force based on the system's state, thereby minimizing the acceleration of the vibration isolation platform.
[0063] In this embodiment of the invention, in a standard deep Q-network, training is achieved by minimizing the mean squared time difference error, and the training loss function is:
[0064] ;
[0065] Where L(θ) represents the loss function, which measures the quality of the evaluation value Q(S(t), A(t)) output by the current reinforcement learning model (deep Q-network). The smaller the loss, the more accurate the evaluation of the action value. θ represents the parameters of the online network in the deep Q-network. -Let A' and S' represent the parameters of the target network in the deep Q-network, respectively, and let E represent the compensation force and evaluation value of the target network output.
[0066] In reinforcement learning, the ε-greedy policy controls the probability of selecting random actions (exploration). It typically starts with a high value (e.g., 1.0) and decays to a low value (e.g., 0.01) over time or steps. The initial high exploration helps discover effective compensation strategies under various load positions and perturbations. Decaying to a low value ensures stable execution of the learned optimal policy later. The decay scheme (e.g., linear decay, exponential decay) and speed are also factors that need to be set. Simultaneously, the online network parameters are replicated at a fixed frequency using a soft update method, i.e., a very small update rate τ. This provides a more stable learning target; the smaller update rate τ makes the target network change slowly, resulting in more stable training. Each time, a batch of experiences is sampled from the experience replay pool for gradient updates. Larger batches make gradient estimation more accurate and training more stable, but increase memory consumption and computation. Therefore, a trade-off between stability and efficiency needs to be struck and adaptively adjusted. The hyperparameters used in the training process include the base learning rate α, which is adaptively adjusted according to the actual situation. The Adam optimizer is used, with the first moment estimate of the exponential decay rate β1=0.9, the second moment estimate of the exponential decay rate β2=0.999, and the numerical stability constant ε=1e-8.
[0067] S4: Based on the vibration error of the vibration isolation platform, obtain the feedback signal for feedback control of the vibration isolation platform; use the compensation force of the feedforward vibration isolation obtained in step S3 and the feedback signal to perform composite control of the vibration isolation platform.
[0068] This invention combines feedforward control and feedback control to achieve a precise and stable control strategy. Feedforward control, based on real-time load centroid position and driving force signals, predicts and compensates for disturbances in advance through a data-driven model. The feedback control section employs a PI controller with a bandpass filter to correct for residual disturbances. Specifically, in some embodiments, step S4 includes:
[0069] S41: Based on the vibration error of the vibration isolation platform, a PI controller outputs a feedback signal. The principle of the PI controller in this embodiment of the invention is as follows:
[0070] ;
[0071] Where μ(t) represents the feedback signal output by the PI controller at time t, and K p and K lLet represent the proportional gain and integral gain in the PI controller, respectively, and ε(t) represent the vibration error of the vibration isolation platform at time t. The vibration error is essentially the negative of the acceleration signal. The error value is equal to the ideal value minus the actual value a(t). If the ideal value is 0, then 0 - a(t) = -a(t) = ε(t).
[0072] In some embodiments, to further improve the system's response capability and suppress high-frequency noise, step S41 further includes bandpass filtering the feedback signal output by the PI controller. In this embodiment, bandpass filtering can be expressed by the following formula:
[0073] H(s) = s / (s 2 +2λω n s+ω n 2 );
[0074] Where H(s) represents the bandpass filter, λ represents the damping coefficient, and ω n This represents the center frequency of the bandpass filter. Through the bandpass filter, the system can effectively filter out low-frequency and high-frequency noise, retaining only signals related to the target control frequency, thereby improving the accuracy and response speed of feedback control. This composite control strategy ensures that the system applies compensating force in time before disturbances occur, and simultaneously corrects the control input through feedback after a disturbance, ensuring the platform quickly returns to a stable state.
[0075] S42: The feedback signal is added to the compensation force to obtain the vibration isolation force for composite control of the vibration isolation platform. This can be understood as adding the bandpass-filtered feedback signal to the compensation force to obtain the composite control vibration isolation force. In this embodiment of the invention, the composite control vibration isolation force can be obtained by the following formula:
[0076] F M (t)=μ H (t)+A(t);
[0077] Among them, F M μ represents the vibration isolation force exerted on the center point of the vibration isolation platform at time t. H (t) represents the feedback signal after bandpass filtering.
[0078] This invention also provides an adaptive feedforward disturbance compensation control system based on reinforcement learning, such as... Figure 3As shown, the system includes a vibration isolation platform 1, position sensors, multiple vibration isolation components 2, and a control center. The vibration isolation platform 1 is used to place a load 3; the position sensors are used to sense the position of the load 3 on the vibration isolation platform 1; each vibration isolation component 2 includes a voice coil motor and an acceleration sensor group. The acceleration sensor group senses the horizontal and vertical acceleration of the vibration isolation platform 1, and the voice coil motor provides horizontal and vertical vibration isolation forces to the vibration isolation platform 1; the control center collects the position information output by the position sensors, the acceleration information output by the acceleration sensor group, and the vibration isolation force output by the voice coil motor group, and uses the reinforcement learning-based adaptive feedforward disturbance compensation control method provided by this invention to output the vibration isolation force to the voice coil motor group. The position sensor can be adaptively selected according to the actual situation. In this embodiment, the position sensor can be a grating ruler, a laser displacement sensor, an inertial measurement unit (IMU), etc.
[0079] In this embodiment of the invention, the load 3 moves on a motion platform 4 (such as a workpiece stage for processing the load 3, or a laser scanning stage for scanning the load 3, etc.). The motion platform 4 is mounted on the upper surface of the vibration isolation platform 1, and specifically, four vibration isolation components 2 support the vibration isolation platform 1 and isolate the load 3 on the vibration isolation platform 1 from vibration. Meanwhile, each vibration isolation component 2 designed in this embodiment of the invention is as follows: Figure 4 As shown, the system includes an upper support plate 5, a lower support plate 6, two voice coil motors 7, and two acceleration sensors 8. The upper support plate 5 contacts the lower surface of the vibration isolation platform 1. The two voice coil motors 7 are positioned between the upper support plate 5 and the lower support plate 6 to support the upper support plate 5. The two voice coil motors 7 are fixed diagonally on the upper support plate 5, providing active vibration isolation forces in the horizontal and vertical directions, respectively. The two acceleration sensors 8 are located near the two voice coil motors 7 and connected to the lower surface of the upper support plate 5, used to sense the horizontal and vertical accelerations of the vibration isolation platform 1 through the upper support plate 5. Furthermore, in this embodiment, a support spring 9 for passive vibration isolation is also provided between the two voice coil motors 7. The two ends of the support spring 9 are connected to the upper support plate 5 and the lower support plate 6, respectively. In this case, the two voice coil motors 7 and the support spring 9 enable the vibration isolation assembly 2 to simultaneously achieve active and passive vibration isolation.
[0080] In this embodiment of the invention, four vibration isolation components 2 are set to achieve active and passive vibration isolation. The control system provided in this embodiment includes eight voice coil motors 7 and eight acceleration sensors 8. At this time, the vibration isolation force F output by the eight voice coil motors 7 is... F It can be represented as F F =[F1,F2,F3,...,F8], the direction of the vibration isolation force output by each of the 8 voice coil motors 7 is as follows: Figure 5 As shown; acceleration A sensed by 8 accelerometers 8 F It can be represented as A F =[SA1 ,S A2 ,S A3 ,...,S A8 The directions of acceleration sensed by each of the eight accelerometers are as follows: Figure 6 As shown.
[0081] In some embodiments, the vibration isolation force and acceleration output by the plurality of vibration isolation components 2 respectively satisfy the following relationships with the vibration isolation force and acceleration experienced at the center point of the vibration isolation platform:
[0082] F F =R F ×F M ;
[0083] A F =R A ×A M ;
[0084] Among them, F F R represents the vibration isolation force output by multiple vibration isolation components 2. F Let A represent the allocation matrix. F R represents the acceleration collected by multiple vibration isolation components 2. A A represents the transformation matrix between the accelerations collected by multiple vibration isolation components 2 and the acceleration at the center point of the vibration isolation platform. M This represents the acceleration at the center point of the vibration isolation platform. Both the distribution matrix and the transformation matrix are calculated using geometric relationships.
[0085] The innovation of this invention lies in incorporating the position of the load's center of mass on the vibration isolation platform into the state. Specifically, this invention integrates the load's position on the vibration isolation platform into the state, which is not merely adding a sensor signal, but rather changing the control paradigm: feedforward compensation is based on a fixed physical model that assumes the load's center of mass remains constant; when the load's position on the vibration isolation platform changes, the model fails. Feedforward compensation is generated by a data-driven reinforcement learning model with the load's position on the vibration isolation platform as the key input. Through training, it learns how changes in the load's position on the vibration isolation platform affect the system dynamics and dynamically adjusts the compensation strategy.
[0086] The adaptive feedforward compensation force A(t) is directly output by the trained reinforcement learning model. Its "adaptive" capability stems from reward function-driven and model-free learning. Specifically, in reward function-driven learning, the model is trained with minimizing the next platform acceleration as the reward function. This means the model is forced to learn how to predict and counteract upcoming vibrations based on the current state (including the real-time position of the load on the vibration isolation platform). In model-free learning, this invention transforms control from relying on a fixed physical model to an online learning process based on real-time data. Traditional methods show a significant decrease in vibration suppression effectiveness when the load's center of mass changes rapidly. However, this invention, by introducing the load position state, enables the model to avoid limited compensation for nonlinear disturbances caused by changes in the load's center of mass, thereby achieving real-time and accurate compensation. This significantly improves the system's performance in highly dynamic scenarios.
[0087] The control method and control system provided by this invention can easily adapt to different application requirements. For example, in the aerospace field, it can be used for vibration control of equipment such as satellites and rockets; in the automotive industry, it can be used for powertrain and comfort control of high-end vehicles; in earthquake engineering and building structures, it can be used for seismic control of large buildings, bridges, etc. Even in the field of high-performance robotics, the control method of this invention can be effectively applied to precision and micro-manipulation tasks, ensuring the accurate stability of the system under dynamic disturbances. Therefore, this invention has strong substitutability and cross-industry application potential.
[0088] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.
[0089] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. An adaptive feedforward disturbance compensation control method based on reinforcement learning, characterized in that, include: S1: Place the load to be vibrated on the vibration isolation platform; S2: Control the vibration isolation platform to isolate the moving load, and at the same time collect the driving force of the current vibration isolation platform on the load, the position of the center of mass of the load on the vibration isolation platform, and the acceleration of the vibration isolation platform; S3: The three types of information collected in step S2 constitute the state of the vibration isolation platform; the current state is input into the trained reinforcement learning model to obtain the compensation force for feedforward vibration isolation of the load; S4: Based on the vibration error of the vibration isolation platform, obtain the feedback signal for feedback control of the vibration isolation platform; use the compensation force of the feedforward vibration isolation obtained in step S3 and the feedback signal to perform composite control of the vibration isolation platform.
2. The adaptive feedforward disturbance compensation control method based on reinforcement learning according to claim 1, characterized in that, During the training of the reinforcement learning model in step S3: The reward function is constructed using the acceleration of the vibration isolation platform, and the evaluation value of the reinforcement learning model is updated using the Bellman equation. A greedy strategy is used to select the compensation force corresponding to the maximum evaluation value, and the compensation force at this time is output.
3. The adaptive feedforward disturbance compensation control method based on reinforcement learning according to claim 2, characterized in that, The reward function is: R(t)=-|a(t+1)| 2 ; Where R(t) represents the reward function in the t-th training step, and a(t) represents the acceleration of the vibration isolation platform in the t-th training step.
4. The adaptive feedforward disturbance compensation control method based on reinforcement learning according to claim 2, characterized in that, The Bellman equation is: ; Where Q(S(t),A(t)) represents the evaluation value corresponding to the state S(t) and the compensating force A(t) at the t-th training step, α represents the learning rate, and γ represents the discount factor. This represents the maximum evaluation value of the state during the (t+1)th training step.
5. The adaptive feedforward disturbance compensation control method based on reinforcement learning according to claim 1, characterized in that, In step S4: Based on the vibration error of the vibration isolation platform, a PI controller is used to output a feedback signal; By adding the feedback signal to the compensation force, the vibration isolation force for composite control of the vibration isolation platform is obtained.
6. The adaptive feedforward disturbance compensation control method based on reinforcement learning according to claim 5, characterized in that, Step S4 also includes bandpass filtering the feedback signal output by the PI controller before adding it to the compensation force.
7. An adaptive feedforward disturbance compensation control system based on reinforcement learning, characterized in that, include: Vibration isolation platform, used to place loads; Position sensors are used to detect the position of the load on the vibration isolation platform; Multiple vibration isolation components, each including a voice coil motor assembly and an acceleration sensor assembly. The acceleration sensor assembly is used to sense the horizontal and vertical acceleration of the vibration isolation platform, and the voice coil motor assembly is used to provide horizontal and vertical vibration isolation forces to the vibration isolation platform. The control center is used to collect position information output by the position sensor, acceleration information output by the acceleration sensor group, and vibration isolation force output by the voice coil motor group. It uses the adaptive feedforward disturbance compensation control method based on reinforcement learning as described in any one of claims 1 to 6 to output vibration isolation force to the voice coil motor group.
8. The adaptive feedforward disturbance compensation control system based on reinforcement learning according to claim 7, characterized in that, The vibration isolation force output by multiple vibration isolation components satisfies the following relationship with the vibration isolation force acting on the center point of the vibration isolation platform: F F =R F ×F M ; Among them, F F R represents the vibration isolation force output by multiple vibration isolation components. F Let F represent the allocation matrix. M This indicates the vibration isolation force acting on the center point of the vibration isolation platform.
9. The adaptive feedforward disturbance compensation control system based on reinforcement learning according to claim 7, characterized in that, The accelerations collected by multiple vibration isolation components satisfy the following relationship with the acceleration at the center point of the vibration isolation platform: A F =R A ×A M ; Among them, A F R represents the acceleration collected by multiple vibration isolation components. A A represents the transformation matrix between the accelerations collected by multiple vibration isolation components and the acceleration at the center point of the vibration isolation platform. M This indicates the acceleration at the center point of the vibration isolation platform.