Adaptive linear active-disturbance-rejection magnetic compensation control method and system based on reinforcement learning
By adopting an adaptive linear active disturbance rejection magnetic compensation control method based on reinforcement learning, a linear extended state observer and a linear state error feedback control law are designed. Combined with the parameter adaptive module of reinforcement learning, the shortcomings of the shielding performance of the passive magnetic shielding device and the PID controller are solved, and the effective suppression of magnetic field disturbances and the improvement of system robustness are achieved.
Patent Information
- Application Number
- CN202511529081.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2026-01-23
AI Technical Summary
Existing passive magnetic shielding devices have limited shielding performance against the Earth's magnetic field, and PID controllers are insufficient in terms of control accuracy and robustness, making it difficult to adapt to environmental changes.
An adaptive linear active disturbance rejection control method based on reinforcement learning is adopted. By designing a linear extended state observer and a linear state error feedback control law, combined with a parameter adaptive module based on reinforcement learning, the magnetic field disturbance is effectively suppressed.
It improves the ability to suppress magnetic field disturbances and the robustness of the system, and can automatically adjust parameters when external disturbances change, thereby improving control accuracy and adaptability.
Smart Images

Figure CN121386397A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of active magnetic compensation technology of magnetic shielding devices, and particularly relates to a self-adaptive linear active disturbance rejection magnetic compensation control method and system based on reinforcement learning. BACKGROUND
[0002] With the development of social economy, the global people's life has changed greatly, and cardiovascular diseases have become the main cause of death. Early detection of heart disease is crucial to prevent fatal complications. Magnetocardiogram (MCG) measurement based on quantum magnetometer is a new non-invasive heart diagnosis technology, which can reflect some subtle changes and abnormalities of the heart that some electrocardiogram signals cannot reflect, and therefore can help doctors diagnose early lesions of the cardiovascular system. However, compared with the uT order of magnitude of the geomagnetic signal, the MCG signal is only pT order of magnitude, so the geomagnetic field needs to be shielded to enhance the signal-to-noise ratio of the MCG signal measurement.
[0003] Now there are many passive magnetic shielding devices that use material properties to shield the geomagnetic field. High permeability materials such as permalloy, ferrite or nanocrystalline are used to shield low-frequency magnetic fields; high conductivity materials such as aluminum alloy are used to eliminate high-frequency magnetic fields through eddy current effect. However, due to the limited performance of passive magnetic shielding, an active magnetic field suppression system is needed to further reduce the magnetic field disturbance inside the magnetic shielding device by generating a magnetic field opposite to the magnetic field disturbance through a control coil. Most current systems use proportional-integral-derivative (PID) controllers, which are easy to debug but have limited control accuracy and robustness.
[0004] A linear active disturbance rejection controller can estimate disturbances through a linear extended state observer, achieve disturbance-free estimation through appropriate observer gain adjustment, and achieve accurate disturbance suppression through a linear state error feedback control law. However, the parameter adjustment of the linear active disturbance rejection controller is complex and difficult to manually adjust to an optimal state. When the environment changes, the system's disturbance rejection capability also changes. By introducing an adaptive mechanism, the system can automatically adjust to an optimal state and adapt to environmental changes when external disturbances change. SUMMARY
[0005] In order to enable the active magnetic compensation system to adaptively adjust the optimal control parameters according to the disturbance changes, the application provides a self-adaptive linear active disturbance rejection magnetic compensation control method and system based on reinforcement learning, which aims to achieve effective suppression of magnetic field disturbances by designing a linear extended state observer based on system parameters and a linear state error feedback control law and adjusting the parameters. A parameter adaptive module based on reinforcement learning is designed to ensure good robustness of the system.
[0006] To achieve the above objectives, the present invention provides the following solution: An adaptive linear active disturbance rejection (ADRR) compensation control method based on reinforcement learning, the method comprising: Establish a mathematical model for the active magnetic compensation control system; Design of a linear extended state observer based on a mathematical model of an active magnetic compensation control system; Based on the constructed mathematical model of the active magnetic compensation control system and the linear extended state observer, a linear error feedback control law is designed. Parameter tuning is performed based on the designed linear extended state observer and linear error feedback control law; By using the tuning parameters as actions in reinforcement learning, an adaptive reinforcement learning mechanism is designed to perform parameter adaptation and thus suppress disturbances.
[0007] Preferred methods for establishing mathematical models of active magnetic compensation control systems include: ; in, , The total disturbance includes both external environmental disturbances and system parameter fluctuations. This refers to the conversion ratio coefficient of the digital-to-analog converter. The filter time constant is This is the voltage-controlled current source conversion coefficient. The coil constant represents the conversion ratio between current and magnetic field. It is the conversion ratio of the analog-to-digital converter. This is the conversion ratio coefficient between the magnetic field value and the voltage output. The time constant of a first-order inertial element. For magnetic field disturbance, y The output of the feedback channel, For controller output, for The first derivative, for The second derivative, , and All of these are model-related parameters. It is a complex frequency variable.
[0008] Preferably, the method for designing a linearly extended state observer includes: ; in, , , and For the system parameter matrix, The system state matrix, , and are the estimated values of state variables and , is expressed as wherein , , are the estimated values of , and , is an observer gain matrix, wherein , and are observer gains, is the first derivative of the estimated value of state variable .
[0009] Preferably, the method for designing a linear state error feedback control law comprises: ; wherein is a reference value, and are two adjustable controller gains.
[0010] Preferably, the parameter tuning of the extended state observer comprises: ; pole assignment to to obtain an observer gain matrix is expressed as: ; When the observer gain satisfies the preset requirement, the unknown disturbance is estimated, so that , the three adjustable parameters of the observer gain matrix are tuned to one adjustable parameter ; The parameter tuning of the linear state error feedback control law comprises: ; The transfer function is written as , and pole assignment is performed to to ensure the stability of the system, so that the controller gain matrix is: ; The two adjustable parameters of the controller gain matrix are tuned to one adjustable parameter .
[0011] Preferably, the designed reinforcement learning adaptive mechanism comprises: an action space, an observation space, a reward function, and a reinforcement learning module.
[0012] Preferred methods for designing motion space include: ; in, For the observer bandwidth, For controller bandwidth; The design methods for observation spaces include: ; in, This refers to the digital quantity of voltage converted from data collected by the sensor. Three state variables estimated for a linearly extended state observer; The design methods for reward functions include: ; in, It is the weight of the integral error. Indicates the integral error. It is a classification error; The design methods for reinforcement learning modules include: ; in, The output of the critics' network This is the current state. For action, As a reward, It is a regularization coefficient that represents the weight of entropy. Indicates the first The network weight parameters of a commentator network. Indicates the current weight The following strategy.
[0013] The present invention also provides an adaptive linear active disturbance rejection compensation control system based on reinforcement learning. The system is used to implement the aforementioned method and includes: a construction module, a first design module, a second design module, a parameter tuning module, and a third design module. The building module is used to establish a mathematical model of the active magnetic compensation control system; The first design module is used to design a linear extended state observer based on the mathematical model of the active magnetic compensation system; The second design module is used to design a linear error feedback control law based on the constructed mathematical model of the active magnetic compensation system and the linear extended state observer; The parameter tuning module is used to tune parameters based on the designed extended state observer and linear state error feedback control law. The third design module is used for designing the reinforcement learning adaptive mechanism as an action of the reinforcement learning with the setting parameter, performing parameter adaptation and thus performing disturbance suppression.
[0014] Compared with the prior art, the application has the following beneficial effects: The linear active disturbance rejection magnetic compensation control system designed in the application can effectively estimate and suppress magnetic field disturbance, and improve the disturbance suppression capability of the system.
[0015] The application involves the parameter adaptive mechanism based on reinforcement learning, the system can actively adjust the adjustable parameters of the system when the external disturbance changes, effectively suppress the disturbance, and effectively improve the robustness and precision of the active magnetic compensation system. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the application, the drawings needed in the embodiments are briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0017] Figure 1 The control system design flowchart of the adaptive linear active disturbance rejection magnetic compensation control system based on reinforcement learning in the embodiment of the application is shown in the figure. Figure 2 The mathematical model diagram of the active magnetic compensation control system of the magnetocardiography device in the embodiment of the application is shown in the figure. Figure 3 The schematic diagram of the linear active disturbance rejection controller in the embodiment of the application is shown in the figure. Figure 4 The schematic diagram of the reinforcement learning parameter adaptive mechanism in the embodiment of the application is shown in the figure. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only some of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.
[0019] In order to make the above-mentioned purposes, features and advantages of the application more obvious and easy to understand, the application will be further described in detail below with reference to the drawings and specific embodiments.
[0020] Embodiment one As Figure 1As shown, the application provides a kind of adaptive linear active disturbance rejection magnetic compensation control method based on reinforcement learning, method includes: establishing active magnetic compensation control system mathematical model;Linear extended state observer is designed based on active magnetic compensation system mathematical model;Linear error feedback control law is designed based on the model and linear extended state observer constructed;Parameter setting is carried out based on the designed extended state observer and linear state error feedback control law;The setting parameter is as the action of reinforcement learning, and reinforcement learning adaptive mechanism is designed, and parameter adaptation is carried out to carry out disturbance suppression.
[0021] As Figure 2 Shown, the mathematical model of active magnetic compensation control system of the application is constructed as follows: Active magnetic compensation system is composed of controller, digital-analog converter, current source, coil, quantum magnetometer sensor, analog-digital converter.Wherein G c (s) is the transfer function of controller, G da (s) is the transfer function of digital-analog converter, which is composed of low-pass filter and proportional link, and is expressed as: ; Wherein, is the conversion proportional coefficient of digital-analog converter, is filter time constant, is complex frequency variable.The coil can be approximated as proportional link, and is expressed as , and feedback sensor is expressed as , which is expressed as proportional link and first-order inertia link: ; Wherein, is the conversion proportional coefficient of magnetic field value and voltage output, is the time constant of first-order inertia link.Analog-digital converter conversion proportional coefficient is expressed as , and reference input is , and magnetic field disturbance is , and system output is The disturbance transfer function of system can be written as: ; Wherein, is the conversion coefficient of voltage-controlled current source; The output of controller output to feedback channel y It is described as: ; Wherein is the controller output, and the parameter information of model is substituted to obtain: ; Wherein, isthe first derivative of , is , , , , , , and are model-related parameters, The total disturbance containing the external environmental disturbance of the system and the fluctuation of the system parameters, the state variable of the system is defined as The state equation of the system can be written as: ; wherein , and are system state variables, , and are the first derivatives thereof, is the first derivative of the disturbance term .
[0022] As shown in Figure 3 , the linear active disturbance rejection controller design process of the application is as follows: Step (1), designing a full-order linear extended state observer includes: ; wherein , , and are system parameter matrices, is a system state matrix, , and are the estimated values of state variables and , is represented as , wherein , , represent the estimated values of , and , is an observer gain matrix, wherein , and are observer gains, is the first derivative of the estimated value of the state variable .
[0023] Step (2), designing a linear state error feedback control law includes: ; where, is the reference value, controller gain is the adjustable parameter, there are two adjustable controller gains and .
[0024] Step (3), parameter setting of the extended state observer includes: ; pole assignment to get the observer gain matrix is expressed as: ; Step (4), parameter setting of the linear state error feedback control law includes: ; Write the transfer function as , pole assignment to to ensure the stability of the system, the controller gain matrix is: ; Set the two adjustable parameters of the controller gain matrix to one adjustable parameter .
[0025] As shown in Figure 4 , the reinforcement learning parameter adaptive mechanism includes: 1) Action space. The selection of the action space is the bandwidth of the observer and the bandwidth of the controller . The action space is: ; 2) Observation space. The observation space is composed of three state quantities estimated by the linear extended state observer, which is converted from the data collected by the sensor into a digital quantity of voltage , and is expressed as: ; 3) Reward function. The reward function is defined as: ; where is the weight of the integral error, represents the integral error, which is expressed as: ; where represents time, represents the basic error of the actual magnetic field and the expected magnetic field, which is expressed as: ; It is the grading error, expressed as: ; in, These are the weighting coefficients for the residual disturbances in different environments. It is a coefficient, represented as 4) Reinforcement Learning Module. The reinforcement learning module employs a soft actor-critic framework, learning from data acquired from the environment, including the current state. ,action ,award and the state at the next moment The role of the critic network is to approximate... The function is represented as: ; in express The state space at any given moment, express The space of action at any moment This represents the state currently obtained by the agent. Actions performed by the intelligent agent The rewards obtained below Indicates the current real-time policy The expectations below It is based on strategy The actions performed Indicating in strategy Entropy under the given conditions, and its representation strategy The degree of randomness, It is a regularization coefficient, representing the weight of entropy, which affects the exploration level of the reinforcement learning agent and increases... It helps accelerate policy learning and reduce the likelihood of policies getting stuck in local errors.
[0026] The role of the actor network is to complement the output of the critic network. Optimize the action distribution to make it closer to the actual distribution, with weights of 1. , The updated expression is: .
[0027] in Indicates the first The network weight parameters of a commentator network. Indicates the current weight The following strategy.
[0028] Embodiment two The application also provides an adaptive linear active disturbance rejection magnetic compensation control system based on reinforcement learning, which is used to implement the method of embodiment one, and comprises a construction module, a first design module, a second design module, a parameter tuning module and a third design module. The construction module is used to establish a mathematical model of the active magnetic compensation control system. The first design module is used to design a linear extended state observer based on the mathematical model of the active magnetic compensation system. The second design module is used to design a linear error feedback control law based on the established mathematical model of the active magnetic compensation system and the linear extended state observer. The parameter tuning module is used to tune parameters based on the designed extended state observer and linear state error feedback control law. The third design module is used to design a reinforcement learning adaptive mechanism with the tuning parameters as actions of reinforcement learning, to perform parameter adaptation and disturbance suppression.
[0029] The above embodiments are only used to describe the preferred modes of the application, and do not limit the scope of the application. Without departing from the design spirit of the application, various modifications and improvements to the technical solutions of the application made by those skilled in the art shall fall within the protection scope of the claims of the application.
Claims
1. A method for adaptive linear active disturbance rejection magnetic compensation control based on reinforcement learning, characterized in that, The method comprises: establishing a mathematical model of the active magnetic compensation control system; designing a linear extended state observer based on the mathematical model of the active magnetic compensation control system; designing a linear error feedback control law based on the established mathematical model of the active magnetic compensation control system and the linear extended state observer; performing parameter tuning based on the designed linear extended state observer and the linear error feedback control law; designing a reinforcement learning adaptive mechanism by taking the tuned parameters as actions of the reinforcement learning, performing parameter adaptation, and thus suppressing disturbances.
2. The method of claim 1, wherein, The method for establishing the mathematical model of the active magnetic compensation control system comprises: ; wherein , total disturbance comprising external environmental disturbances of the system and system parameter fluctuations, conversion scale factor for a digital-to-analog converter, filter time constant, conversion scale factor for a voltage-controlled current source, conversion scale factor for the coil constant representing the conversion of current and magnetic field, conversion scale factor for an analog-to-digital converter, conversion scale factor for the magnetic field value to the voltage output, time constant for a first-order inertial element, magnetic field disturbance, y output of the feedback channel, controller output, first derivative of second derivative of model-related parameters, model-related parameters, , and are model-related parameters, complex frequency variable.
3. The method of claim 2, wherein, The method for designing the linear extended state observer comprises: ; wherein , , and are system parameter matrices, is a system state matrix, , and are the estimated values of state quantities and , denotes wherein , , denote the estimated values of , and , is an observer gain matrix, wherein , and are observer gains, is the first derivative of the estimated value of the state quantity .
4. The method of claim 3, wherein, The method for designing the linear error feedback control law comprises: ; wherein is a reference value, and are two adjustable controller gains.
5. The method of claim 4, wherein, The parameter tuning of the extended state observer comprises: ; are configured to the poles of obtaining an observer gain matrix is represented as: ; When the observer gain satisfies the preset requirement, the unknown disturbance is estimated, so that three adjustable parameters of the observer gain matrix are set as one adjustable parameter ; The parameter tuning of the linear error feedback control law comprises: ; The transfer function is written as The pole assignment is performed to To ensure the stability of the system, the controller gain matrix is obtained as ; tuning two adjustable parameters of a controller gain matrix to one adjustable parameter .
6. The method of claim 1, wherein, The designed reinforcement learning adaptive mechanism comprises: an action space, an observation space, a reward function, and a reinforcement learning module.
7. The method of claim 6, wherein, The method for designing the action space comprises: ; wherein, is the observer bandwidth, is the controller bandwidth; The method for designing the observation space comprises: ; wherein, is a digital quantity converted from the data collected by the sensor into a voltage, are three state quantities estimated by the linear extended state observer; The method for designing the reward function comprises: ; wherein is a weight of the integral error, denotes the integral error, is the hierarchical error; The method for designing the reinforcement learning module comprises: ; wherein, the output of the critic network, for the current state, for the action, for the reward, is a regularization coefficient representing the weight of the entropy, denotes the network weight parameters of the th critic network, denotes the policy under the current weights .
8. A reinforcement learning based adaptive linear active disturbance rejection magnetic compensation control system for implementing the method of any one of claims 1-7, characterized in that, The system comprises: a construction module, a first design module, a second design module, a parameter tuning module, and a third design module; The construction module is configured to establish a mathematical model of the active magnetic compensation control system; The first design module is configured to design a linear extended state observer based on the mathematical model of the active magnetic compensation control system; The second design module is configured to design a linear error feedback control law based on the established mathematical model of the active magnetic compensation control system and the linear extended state observer; The parameter tuning module is configured to perform parameter tuning based on the designed linear extended state observer and the linear error feedback control law; The third design module is configured to design a reinforcement learning adaptive mechanism by taking the tuned parameters as actions of the reinforcement learning, perform parameter adaptation, and thus suppress disturbances.