Laser power stability control method and system
By combining the optical frequency shift principle and reinforcement learning model, the stability problem of laser power control in complex environments is solved, achieving high-precision and robust control of laser power, and adapting to the dynamic changes of nonlinear and time-varying systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING INST OF RADIO METROLOGY & MEASUREMENT
- Filing Date
- 2025-12-26
- Publication Date
- 2026-05-08
AI Technical Summary
Existing laser power control methods have weak stability and robustness in highly fluctuating, nonlinear, and time-varying systems, and cannot effectively cope with complex environmental changes, resulting in inaccurate laser power control.
A laser power measurement system based on the principle of optical frequency shift is built. It is trained and optimized by combining a reinforcement learning model. High-precision measurement is achieved through the optical frequency shift effect. Feedback control is performed using the state space, action space and reward function of the reinforcement learning model to optimize the driving signal and stabilize the laser power.
It achieves highly robust control of laser power in dynamic environments, has stronger adaptive capabilities, and can maintain the stability and accuracy of laser power in nonlinear and time-varying systems.
Smart Images

Figure CN122000778A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of optoelectronic technology, and in particular to a laser power stabilization control method and system. Background Technology
[0002] Lasers are widely used in precision measurement, quantum computing, quantum frequency standards and other research. However, even small fluctuations in laser power have a significant impact on the accuracy and reliability of experimental results. Therefore, how to achieve precise and stable control of laser power has become a key research focus in related technical fields.
[0003] Currently, laser power stabilization control typically employs traditional methods such as PID control, fuzzy control, and model predictive control (MPC). PID control, as a common feedback regulation method, can adjust laser power by modifying the gain parameter; however, its stability and robustness are weak in highly fluctuating, nonlinear, and time-varying systems, especially under complex external disturbances where it cannot effectively maintain precise control. Furthermore, existing control methods fail to adequately consider adaptive adjustment capabilities due to environmental changes and variations in light source characteristics, thus limiting their effectiveness in variable environments.
[0004] A laser power measurement method based on the optical frequency shift principle (AC Stark effect) can provide high-precision power measurement. However, how to further utilize it for precise feedback control to achieve long-term stability of laser power remains a technical challenge. Summary of the Invention
[0005] This invention provides a laser power stabilization control method and system to solve the problem of accurate and stable control of laser power in dynamic and complex interference environments.
[0006] In a first aspect, the present invention provides a laser power stabilization control method, comprising:
[0007] A laser power measurement system is built based on the optical frequency shift effect. The laser power measurement system includes a laser, an acousto-optic driving module, an atomic clock, and a frequency measurement module. The acousto-optic driving module is used to adjust the laser power, and the frequency measurement module is used to measure the output frequency of the atomic clock.
[0008] A reinforcement learning model is trained and optimized. The state space of the reinforcement learning model includes the frequency deviation of the current control cycle, the rate of frequency change of the current control cycle relative to the previous control cycle, and the driving signal of the previous control cycle. The frequency deviation of the current control cycle is the difference between the output frequency of the atomic clock in the current control cycle and the standard frequency. The rate of frequency change of the current control cycle relative to the previous control cycle is the difference between the frequency deviation of the current control cycle and the frequency deviation of the previous control cycle. The action space of the reinforcement learning model includes the increment of the driving signal in the current control cycle. The action space is obtained based on the control policy and the state space and is within a preset range. The reward function of the reinforcement learning model is used to minimize the frequency deviation and suppress rapid frequency changes.
[0009] The driving signal for the current control cycle is obtained based on the driving signal of the previous control cycle and the action space of the optimized reinforcement learning model.
[0010] The acousto-optic drive module is controlled based on the drive signal of the current control cycle to stabilize the laser power.
[0011] Optionally, the state space of the reinforcement learning model includes the frequency deviation of the current control cycle, the rate of change of the frequency of the current control cycle relative to the previous control cycle, and the driving signal of the previous control cycle, including:
[0012] s k =[Δv k ,Δv k -Δv k-1 ,u k-1 ]
[0013] Among them, s k Let Δv be the state space of the current control cycle. k The frequency deviation of the current control cycle, Δv k-1 The frequency deviation of the previous control cycle, Δv k -Δv k-1 u is the rate of change of frequency of the current control cycle relative to the previous control cycle. k-1 This refers to the drive signal of the previous control cycle.
[0014] Optionally, the action space is obtained based on the control strategy and the state space, and includes, within a preset range:
[0015] a k =π θ (s k )
[0016] Among them, a k For the action space of the current control cycle, a k ∈[amin ,a max ],π θ Let θ be the control strategy, and θ be the strategy parameter.
[0017] Optionally, the reward function of the reinforcement learning model for minimizing frequency bias and suppressing rapid frequency changes includes:
[0018]
[0019] Where, r k The reward signal for the current control cycle is represented by α and β, which are weighting coefficients.
[0020] Optionally, obtaining the driving signal for the current control cycle based on the driving signal of the previous control cycle and the action space of the optimized reinforcement learning model includes:
[0021] u k =u k-1 +a k
[0022] Among them, u k This is the drive signal for the current control cycle.
[0023] In a second aspect, the present invention provides a laser power stabilization control system, comprising a laser power measurement system, a first processing module, a reinforcement learning module, and a second processing module, wherein:
[0024] The laser power measurement system is built based on the optical frequency shift effect and includes a laser, an acousto-optic driving module, an atomic clock, and a frequency measurement module. The acousto-optic driving module is used to adjust the laser power, and the frequency measurement module is used to measure the output frequency of the atomic clock.
[0025] The first processing module is used to calculate the frequency deviation of the current control cycle and the frequency change rate of the current control cycle relative to the previous control cycle. The frequency deviation of the current control cycle is the difference between the output frequency of the atomic clock in the current control cycle and the standard frequency. The frequency change rate of the current control cycle relative to the previous control cycle is the difference between the frequency deviation of the current control cycle and the frequency deviation of the previous control cycle.
[0026] The reinforcement learning module is used to deploy the trained and optimized reinforcement learning model. The state space of the reinforcement learning model includes the frequency deviation of the current control cycle, the rate of change of the frequency of the current control cycle relative to the previous control cycle, and the driving signal of the previous control cycle. The action space of the reinforcement learning model includes the increment of the driving signal of the current control cycle. The action space is obtained based on the control policy and the state space and is within a preset range. The reward function of the reinforcement learning model is used to minimize the frequency deviation and suppress rapid changes in frequency.
[0027] The second processing module is used to obtain the driving signal of the current control cycle based on the driving signal of the previous control cycle and the action space of the optimized reinforcement learning model. The driving signal of the current control cycle is used to control the acousto-optic driving module to stabilize the laser power.
[0028] Optionally, the state space of the reinforcement learning model includes the frequency deviation of the current control cycle, the rate of change of the frequency of the current control cycle relative to the previous control cycle, and the driving signal of the previous control cycle, including:
[0029] s k =[Δv k ,Δv k -Δv k-1 ,u k-1 ]
[0030] Among them, s k Let Δv be the state space of the current control cycle. k The frequency deviation of the current control cycle, Δv k-1 The frequency deviation of the previous control cycle, Δv k -Δv k-1 u is the rate of change of frequency of the current control cycle relative to the previous control cycle. k-1 This refers to the drive signal of the previous control cycle.
[0031] Optionally, the action space is obtained based on the control strategy and the state space, and includes, within a preset range:
[0032] a k =π θ (s k )
[0033] Among them, a k For the action space of the current control cycle, a k ∈[a min ,a max ],π θ Let θ be the control strategy, and θ be the strategy parameter.
[0034] Optionally, the reward function of the reinforcement learning model for minimizing frequency bias and suppressing rapid frequency changes includes:
[0035]
[0036] Where, r k The reward signal for the current control cycle is represented by α and β, which are weighting coefficients.
[0037] Optionally, obtaining the driving signal for the current control cycle based on the driving signal of the previous control cycle and the action space of the optimized reinforcement learning model includes:
[0038] u k =u k-1 +a k
[0039] Among them, u k This is the drive signal for the current control cycle.
[0040] The above scheme achieves high-precision measurement of laser power through the optical frequency shift principle (AC Stark effect), and combines it with a reinforcement learning model to optimize power stability. The control strategy is continuously trained and adjusted under different environmental changes to ensure optimal control performance and stabilize laser power. Compared with traditional control methods, this scheme has stronger adaptive capabilities, not only able to handle nonlinear and time-varying systems, but also achieving highly robust control performance in dynamic environments. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a schematic flowchart of a laser power stabilization control method provided in Embodiment 1 of the present invention;
[0043] Figure 2 This is a schematic diagram of a laser power stabilization control system provided in Embodiment 2 of the present invention. Detailed Implementation
[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0045] Example 1
[0046] like Figure 1 As shown, this embodiment provides a laser power stabilization control method, including:
[0047] S1: A laser power measurement system is built based on the optical frequency shift effect. The laser power measurement system includes a laser, an acousto-optic drive module, an atomic clock, and a frequency measurement module. The acousto-optic drive module is used to adjust the laser power, and the frequency measurement module is used to measure the output frequency of the atomic clock.
[0048] When a laser beam strikes an atomic clock, it causes a frequency shift, resulting in a change in the clock's output frequency. This change in output frequency is correlated with the change in laser power; therefore, changes in laser power can be reflected by monitoring the changes in the atomic clock's output frequency. Based on a pre-calibrated relationship between frequency deviation and laser power, the system can estimate the current laser power according to the frequency deviation.
[0049] By adjusting the drive signal of the acousto-optic drive module, closed-loop control of the laser power is achieved, thereby enabling precise adjustment and long-term stability of the laser power.
[0050] Specifically, the acousto-optic driving module is an acousto-optic modulator (AOM).
[0051] S2: Train and optimize the reinforcement learning model. The state space of the reinforcement learning model includes the frequency deviation of the current control cycle, the rate of change of the frequency of the current control cycle relative to the previous control cycle, and the driving signal of the previous control cycle. The frequency deviation of the current control cycle is the difference between the output frequency of the atomic clock in the current control cycle and the standard frequency. The rate of change of the frequency of the current control cycle relative to the previous control cycle is the difference between the frequency deviation of the current control cycle and the frequency deviation of the previous control cycle. The action space of the reinforcement learning model includes the increment of the driving signal in the current control cycle. The action space is obtained based on the control policy and the state space and is within a preset range. The reward function of the reinforcement learning model is used to minimize the frequency deviation and suppress rapid frequency changes.
[0052] This invention employs a discrete-time control framework, taking one complete working cycle of an atomic clock as a control period, with the start time of the kth control period denoted as time k.
[0053] The state space is used to characterize system information related to laser power control and serves as the input signal for the reinforcement learning model.
[0054] Specifically, the state space of the reinforcement learning model includes the frequency deviation of the current control cycle, the rate of change of the frequency of the current control cycle relative to the previous control cycle, and the driving signal of the previous control cycle, including:
[0055] s k =[Δv k ,Δv k -Δv k-1 ,u k-1 ]
[0056] Among them, sk Let Δv be the state space of the current control cycle. k The frequency deviation of the current control cycle, Δv k-1 The frequency deviation of the previous control cycle, Δv k -Δv k-1 u is the rate of change of frequency of the current control cycle relative to the previous control cycle. k-1 This is the drive signal for the previous control cycle.
[0057] Depending on the actual needs, other observations such as estimated laser power can also be added to the state space.
[0058] The action space describes the actions an agent performs on driving signals in each control cycle. As the output signal of the reinforcement learning model, its value can be positive or negative.
[0059] By limiting the action space within a preset range, the amplitude of a single control can be limited to protect the device and avoid over-adjustment.
[0060] The action space can be a continuous space or it can be discretized into several fixed increments.
[0061] A reinforcement learning agent contains a control policy that represents the mapping from the state space to the action space.
[0062] Specifically, the action space is obtained based on the control strategy and the state space, and includes the following within a preset range:
[0063] a k =π θ (s k )
[0064] Among them, a k For the action space of the current control cycle, a k ∈[a min ,a max ],π θ Here, θ represents the control strategy, and θ represents the strategy parameters.
[0065] The reward function is used to quantify the laser power control effect in each control cycle and serves as the basis for optimizing the reinforcement learning algorithm.
[0066] Specifically, the reward function of a reinforcement learning model, used to minimize frequency bias and suppress rapid changes in frequency, includes:
[0067]
[0068] Where, r k The reward signal for the current control cycle is α, and β are weighting coefficients used to balance steady-state error and dynamic smoothness.
[0069] This reward function penalizes frequency deviation and excessively rapid frequency changes. When the output frequency is close to the standard frequency for a long time and changes steadily, the penalty term is small and the reward signal value is large. When the frequency deviation increases or the frequency fluctuates drastically, the penalty term increases and the reward signal value decreases. This ensures that the output frequency is close to the standard frequency for a long time, while avoiding over-adjustment of the system and improving overall stability.
[0070] During the training phase, the reinforcement learning AI will adjust its state based on the current state space s in each control cycle. k Choose an action space a k The driving signal is adjusted to change the laser power, which in turn affects the output frequency of the atomic clock, and a new state space s is formed in the next control cycle. k+1 At the end of each control cycle, the reward signal r is calculated according to the reward function described above. k And through the reward signal r k For control strategy π θ The performance of the policy is evaluated, and the policy parameters θ (such as neural network weights) are updated using a predetermined reinforcement learning algorithm, so that the control policy π improves after multiple interactions with the environment. θ The output action tends to be optimized under all possible conditions, thereby achieving stable control of laser power and minimizing frequency fluctuations.
[0071] As training progresses, the agent continuously optimizes its control strategy. θ By exploring (trying different actions) and utilizing (adopting the current better control strategy π) θ The system balances the parameters π and π, gradually converging to a set of approximately optimal policy parameters θ, thus obtaining the optimal control policy π. θ .
[0072] After training, the reinforcement learning model needs to be evaluated for performance. Long-term testing is conducted to observe the fluctuation range of the atomic clock output frequency and corresponding Allan variance, etc., to verify the stability of the control strategy for laser power under real-world conditions. Simultaneously, the stability and robustness of the reinforcement learning model are tested under different environmental disturbances (such as temperature changes). During the evaluation phase, model parameters can be appropriately adjusted based on the test results, including optimizing the reward function weight coefficients and the control strategy network structure, to further improve the model's performance and stability.
[0073] Finally, the optimized reinforcement learning model is deployed and applied. During online runtime, the control policy π... θ The model is fixed and based on the real-time acquired state space s k Calculate the action space a kThe model adjusts the drive signal in real time to maintain the laser power within the desired range. Simultaneously, it monitors the control results based on real-time feedback to address potential changes and disturbances, ensuring the stability of the laser power and atomic clock output frequency during long-term operation.
[0074] S3: The driving signal for the current control cycle is obtained based on the driving signal of the previous control cycle and the action space of the optimized reinforcement learning model.
[0075] Specifically, the driving signals for the current control cycle are obtained based on the driving signals from the previous control cycle and the action space of the optimized reinforcement learning model. These signals include:
[0076] u k =u k-1 +a k
[0077] Among them, u k This is the drive signal for the current control cycle.
[0078] S4: Controls the acousto-optic drive module based on the drive signal of the current control cycle to stabilize the laser power.
[0079] The driving signal is applied to the acousto-optic driving module. When the driving signal changes, the laser power through the acousto-optic driving module also changes. Ultimately, the driving signal is adjusted by frequency change feedback to maintain the stability of the laser power.
[0080] The above scheme achieves high-precision measurement of laser power through the optical frequency shift principle (AC Stark effect), and combines it with a reinforcement learning model to optimize power stability. The control strategy is continuously trained and adjusted under different environmental changes to ensure optimal control performance and stabilize laser power. Compared with traditional control methods, this scheme has stronger adaptive capabilities, not only able to handle nonlinear and time-varying systems, but also achieving highly robust control performance in dynamic environments.
[0081] Example 2
[0082] like Figure 2 As shown, this embodiment provides a laser power stabilization control system, including a laser power measurement system, a first processing module, a reinforcement learning module, and a second processing module, wherein:
[0083] The laser power measurement system is built based on the optical frequency shift effect and includes a laser, an acousto-optic drive module, an atomic clock, and a frequency measurement module. The acousto-optic drive module is used to adjust the laser power, and the frequency measurement module is used to measure the output frequency of the atomic clock.
[0084] The first processing module is used to calculate the frequency deviation of the current control cycle and the frequency change rate of the current control cycle relative to the previous control cycle. The frequency deviation of the current control cycle is the difference between the output frequency of the atomic clock in the current control cycle and the standard frequency. The frequency change rate of the current control cycle relative to the previous control cycle is the difference between the frequency deviation of the current control cycle and the frequency deviation of the previous control cycle.
[0085] The reinforcement learning module is used to deploy and optimize the trained reinforcement learning model. The state space of the reinforcement learning model includes the frequency deviation of the current control cycle, the rate of change of the frequency of the current control cycle relative to the previous control cycle, and the driving signal of the previous control cycle. The action space of the reinforcement learning model includes the increment of the driving signal of the current control cycle. The action space is obtained based on the control policy and the state space and is within a preset range. The reward function of the reinforcement learning model is used to minimize the frequency deviation and suppress rapid changes in frequency.
[0086] The second processing module is used to obtain the driving signal for the current control cycle based on the driving signal of the previous control cycle and the action space of the optimized reinforcement learning model. The driving signal of the current control cycle is used to control the acousto-optic driving module to stabilize the laser power.
[0087] Specifically, the state space of the reinforcement learning model includes the frequency deviation of the current control cycle, the rate of change of the frequency of the current control cycle relative to the previous control cycle, and the driving signal of the previous control cycle, including:
[0088] s k =[Δv k ,Δv k -Δv k-1 ,u k-1 ]
[0089] Among them, s k Let Δv be the state space of the current control cycle. k The frequency deviation of the current control cycle, Δv k-1 The frequency deviation of the previous control cycle, Δv k -Δv k-1 u is the rate of change of frequency of the current control cycle relative to the previous control cycle. k-1 This is the drive signal for the previous control cycle.
[0090] Specifically, the action space is obtained based on the control strategy and the state space, and includes the following within a preset range:
[0091] a k =π θ (s k )
[0092] Among them, a k For the action space of the current control cycle, ak ∈[a min ,a max ],π θ Here, θ represents the control strategy, and θ represents the strategy parameters.
[0093] Specifically, the reward function of a reinforcement learning model, used to minimize frequency bias and suppress rapid changes in frequency, includes:
[0094]
[0095] Where, r k The reward signal for the current control cycle is represented by α and β, which are weighting coefficients.
[0096] Specifically, the driving signals for the current control cycle are obtained based on the driving signals from the previous control cycle and the action space of the optimized reinforcement learning model. These signals include:
[0097] u k =u k-1 +a k
[0098] Among them, u k This is the drive signal for the current control cycle.
[0099] The system disclosed in the embodiments is described in a relatively simple manner because it corresponds to the method disclosed in the embodiments. For relevant details, please refer to the method section.
[0100] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A laser power stabilization control method, characterized in that, include: A laser power measurement system is built based on the optical frequency shift effect. The laser power measurement system includes a laser, an acousto-optic driving module, an atomic clock, and a frequency measurement module. The acousto-optic driving module is used to adjust the laser power, and the frequency measurement module is used to measure the output frequency of the atomic clock. The reinforcement learning model is trained and optimized. The state space of the reinforcement learning model includes the frequency deviation of the current control cycle, the frequency change rate of the current control cycle relative to the previous control cycle, and the driving signal of the previous control cycle. The frequency deviation of the current control cycle is the difference between the output frequency of the atomic clock in the current control cycle and the standard frequency. The frequency change rate of the current control cycle relative to the previous control cycle is the difference between the frequency deviation of the current control cycle and the frequency deviation of the previous control cycle. The action space of the reinforcement learning model includes the driving signal increment of the current control cycle. The action space is obtained based on the control policy and the state space and is within a preset range. The reward function of the reinforcement learning model is used to minimize frequency bias and suppress rapid changes in frequency; The driving signal for the current control cycle is obtained based on the driving signal of the previous control cycle and the action space of the optimized reinforcement learning model. The acousto-optic drive module is controlled based on the drive signal of the current control cycle to stabilize the laser power.
2. The method according to claim 1, characterized in that, The state space of the reinforcement learning model includes the frequency deviation of the current control cycle, the rate of change of the frequency of the current control cycle relative to the previous control cycle, and the driving signal of the previous control cycle, including: s k =[Δv k ,Δv k -Δv k-1 ,u k-1 ] Among them, s k Let Δv be the state space of the current control cycle. k The frequency deviation of the current control cycle, Δv k-1 The frequency deviation of the previous control cycle, Δv k -Δv k-1 u is the rate of change of frequency of the current control cycle relative to the previous control cycle. k-1 This refers to the drive signal of the previous control cycle.
3. The method according to claim 2, characterized in that, The action space is obtained based on the control strategy and the state space, and includes the following within a preset range: a k =π θ (s k ) Among them, a k For the action space of the current control cycle, a k ∈[a min ,a max ],π θ Let θ be the control strategy, and θ be the strategy parameter.
4. The method according to claim 2, characterized in that, The reward function of the reinforcement learning model, used to minimize frequency bias and suppress rapid frequency changes, includes: Where, r k The reward signal for the current control cycle is represented by α and β, which are weighting coefficients.
5. The method according to claim 3, characterized in that, The process of obtaining the driving signal for the current control cycle based on the driving signal of the previous control cycle and the action space of the optimized reinforcement learning model includes: ux=u k-1 +a k Among them, u k This is the drive signal for the current control cycle.
6. A laser power stabilization control system, characterized in that, It includes a laser power measurement system, a first processing module, a reinforcement learning module, and a second processing module, wherein: The laser power measurement system is built based on the optical frequency shift effect and includes a laser, an acousto-optic driving module, an atomic clock, and a frequency measurement module. The acousto-optic driving module is used to adjust the laser power, and the frequency measurement module is used to measure the output frequency of the atomic clock. The first processing module is used to calculate the frequency deviation of the current control cycle and the frequency change rate of the current control cycle relative to the previous control cycle. The frequency deviation of the current control cycle is the difference between the output frequency of the atomic clock in the current control cycle and the standard frequency. The frequency change rate of the current control cycle relative to the previous control cycle is the difference between the frequency deviation of the current control cycle and the frequency deviation of the previous control cycle. The reinforcement learning module is used to deploy the trained and optimized reinforcement learning model. The state space of the reinforcement learning model includes the frequency deviation of the current control cycle, the rate of change of the frequency of the current control cycle relative to the previous control cycle, and the driving signal of the previous control cycle. The action space of the reinforcement learning model includes the increment of the driving signal of the current control cycle. The action space is obtained based on the control policy and the state space and is within a preset range. The reward function of the reinforcement learning model is used to minimize the frequency deviation and suppress rapid changes in frequency. The second processing module is used to obtain the driving signal of the current control cycle based on the driving signal of the previous control cycle and the action space of the optimized reinforcement learning model. The driving signal of the current control cycle is used to control the acousto-optic driving module to stabilize the laser power.
7. The system according to claim 6, characterized in that, The state space of the reinforcement learning model includes the frequency deviation of the current control cycle, the rate of change of the frequency of the current control cycle relative to the previous control cycle, and the driving signal of the previous control cycle, including: s k =[Δv k ,Δv k -Δv k-1 ,u k-1 ] Among them, s k Let Δv be the state space of the current control cycle. k The frequency deviation of the current control cycle, Δv k-1 The frequency deviation of the previous control cycle, Δv k -Δv k-1 u is the rate of change of frequency of the current control cycle relative to the previous control cycle. k-1 This refers to the drive signal of the previous control cycle.
8. The system according to claim 7, characterized in that, The action space is obtained based on the control strategy and the state space, and includes the following within a preset range: a k =π θ (s k ) Among them, a k For the action space of the current control cycle, a k ∈[a min ,a max ],π θ Let θ be the control strategy, and θ be the strategy parameter.
9. The system according to claim 7, characterized in that, The reward function of the reinforcement learning model, used to minimize frequency bias and suppress rapid frequency changes, includes: Where, r k The reward signal for the current control cycle is represented by α and β, which are weighting coefficients.
10. The system according to claim 8, characterized in that, The process of obtaining the driving signal for the current control cycle based on the driving signal of the previous control cycle and the action space of the optimized reinforcement learning model includes: u k u k-1 +a k Among them, u k This is the drive signal for the current control cycle.