An adaptive laser frequency stabilization control method based on reinforcement learning algorithm

By introducing an adaptive laser frequency stability control method based on a reinforcement learning algorithm, the problems of difficult parameter adjustment and poor robustness of the traditional PID algorithm in laser frequency stability control are solved, and high-precision and adaptive frequency control is achieved, which is suitable for high-precision equipment such as cold atom gravimeters.

CN120085548BActive Publication Date: 2025-09-12JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510248242.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-09-12
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

Traditional PID feedback control algorithms have problems in controlling laser frequency stability, such as difficult parameter adjustment, poor robustness, and insufficient adaptability. They are unable to meet the high frequency stability requirements of high-precision equipment such as cold atom gravimeters.

Method used

An adaptive laser frequency stabilization control method based on reinforcement learning algorithm is adopted. Through interactive learning between the SAC agent and the environment, the frequency control strategy is automatically optimized, a specific reward function is designed to minimize the frequency error and limit the action amplitude, and a feedforward neural network model is combined to simulate the dynamic behavior of the laser frequency error.

Benefits of technology

It significantly improves the adaptability and robustness of laser frequency control, can achieve high-precision frequency stabilization in complex environments, adapt to the laser frequency adjustment needs under various environmental conditions, and meet the high-precision and real-time control requirements of precision instruments such as cold atom gravimeters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120085548B_ABST
    Figure CN120085548B_ABST
Patent Text Reader

Abstract

An adaptive laser frequency stabilization control method based on a reinforcement learning algorithm is disclosed, which addresses the shortcomings of traditional PID control in high-precision frequency regulation. This method introduces deep reinforcement learning based on the SAC algorithm. The intelligent agent automatically learns and optimizes the control strategy through continuous interaction with the environment, achieving high-precision and adaptive regulation of the laser frequency. This method takes the laser's frequency error as the core optimization objective and minimizes the frequency error by adjusting the two parameters of current and temperature. The action space is designed as a temperature adjustment stage and a current fine-tuning stage to avoid frequent large changes that adversely affect the system. The reward function takes into account penalties for frequency error, current and temperature adjustment amplitudes, and frequency change rate, aiming to balance control accuracy and system stability. This method is suitable for high-precision control requirements in fields such as cold atom gravimeters, quantum computing, and fiber-optic communications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a laser frequency stabilization control method, specifically a laser frequency stabilization control method based on a reinforcement learning algorithm. By introducing reinforcement learning technology, the present invention improves the dynamic response speed of laser frequency stabilization control. The present invention is applicable to cold atom gravimeters and other precision instruments with strict frequency control requirements. Background Art

[0002] In precision instruments such as cold atom gravimeters, laser frequency stability is a key technology that influences measurement accuracy and instrument performance. Traditional PID feedback control algorithms have been widely used for laser frequency stability control due to their simple structure and low implementation cost. PID algorithms achieve good control by adjusting parameters such as the laser's current and temperature to maintain the frequency close to a set reference value. However, the performance of this traditional algorithm relies on precise adjustment of control parameters, which typically require optimization based on human experience.

[0003] When faced with complex dynamic environments or nonlinear characteristics, PID algorithms have significant shortcomings, including difficulty in parameter adjustment, poor robustness, and insufficient adaptability. For example, in high-precision measurement equipment such as cold atom gravimeters, disturbances such as ambient temperature and mechanical vibration can significantly affect the frequency stability of the laser. PID algorithms struggle to quickly respond to these complex environmental changes, limiting the instrument's measurement accuracy and making it difficult to meet the higher frequency stability requirements of precision instruments.

[0004] In recent years, reinforcement learning (RL) algorithms have become a research hotspot in the control field due to their powerful adaptive learning capabilities. RL can effectively overcome the inherent shortcomings of traditional control algorithms by dynamically adjusting control strategies through interactive learning with the environment.

[0005] Frequency control methods based on reinforcement learning can provide efficient and precise control solutions in diverse and dynamic environments. Compared to PID algorithms, this approach does not rely on manual experience to adjust parameters, but instead achieves dynamic adaptation through algorithmic self-optimization. However, research on reinforcement learning algorithms in the field of laser frequency control is still in its infancy, especially for high-precision equipment such as cold atom gravimeters, where related technical research is still immature. Summary of the Invention

[0006] The present invention provides an adaptive laser frequency stabilization control method based on a reinforcement learning algorithm, aiming to solve the problems of limited frequency regulation accuracy, poor robustness to environmental disturbances, and complex parameter adjustment in the prior art.

[0007] An adaptive laser frequency stabilization control method based on reinforcement learning algorithm is implemented by the following steps:

[0008] Step 1: Record the laser current, temperature, and frequency error signal corresponding to the laser output in real time, and normalize the recorded current, temperature, and frequency error signals to obtain a normalized data set;

[0009] Step 2: construct a feedforward neural network model, and establish a relationship model between laser frequency, current and temperature through the feedforward neural network model to simulate the dynamic behavior of laser frequency error;

[0010] Step 3: Environment modeling and reward function design to achieve interactive updates between the SAC agent and the neural network model described in Step 2. The specific process is as follows:

[0011] Step 3.1: Set the state space to the current working state of the input laser and construct the state space s, which is expressed as:

[0012] s=[I n T n f e ]

[0013] Where, I n is the normalized current value; T n is the normalized temperature value; f e is the normalized frequency error;

[0014] Step 32: Continuously adjust the current and temperature through the SAC agent to minimize the frequency error f e , so that the laser frequency is stabilized at the target frequency f t Nearby, while satisfying the physical constraints:

[0015] ∣ΔI∣≤ΔI max

[0016] ∣ΔT∣≤ΔT max

[0017] Where, ΔI max and ΔT max are the maximum changes of current and temperature in each time step respectively;

[0018] Step 3. Set the action space as the output of the SAC agent, A = [ΔI, ΔT]. The SAC agent's goal is to minimize the frequency error by adjusting the laser's current and temperature. The action space is divided into a temperature adjustment phase and a current fine-tuning phase. The SAC agent selects the control variable according to the current state in each phase.

[0019] Step 3 and 4: Set the comprehensive reward function to minimize the frequency error and design the comprehensive reward function: Based on the comprehensive penalty term of normalized frequency error, current and temperature change, and frequency change rate, design the comprehensive reward function r, which can be expressed as follows:

[0020]

[0021] Where ΔT and ΔI are the temperature change and current change respectively. is the frequency change rate, α is the penalty factor for frequency error, β T and β I are the penalty factors for temperature change and current change, respectively, and γ is the penalty factor for frequency change rate;

[0022] Step 4: Use the neural network model described in Step 2 to simulate the real working environment of the laser, providing a realistic training environment for the SAC agent training; during the SAC agent training process, the goal is to maximize the average reward of the SAC agent in the process of interacting with the environment; in each round of training, the SAC agent samples multiple sets of data from the environment to update the SAC agent's strategy; and finally achieve real-time frequency control of the laser.

[0023] Beneficial effects of the present invention:

[0024] This invention offers significant innovation and technical advantages in laser frequency control methods. Traditional laser frequency adjustment methods rely primarily on manual experience or fixed PID algorithms, often lacking adaptability and unable to cope with complex, dynamically changing environments. Furthermore, traditional methods require manual parameter adjustment to optimize control, which is time-consuming and susceptible to human error, making it difficult to meet the high-precision and real-time control requirements of precision instruments such as cold atom gravimeters.

[0025] In contrast, this paper proposes an intelligent laser frequency control strategy by introducing a reinforcement learning algorithm, specifically a deep reinforcement learning method based on the SAC algorithm. The SAC agent, through continuous interaction and learning with the environment, automatically optimizes the frequency control strategy, eliminating the tedious process of manual parameter adjustment and significantly improving control efficiency and accuracy. This method demonstrates enhanced robustness and adaptability, particularly in complex, high-noise environments.

[0026] Furthermore, the present invention designs a specific reward function for laser frequency regulation, including penalty terms for frequency error, temperature variation, and current adjustment. This allows the SAC agent to optimize control accuracy while also balancing system stability and energy efficiency. This multi-objective optimization mechanism ensures that the laser frequency is stabilized near the target value with minimal error, while preventing system disruptions caused by frequent parameter changes.

[0027] The advantages of this method lie not only in its adaptability and high-precision control, but also in its strong generalization capabilities, adapting to the needs of laser frequency regulation under various environmental conditions. This innovative frequency control method provides new ideas for addressing the shortcomings of traditional methods and also provides a technical reference for other fields requiring high-precision control, such as cold atom gravimeters, quantum computing, and fiber-optic communications. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is a system block diagram of a laser frequency stabilization control method based on a reinforcement learning algorithm according to the present invention;

[0029] Figure 2 This is an action space phase switching diagram of a laser frequency stabilization control method based on a reinforcement learning algorithm according to the present invention;

[0030] Figure 3 This is a flow chart of the SAC intelligent agent environment training of the laser frequency stabilization control method based on the reinforcement learning algorithm described in the present invention. DETAILED DESCRIPTION

[0031] Combine Figures 1 to 3 This embodiment describes a reinforcement learning-based laser frequency stability control method. This method incorporates reinforcement learning algorithms, specifically deep reinforcement learning methods based on the SAC (Soft Actor-Critic) algorithm, to achieve adaptive, high-precision control of the laser frequency. Through continuous interaction with the environment, the SAC agent automatically learns and optimizes the control strategy to maximize frequency stability, eliminating the need for tedious manual parameter adjustment and overcoming the shortcomings of traditional PID control.

[0032] like Figure 1 As shown, Figure 1This is the experimental system structure of the adaptive laser frequency stabilization control method based on the reinforcement learning algorithm of this embodiment. The core of the system is the RL_agent controller, which adjusts the frequency output of the laser through the current control circuit and the temperature control circuit. The laser signal emitted by the DFB laser enters the modulation transfer spectrum optical path, which adjusts the transmission and absorption characteristics of the signal, ultimately affecting the stability of the laser frequency. The feedback information of the modulation transfer spectrum optical path is collected by the laser receiving circuit and enters the frequency feedback data path. The frequency data is transmitted to the computer platform. After receiving this data, the computer platform processes the signal through the data acquisition and recording module and communicates with the RL_agent controller through the communication module. Based on the reinforcement learning algorithm, the RL_agent controller continuously adjusts the control strategy according to the frequency feedback data and system status, thereby optimizing the frequency stability of the laser. The output signal of the control strategy is used to fine-tune the operating state of the laser through the current and temperature control circuits to ensure accurate and stable frequency. At the same time, the acquisition card (based on the FPGA main control chip) is responsible for real-time data extraction and transmission, supporting real-time data processing and control signal feedback of the entire system.

[0033] The specific implementation process of the control method of this embodiment is as follows:

[0034] Step 1: Use the laser and laser driver together and run the system in free-running mode. Use MTS (Modulation Transfer Spectroscopy) technology to collect the frequency error signal Δf of the output laser in real time.

[0035] In this embodiment, in order to achieve high-precision frequency control, the target frequency of the laser is set to f t , the actual frequency is f, and the frequency error signal Δf is defined as:

[0036] Δf=|ff t ∣

[0037] The data acquisition frequency was set to 1 kHz and the sampling duration was several hours to obtain sufficient sample size for subsequent analysis and modeling.

[0038] Step 2: Generate a random variation sequence within the physically permissible range of the laser current and temperature. Simultaneously, through a reasonable control strategy, limit the amplitude of current and temperature changes to avoid large parameter jumps that could lead to system instability. The random variation sequence is used to stimulate the laser into different operating states, and the corresponding current and temperature values ​​are recorded. Furthermore, the frequency error signal Δf corresponding to the laser output is synchronously recorded to form a complete data set.

[0039] In this implementation, the stability of the laser frequency depends primarily on the laser's internal operating environment, with current and temperature being key factors influencing frequency variations. Fluctuations in operating current and temperature can cause subtle shifts in the laser's output wavelength, leading to variations in the laser frequency. Therefore, in frequency stabilization control, the SAC agent must precisely adjust both current and temperature to ensure the laser's output frequency remains close to the target value.

[0040] The relationship between the laser's output wavelength λ and the laser's internal current I and temperature T is approximately a linear model. This relationship can be expressed as:

[0041] λ=λ0+k I (I-I0)+k T (T-T0)

[0042] Where: λ0 is the initial wavelength of the laser; I0 and T0 are the reference current and reference temperature respectively; k I and k T They represent the adjustment coefficients of current and temperature on wavelength respectively.

[0043] The relationship between the actual frequency f and the current and temperature can be derived using the relationship formula between the laser frequency and wavelength:

[0044]

[0045] Step 3: Process the collected data and normalize them to ensure that the input and output values ​​are within a unified range to form a data set;

[0046] In practical applications, the actual frequency f will produce errors due to changes in current and temperature. Therefore, we can collect the frequency deviation pairs corresponding to the temperature and current in the actual laser, and process and normalize the collected data to ensure that the input and output values ​​are within the same range. The specific process is as follows:

[0047] Step 3-1: Perform preliminary filtering on the data to remove abnormal data caused by system noise and external environment fluctuations;

[0048] Step 3-2: To quantify the impact of frequency error on the SAC agent, the frequency error is normalized and defined as the normalized frequency error: Normalized frequency error f e The smaller it is, the closer the laser frequency is to the target frequency. The goal of the SAC agent is to minimize the normalized frequency error f e The agent's goal is to minimize the normalized frequency error and thus achieve stable frequency control.

[0049] According to the above formula, the specific control target of the agent can be expressed as:

[0050]

[0051] Step 3-3: Perform maximum and minimum normalization on the current and temperature data to ensure that the input features are on the same scale and avoid the impact of different dimensions on model training.

[0052] The normalized current value is:

[0053] The normalized temperature value is:

[0054] Where, I n is the normalized current value, ranging from [0,1]; T n is the normalized temperature value, ranging from [0,1], to form a data pair: {(I n ,T n )→f e}.

[0055] In this embodiment, the optimization goal of the control strategy is to continuously adjust the value of the current I and the temperature T so that the normalized frequency error f e Gradually reduce the frequency, ultimately stabilizing it near the target frequency. In real-world applications, current and temperature regulation are subject to physical constraints. Therefore, the proxy control strategy needs to be optimized within these physical constraints to minimize the frequency error without exceeding the allowable range of the real system.

[0056] Step 4: Construct a feedforward neural network (FNN) model to simulate the dynamic behavior of the laser frequency error.

[0057] The network model structure is designed as follows: input layer, hidden layer, activation layer and output layer; the input layer contains I n and T n The hidden layer uses three fully connected layers, each containing 128 neurons to balance complexity and accuracy. The activation function is ReLU. The output layer is the predicted normalized frequency error, and its loss function is defined as the mean squared error (MSE). The neural network model is trained using the PyTorch framework, with the data divided into an 80% training set, a 10% validation set, and a 10% test set. Finally, the generated neural network model is evaluated using the mean absolute error and maximum absolute error. The specific implementation process is as follows:

[0058] Step 4-1: Introduce the feedforward neural network to build the real laser model. Its input layer includes I n and T n ; The hidden layer uses three layers of full connection; the output layer is the predicted normalized frequency error Its loss function is defined as mean square error MSE:

[0059]

[0060] in: is the true normalized frequency error; The loss function optimizes the neural network model parameters by minimizing the frequency error, enabling it to accurately predict the frequency error.

[0061] Step 4-2: Use the PyTorch framework to train a neural network model to approximate the nonlinear function output of a real laser under free operation, creating a realistic training environment for the SAC agent. The obtained dataset is divided into an 80% training set, a 10% validation set, and a 10% test set to effectively evaluate the training effect and generalization ability of the neural network model.

[0062] In this implementation, the Adam optimizer is used for training with a learning rate of 0.001. The backpropagation algorithm is used during training to update network parameters by minimizing the loss function. A training cycle is set, and training is terminated early if the loss function on the validation set stabilizes to avoid overfitting.

[0063] Step 4-3: Use the test set to evaluate the trained neural network model and calculate the mean absolute error and maximum absolute error. These indicators reflect the accuracy and robustness of the neural network model. The evaluation results will determine whether to enter the next step of reinforcement learning training.

[0064] In this implementation, the trained feedforward neural network model accurately predicts the normalized frequency error of the laser under given current and temperature conditions, providing accurate frequency error information to the SAC agent, enabling it to adjust current and temperature parameters to minimize the frequency error during control. This feedforward neural network model effectively simulates the complex nonlinear behavior of the laser under varying current and temperature conditions, providing a realistic training environment for SAC agent training.

[0065] Step 5: Implement interactive updates between the SAC agent and the neural network model described in step 4 through environment modeling and reward function design.

[0066] The environment modeling includes setting the state space and action space. The state space is set as the current working state of the input laser, the action space is set as the output of the SAC agent, and the reward function is set to minimize the frequency error while limiting the action amplitude. The specific implementation process is as follows:

[0067] Step 5-1. The state space is defined as a multidimensional space, where each dimension corresponds to an important characteristic of the system, including the current and temperature information of the laser, and also taking into account the current frequency error and frequency change rate, as follows:

[0068] s=[I n T n f e ]

[0069] Among them: I n is the normalized current value, ranging from [0,1]; T n is the normalized temperature value, ranging from [0,1]; f e is the normalized frequency error, ranging from -1 to 1.

[0070] According to the above formula, the state vector of the current system can be obtained, which provides the SAC agent with sufficient environmental information to make the next decision.

[0071] Step 5-2: Set the current and temperature changes controlled by the agent. The maximum range of change is determined by the actual system parameters, and we get:

[0072] ∣ΔI∣≤ΔI max

[0073] ∣ΔT∣≤ΔT max

[0074] Where: ΔI max and ΔT max are the maximum changes in current and temperature in each time step, respectively.

[0075] In actual control, the agent will select the appropriate adjustment amplitude based on the current current and temperature conditions to ensure that the frequency error is minimized while meeting the physical constraints.

[0076] Step 5-3. Define the action space for the output of the SAC agent as: A = [ΔI, ΔT]. The SAC agent's goal is to minimize the frequency error by adjusting the laser's current and temperature. To ensure system stability, the action space is divided into two stages: temperature adjustment and current fine-tuning. The design is as follows:

[0077] Phase 1: Temperature adjustment phase;

[0078] The SAC agent can only adjust the temperature of the laser. When the normalized frequency error f e When the temperature is greater than or equal to the conversion threshold, the SAC intelligent body will enter the temperature adjustment phase. The temperature adjustment amount ΔT ranges from -0.2℃ to 0.2℃ to ensure that the temperature change is within a reasonable range and avoid laser system instability and mode hopping caused by excessive temperature changes.

[0079] Phase 2: Current fine-tuning phase;

[0080] The SAC agent can only adjust the laser current. When the normalized frequency error f e When the current is less than the conversion threshold, the SAC intelligent body will enter the current adjustment stage. The current adjustment amount ΔI ranges from -1.0mA to 1.0mA to ensure that the current change does not exceed the maximum current change limit.

[0081] In this embodiment, the conversion threshold is set to 0.1. The conversion threshold is determined experimentally based on the laser frequency error characteristics or optimized during training using validation data. Specifically, when the normalized frequency error is less than 0.1, it indicates that the frequency is close to the target value and it is suitable to enter the current fine-tuning stage.

[0082] To ensure the versatility of the SAC agent model, the action space is adjusted using a reinforcement learning algorithm. For each stage, the action space is continuous and automatically optimized using the action selection strategy.

[0083] Step 5-4: Set the comprehensive reward function to ensure that the SAC agent can adjust according to the frequency error, current and temperature changes, and frequency change rate, thereby maximizing frequency stability and reducing unnecessary energy consumption.

[0084] The reward function is the core component of this implementation, determining the behavior of the SAC agent during its interaction with the environment. A well-designed reward function can effectively guide the SAC agent toward its target state.

[0085] In this implementation, the goal is to minimize the deviation in laser frequency, i.e., the frequency error, and this is achieved by adjusting the current and temperature. The design takes the following aspects into consideration:

[0086] Normalized frequency error f e is the deviation between the current frequency and the target frequency. The smaller the error, the greater the reward should be. Its form is:

[0087] r1=-αf e

[0088] Where: α is the penalty factor of the frequency error. By adjusting the value of α, the frequency error can be finely controlled.

[0089] To avoid system instability caused by frequent large temperature or current changes, a penalty term for temperature and current changes is introduced into the reward function, which is in the form of:

[0090] r2=-β T |ΔT|-βI ∣ΔI∣

[0091] Where: β T and β I are the penalty factors for temperature variation and current variation, respectively. By adjusting these penalty factors, the SAC agent can balance the trade-off between frequency accuracy and temperature and current variations.

[0092] Frequency change rate It reflects the stability of the frequency. If the frequency changes too fast, it means that the SAC agent is unstable during the dynamic adjustment process, so the frequency change rate should be penalized. Its form is:

[0093]

[0094] Where: γ is the penalty factor for the frequency change rate, and a larger value is usually selected to ensure that the frequency is as stable as possible.

[0095] Combining the above penalty terms and reward terms, the comprehensive reward function is expressed as:

[0096]

[0097] Among them, α=1.0, β=0.5, γ=0.5, and δ=0.2. Through optimization, the frequency error is preferentially reduced while limiting rapid changes. The specific parameter values ​​can be adjusted according to the specific application.

[0098] Step 6: Use the neural network model trained in Step 4 to approximate the real laser environment. This neural network model accurately reflects the actual dynamic behavior of the laser in the simulated environment, such as the relationship between temperature, current, and frequency, and their impact on laser frequency stability. This neural network model serves as the basis for the SAC agent's interaction with the environment and provides learning signals for the SAC agent.

[0099] The SAC agent updates its strategy by interacting with the neural network model and uses the Optuna automatic parameter adjustment tool to optimize key hyperparameters (optimizing the learning rate (10 -4 to 10 -2), discount factor (0.9 to 0.99), with the goal of maximizing the average reward of 1000 training sessions. The objective function of automatic parameter adjustment is set to the average reward of the SAC agent during its interaction with the environment to maximize the overall performance of the SAC agent. In each round of training, the SAC agent will sample multiple sets of data from the environment. These data include the control actions selected by the SAC agent, the reward signals fed back by the environment, and the new state information. These data are used to update the SAC agent's strategy (that is, the rule or decision function by which the SAC agent selects the optimal action based on the reward signal under different states, and by optimizing the strategy, the SAC agent can maximize long-term rewards in multiple interactions), and optimize the SAC agent network parameters (parameters of the policy network and the value function network). The optimized parameters will be passed to the next round of training to improve the performance of the SAC agent.

[0100] In this implementation, the Optuna library is used during the SAC agent training process to optimize hyperparameters. This optimizes the hyperparameter combination to enhance the SAC agent's performance and ultimately achieve efficient SAC agent control. After training, the SAC agent's performance is evaluated to determine the optimal control strategy. The SAC agent is then saved and deployed in a practical laser frequency control system.

[0101] This embodiment also includes step 7, verifying the control performance of the SAC agent after training. The verification content includes the SAC agent's ability to suppress the laser frequency error during the control process and its robustness in complex environments. The verification index is set to the normalized frequency error f e The mean and standard deviation of are used to evaluate the control accuracy and stability of the SAC agent;

[0102] Convert and save the trained SAC agent to a lightweight format (such as ONNX) for efficient deployment. Then, deploy the SAC agent to the target hardware environment to achieve real-time frequency control of the laser. During deployment, ensure that high-level hard limits are set to trigger an emergency stop if exceeded. This step ensures that the SAC agent can run efficiently in the embedded system to meet the real-time, safety, and performance requirements of real-world applications.

[0103] The control method described in this embodiment introduces deep reinforcement learning based on the SAC algorithm. The SAC agent automatically learns and optimizes the control strategy through continuous interaction with the environment, thereby achieving high-precision and adaptive adjustment of the laser frequency. In the control method, the frequency error of the laser is used as the core optimization target, and the frequency error is minimized by adjusting the two key parameters of current and temperature. To ensure the stability of the system, the action space is designed as a temperature adjustment stage and a current fine-tuning stage to avoid frequent large changes from adversely affecting the system. The reward function takes into account the penalties for frequency error, current and temperature adjustment amplitude, and frequency change rate, aiming to balance control accuracy and system stability. Through the deep reinforcement learning algorithm, combined with hyperparameter optimization, the SAC agent can efficiently learn in a complex environment and ultimately achieve precise frequency control.

[0104] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0105] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.

Claims

1. An adaptive laser frequency stabilization control method based on reinforcement learning algorithm, characterized by: The method is implemented by the following steps: Step 1: Record the laser current, temperature, and frequency error signal corresponding to the laser output in real time, and normalize the recorded current, temperature, and frequency error signals to obtain a normalized data set; Step 2: construct a feedforward neural network model, and establish a relationship model between laser frequency, current and temperature through the feedforward neural network model to simulate the dynamic behavior of laser frequency error; Step 3: Environment modeling and reward function design to achieve interactive updates between the SAC agent and the neural network model described in Step 2. The specific process is as follows: Step 3.1: Set the state space to the current working state of the input laser and construct the state space s, which is expressed as: s=[I n T n f e ] Where, I n is the normalized current value; T n is the normalized temperature value; f e is the normalized frequency error; Step 32: Continuously adjust the current and temperature through the SAC agent to minimize the frequency error f e , so that the laser frequency is stabilized at the target frequency f t Nearby, while satisfying the physical constraints: ∣ΔI∣≤ΔI max ∣ΔT∣≤ΔT max Where, ΔI max and ΔT max are the maximum changes of current and temperature in each time step respectively; Step 3. Set the action space as the output of the SAC agent, A = [ΔI, ΔT]. The SAC agent's goal is to minimize the frequency error by adjusting the laser's current and temperature. The action space is divided into a temperature adjustment phase and a current fine-tuning phase. The SAC agent selects the control variable according to the current state in each phase. Step 3 and 4: Set the comprehensive reward function to minimize the frequency error and design the comprehensive reward function: Based on the comprehensive penalty term of normalized frequency error, current and temperature change, and frequency change rate, design the comprehensive reward function r, which can be expressed as follows: Where ΔT and ΔI are the temperature change and current change respectively. is the frequency change rate, α is the penalty factor for frequency error, β T and β I are the penalty factors for temperature change and current change, respectively, and γ is the penalty factor for frequency change rate; Step 4: Use the neural network model described in Step 2 to simulate the real working environment of the laser, providing a realistic training environment for the SAC agent training; during the SAC agent training process, the goal is to maximize the average reward of the SAC agent in the process of interacting with the environment; in each round of training, the SAC agent samples multiple sets of data from the environment to update the SAC agent's strategy, and ultimately achieve real-time frequency control of the laser.

2. The adaptive laser frequency stabilization control method based on reinforcement learning algorithm according to claim 1 is characterized in that: In step 1, set the target frequency of the laser to f t , the actual frequency is f, then the frequency error signal Δf is: Δf=∣f-f t ∣ Normalized frequency error f e for: Among them, the normalized frequency error f e The smaller it is, the closer the laser frequency is to the target frequency. The goal of the SAC agent control strategy is to minimize the frequency error f e .

3. The adaptive laser frequency stabilization control method based on reinforcement learning algorithm according to claim 1, characterized in that: The specific process of step 2 is: Step 21: The constructed feedforward neural network model includes an input layer, a hidden layer, an activation layer and an output layer; The normalized current value I obtained from the normalized data set in step 1 is normalized to n and the normalized temperature value T n As input to the input layer of the feedforward neural network model; The hidden layer of the feedforward neural network model is a multi-layer fully connected layer, the activation layer uses the ReLU activation function, and the output layer is the predicted normalized frequency error; Step 22: Using the PyTorch framework to train the feedforward neural network model; The dataset described in step 1 is divided into a training set, a validation set, and a test set. The backpropagation algorithm is used on the training set to update the neural network model parameters by minimizing the loss function. If the loss function on the validation set is stable, the training ends. The trained neural network model is evaluated on the test set to build a realistic training environment for the SAC agent.

4. The adaptive laser frequency stabilization control method based on reinforcement learning algorithm according to claim 1, characterized in that: In step 2, the loss function is defined as the mean square error MSE: in: is the true normalized frequency error value; is the predicted normalized frequency error value; the loss function optimizes the neural network model parameters by minimizing the frequency error so that it can accurately predict the frequency error.

5. The adaptive laser frequency stabilization control method based on reinforcement learning algorithm according to claim 1, characterized in that: In step 33, the temperature adjustment stage and the current fine-tuning stage are respectively: Temperature adjustment stage: When the normalized frequency error f e When the value is greater than or equal to the conversion threshold, the SAC agent enters the temperature adjustment phase. The SAC agent only adjusts the temperature of the laser by ΔT, ranging from -0.2°C to 0.2°C. Current fine-tuning stage: When the normalized frequency error f e When it is less than the conversion threshold, the SAC agent will enter the current adjustment stage. The SAC agent only adjusts the current of the laser with an adjustment amount of ΔI, ranging from -1.0mA to 1.0mA.

6. The adaptive laser frequency stabilization control method based on reinforcement learning algorithm according to claim 5, characterized in that: The conversion threshold is set to 0.1, which is determined based on the frequency error characteristics of the laser.

7. The adaptive laser frequency stabilization control method based on reinforcement learning algorithm according to claim 1, characterized in that: In step 4, during the SAC agent training process, the Optuna library is introduced to perform hyperparameter optimization to achieve SAC agent control; After the training is completed, the performance of the SAC agent is evaluated to determine the optimal control strategy, and the SAC agent is finally saved.

8. The adaptive laser frequency stabilization control method based on reinforcement learning algorithm according to claim 1, characterized in that: Step 4 also includes saving the trained SAC agent in a lightweight format and deploying it to the target hardware environment.

Citation Information

Patent Citations

  • Spectral programmable optical frequency comb generation method based on deep reinforcement learning

    CN117192865A

  • Motor adaptive control method and system based on deep reinforcement learning

    CN118508817A