Ultrafast pulse laser reverse implementation method and system based on reinforcement learning
Through the reinforcement learning method, the reverse design of ultrafast pulse lasers is performed using LSTM and DDPG algorithms, which solves the problem of inefficient reverse design in the prior art and achieves fast and effective laser parameter derivation.
Patent Information
- Application Number
- CN202210415895.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-20
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-04-20
AI Technical Summary
The prior art is difficult to quickly and efficiently perform the reverse design of ultrafast pulse lasers, and the design depends on experience and experiments and is inefficient.
Using reinforcement learning-based method, the traditional distributed Fourier algorithm is replaced by long and short-term memory network model (LSTM), combined with neural network rapid modeling, and reverse design of the laser using reinforcement learning algorithm (DDPG) to deduce laser parameters that meet the output needs.
The rapid reverse design of ultrafast pulse laser is realized, which improves simulation efficiency, reduces the computational complexity, and makes up for the shortcomings of relying on experience and experiments in traditional design methods.
Smart Images

Figure CN114818488B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ultrafast pulse lasers, and in particular to a method and system for realizing ultrafast pulse laser reverse direction based on reinforcement learning. Background Art
[0002] A pulse laser refers to a laser with a single laser pulse width of less than 0.25 seconds and which works only once at a certain interval. It has a large output power and is suitable for laser marking, cutting, ranging, etc.
[0003] Patent document CN113922199A (application number: CN202111141060.1) discloses an anti-return master oscillator power amplified pulse laser, which includes a seed semiconductor laser, a return light processor, a first-level amplifier, a first-level return light processor, a second-level amplifier, a second-level return light processor, ..., an n-level amplifier, and an n-level return light processor; the return light processor is designed to ensure forward transmission of signal light, prevent reverse transmission of signal light, prevent reverse transmission of nonlinear laser, ASE light, and return light outside the optical path, and monitor the return light.
[0004] The existing technology mainly simulates the transmission process of ultrafast lasers through the iterative distributed Fourier traditional algorithm. The algorithm is highly complex and inefficient, and it is difficult to meet the requirements of providing feedback information for rapid reverse design of ultrafast lasers. On the other hand, there is still no technology that uses accurate simulation modeling results to reverse design ultrafast lasers. The design of ultrafast lasers relies almost entirely on experience and experimental attempts. Summary of the invention
[0005] In view of the defects in the prior art, the object of the present invention is to provide a method and system for reverse implementation of ultrafast pulse laser based on reinforcement learning.
[0006] The method for realizing reverse ultrafast pulse laser based on reinforcement learning provided by the present invention includes:
[0007] Step 1: Obtain a seed source through mode locking technology, and determine the pulse shape and spot mode through the seed source;
[0008] Step 2: Use fiber amplification technology to amplify the power of the pulse and determine the final energy and output characteristics of the pulse;
[0009] Step 3: Modularize and integrate the ultrafast laser generation process, and simulate the pulse propagation process in the optical fiber through the distributed Fourier algorithm;
[0010] Step 4: Input the simulation data into the long short-term memory network model LSTM according to the preset window size value, and perform step-by-step training;
[0011] Step 5: Input the training results into the reinforcement learning DDPG to obtain the optimal parameters of the laser so that the laser can quickly output the specified target pulse state.
[0012] Preferably, the action that obtains the maximum value is found through the action estimation network, the action reality network, the state reality network, and the state estimation network, the output pulse characteristics of the seed source in the output laser, the dispersion coefficient of the broadened grating, the gain coefficient and length of the gain fiber.
[0013] Preferably, the pulse output predicted by the model LSTM is used as the input of the reinforcement learning DDPG, different laser parameters are selected as actions, and the Q value is estimated through reinforcement learning. The Q value is the sum of the value of evaluating the current action and the reward value of estimating future actions. Then, the mean square error between the Q value under ideal pulse conditions and the Q value under the current state and action is calculated, and the action that minimizes the mean square error between the simulated output pulse and the target pulse is searched as the optimal laser parameter.
[0014] Preferably, in the simulation, the distance that the optical pulse is transmitted in the optical fiber is set to h, and first the pulse is only subjected to the effect of nonlinear effects, with dispersion and loss being zero; then the nonlinear effect is set to zero, and only the effects of dispersion and loss are considered.
[0015] Preferably, the transmitted pulse amplitude is expressed as:
[0016]
[0017] Where A(z, T) represents the pulse amplitude in the z direction within period T; D represents dispersion; N represents nonlinearity; and the linear operator The calculation formula in the frequency domain is:
[0018]
[0019] Among them, F -1 represents the inverse Fourier transform, ω is the angular frequency, is the complex amplitude.
[0020] The ultrafast pulse laser reverse realization system based on reinforcement learning provided by the present invention includes:
[0021] Module M1: Obtain the seed source through mode locking technology, and determine the pulse shape and spot mode through the seed source;
[0022] Module M2: Use fiber amplification technology to amplify the power of the pulse and determine the final energy and output characteristics of the pulse;
[0023] Module M3: modularizes and integrates the ultrafast laser generation process, and simulates the pulse propagation process in the optical fiber through the distributed Fourier algorithm;
[0024] Module M4: Input the simulation data into the long short-term memory network model LSTM according to the preset window size value, and perform step-by-step training;
[0025] Module M5: Input the training results into the reinforcement learning DDPG to obtain the optimal parameters of the laser so that the laser can quickly output the specified target pulse state.
[0026] Preferably, the action that obtains the maximum value is found through the action estimation network, the action reality network, the state reality network, and the state estimation network, the output pulse characteristics of the seed source in the output laser, the dispersion coefficient of the broadened grating, the gain coefficient and length of the gain fiber.
[0027] Preferably, the pulse output predicted by the model LSTM is used as the input of the reinforcement learning DDPG, different laser parameters are selected as actions, and the Q value is estimated through reinforcement learning. The Q value is the sum of the value of evaluating the current action and the reward value of estimating future actions. Then, the mean square error between the Q value under ideal pulse conditions and the Q value under the current state and action is calculated, and the action that minimizes the mean square error between the simulated output pulse and the target pulse is searched as the optimal laser parameter.
[0028] Preferably, in the simulation, the distance that the optical pulse is transmitted in the optical fiber is set to h, and first the pulse is only subjected to the effect of nonlinear effects, with dispersion and loss being zero; then the nonlinear effect is set to zero, and only the effects of dispersion and loss are considered.
[0029] Preferably, the transmitted pulse amplitude is expressed as:
[0030]
[0031] Where A(z, T) represents the pulse amplitude in the z direction within period T; D represents dispersion; N represents nonlinearity; and the linear operator The calculation formula in the frequency domain is:
[0032]
[0033] Among them, F -1 represents the inverse Fourier transform, ω is the angular frequency, is the complex amplitude.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] (1) The present invention proposes a method for rapid inverse design of an ultrafast pulsed laser based on reinforcement learning. This method first replaces the inefficient traditional split-step Fourier algorithm with a long short-term memory network model (LSTM) to improve the simulation efficiency. Then, according to the output requirements, combined with rapid neural network modeling, the ultrafast laser is inversely designed through reinforcement learning to deduce the laser parameters that meet the output requirements, providing an intelligent reference method for building a real ultrafast pulsed laser system.
[0036] (2) Model the ultrafast pulsed laser through artificial intelligence algorithm modeling to replace the traditional split-step Fourier algorithm, reduce the computational complexity, and improve the simulation efficiency; inversely design the ultrafast pulsed laser through reinforcement learning.
[0037] (3) The present invention uses LSTM neural network modeling to replace the traditional algorithm, shortening the simulation time-consuming and improving the simulation efficiency; adopts reinforcement learning to intelligently inversely design the ultrafast pulsed laser, making up for the defect that the design could only rely on experience and a large number of experiments in the past. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] By reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:
[0039] Figure 1 It is the industrial production flow chart of the ultrafast pulsed laser;
[0040] Figure 2 It is the flow chart of ultrafast pulse amplification;
[0041] Figure 3 It is the schematic diagram of inverse design of the ultrafast pulsed laser based on reinforcement learning;
[0042] Figure 4 It is the schematic diagram of DDPG. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all fall within the protection scope of the present invention.
[0044] Embodiment:
[0045] The output of the ultrafast pulsed laser has an ultrashort time width (which can be as narrow as the femtosecond level), an ultra-high peak power, and an ultra-wide spectral range. Its industrial production process can be roughly divided into three parts: seed source, amplifier, and integration, as Figure 1Among them, the ultrafast laser seed source is generally generated by mode locking technology, which determines the ultrashort pulse shape, spot mode, etc.; the amplifier often uses fiber amplification technology, which determines the final energy and output characteristics of the ultrashort pulse; integration modularizes the ultrafast laser generation process to facilitate industrial production and use.
[0046] In the numerical simulation process of the amplification part, the distributed Fourier algorithm is often used to solve the nonlinear Schrödinger equation satisfied in the pulse propagation process, and the Runge-Kutta algorithm is used to solve the rate equation satisfied in the gain process. Taking the chirped pulse amplification process as an example, Figure 2 The process of ultrafast pulse amplification is demonstrated.
[0047] The traditional iterative distributed Fourier algorithm can model ultrafast pulse lasers more accurately. However, the traditional iterative distributed Fourier algorithm is highly complex and takes a long time to simulate. On the other hand, the design of ultrafast pulse lasers relies almost entirely on experiments. If you want to obtain a pulse output with specified characteristics, you need to fine-tune the parameters of each module of the laser through repeated experiments, which is a huge workload and requires a high level of experience from the experimenter.
[0048] The ultrafast pulse laser reverse design process based on reinforcement learning proposed in the present invention is as follows: Figure 3 As shown in the figure, first, a large amount of simulation data is generated through traditional algorithms, which mainly involves two parts: 1) using the distributed Fourier algorithm to simulate the propagation process of the pulse in the optical fiber; 2) using the fourth-order Runge-Kutta algorithm to solve the rate equation satisfied by the pulse in the gain fiber. Then, the simulation data is input into the neural network LSTM according to the set window size value, and the long short-term memory network model (LSTM) is used for step-by-step training. Compared with the traditional distributed Fourier algorithm, the trained LSTM neural network can greatly shorten the simulation time and improve the simulation efficiency.
[0049] The results of neural network training are input into the reinforcement learning DDPG. The action with the greatest value is found through the action estimation network, action reality network, state reality network and state estimation network, so as to output the ideal parameters of each module in the laser (such as the output pulse characteristics of the seed source, the dispersion coefficient of the broadening grating, the gain coefficient and length of the gain fiber, etc.). Combined with LSTM neural network modeling, the time complexity of the simulation is reduced, and the changes in the pulse light field, time domain and frequency domain of the laser under different laser parameters are calculated in a shorter time. The calculation results are fed back to the reinforcement learning algorithm DPPG, so that the laser can quickly output the specified target pulse state, realize the rapid reverse design of ultrafast lasers, and provide guidance and reference for the construction of real ultrafast laser systems.
[0050] The simulation process is:
[0051] Set the distance that the optical pulse travels in the optical fiber to h. First, let the pulse experience only the nonlinear effect, and the dispersion and loss are zero; then set the nonlinear effect to zero, and only consider the effects of dispersion and loss;
[0052] The transmitted pulse amplitude is expressed as:
[0053]
[0054] Where D stands for dispersion, N stands for nonlinearity, and the linear operator It can be calculated in the frequency domain:
[0055]
[0056] Among them, F -1 represents the inverse Fourier transform, w is the angular frequency, and A is the complex amplitude.
[0057] The fourth-order Runge-Kutta algorithm is used to solve the rate equation satisfied by the pulse in the gain fiber:
[0058]
[0059]
[0060] Among them, h is the solution step length, k 1 , k 2 , k 3 , k 4 is the coefficient.
[0061] The simulation data is input into the long short-term memory network model LSTM according to the preset window size value, and trained step by step; assuming the step size is 10, data No. 0 to 9 in the data set are selected as the input of LSTM and then the 10th data is predicted, and then the 11th data is predicted using 1 to 10, and so on.
[0062] like Figure 4 , combined with LSTM neural network modeling, the time complexity of the simulation is reduced, and the changes in the pulse light field, time domain and frequency domain of the laser under different laser parameters are calculated in a shorter time; the calculation results are fed back to the reinforcement learning algorithm DDPG, the current state is input into the Actor network to select action A, and the state and action are input into the Critic network together to calculate the current Q value. At the same time, the target state is input into the Actor target network, and the obtained action and state are input into the Critic target network together to obtain the target Q value, and the error between the two is calculated. The goal is to minimize the error, so that the laser can quickly output the specified target pulse state, realize the rapid reverse design of ultrafast lasers, and provide guidance and reference for the construction of real ultrafast laser systems.
[0063] Those skilled in the art know that, in addition to implementing the system, device and its various modules provided by the present invention in a purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. Therefore, the system, device and its various modules provided by the present invention can be considered as a hardware component, and the modules included therein for implementing various programs can also be considered as structures within the hardware component; the modules for implementing various functions can also be considered as both software programs for implementing the method and structures within the hardware component.
[0064] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A method for reverse realization of ultrafast pulsed laser based on reinforcement learning, It is characterized in that include: Step 1: Obtain a seed source through mode locking technology, and determine the pulse shape and spot mode through the seed source; Step 2: Use fiber amplification technology to amplify the power of the pulse and determine the final energy and output characteristics of the pulse; Step 3: Modularize and integrate the ultrafast laser generation process, and simulate the pulse propagation process in the optical fiber through the distributed Fourier algorithm; Step 4: Input the simulation data into the long short-term memory network model LSTM according to the preset window size value, and perform step-by-step training; Step 5: Input the training results into the reinforcement learning DDPG to obtain the optimal parameters of the laser so that the laser can quickly output the specified target pulse state.
2. The method for realizing reverse ultrafast pulse laser based on reinforcement learning according to claim 1, It is characterized in that The action with the maximum value is found through the action estimation network, the action reality network, the state reality network and the state estimation network, the output pulse characteristics of the seed source in the output laser, the dispersion coefficient of the broadened grating, the gain coefficient and the length of the gain fiber.
3. The method for realizing reverse ultrafast pulse laser based on reinforcement learning according to claim 1, It is characterized in that The pulse output predicted by the model LSTM is used as the input of the reinforcement learning DDPG. Different laser parameters are selected as actions. The Q value is estimated through reinforcement learning. The Q value is the sum of the value of evaluating the current action and the reward value of estimating future actions. Then, the mean square error between the Q value under ideal pulse conditions and the Q value under the current state and action is calculated. The action that minimizes the mean square error between the simulated output pulse and the target pulse is searched as the optimal laser parameter.
4. The method for realizing reverse ultrafast pulse laser based on reinforcement learning according to claim 1, It is characterized in that In the simulation, the distance that the optical pulse is transmitted in the optical fiber is set to h. First, the pulse is only subjected to the nonlinear effect, and the dispersion and loss are zero. Then the nonlinear effect is set to zero, and only the dispersion and loss are considered.
5. The method for realizing reverse ultrafast pulse laser based on reinforcement learning according to claim 4, It is characterized in that The transmitted pulse amplitude expression is: Where A(z, T) represents the pulse amplitude in the z direction within period T; D represents dispersion; N represents nonlinearity; and the linear operator The calculation formula in the frequency domain is: Among them, F -1 represents the inverse Fourier transform, ω is the angular frequency, is the complex amplitude.
6. An ultrafast pulse laser reverse realization system based on reinforcement learning, It is characterized in that include: Module M1: Obtain the seed source through mode locking technology, and determine the pulse shape and spot mode through the seed source; Module M2: Use fiber amplification technology to amplify the power of the pulse and determine the final energy and output characteristics of the pulse; Module M3: modularizes and integrates the ultrafast laser generation process, and simulates the pulse propagation process in the optical fiber through the distributed Fourier algorithm; Module M4: Input the simulation data into the long short-term memory network model LSTM according to the preset window size value, and perform step-by-step training; Module M5: Input the training results into the reinforcement learning DDPG to obtain the optimal parameters of the laser so that the laser can quickly output the specified target pulse state.
7. The ultrafast pulse laser reverse realization system based on reinforcement learning according to claim 6, It is characterized in that The action with the maximum value is found through the action estimation network, the action reality network, the state reality network and the state estimation network, the output pulse characteristics of the seed source in the output laser, the dispersion coefficient of the broadened grating, the gain coefficient and the length of the gain fiber.
8. The ultrafast pulse laser reverse realization system based on reinforcement learning according to claim 6, It is characterized in that The pulse output predicted by the model LSTM is used as the input of the reinforcement learning DDPG. Different laser parameters are selected as actions. The Q value is estimated through reinforcement learning. The Q value is the sum of the value of evaluating the current action and the reward value of estimating future actions. Then, the mean square error between the Q value under ideal pulse conditions and the Q value under the current state and action is calculated. The action that minimizes the mean square error between the simulated output pulse and the target pulse is searched as the optimal laser parameter.
9. The ultrafast pulse laser reverse realization system based on reinforcement learning according to claim 6, It is characterized in that In the simulation, the distance that the optical pulse is transmitted in the optical fiber is set to h. First, the pulse is only subjected to the nonlinear effect, and the dispersion and loss are zero. Then the nonlinear effect is set to zero, and only the dispersion and loss are considered.
10. The ultrafast pulse laser reverse realization system based on reinforcement learning according to claim 9, It is characterized in that The transmitted pulse amplitude expression is: Where A(z, T) represents the pulse amplitude in the z direction within period T; D represents dispersion; N represents nonlinearity; and the linear operator The calculation formula in the frequency domain is: Among them, F -1 represents the inverse Fourier transform, ω is the angular frequency, is the complex amplitude.
Citation Information
Patent Citations
Anti-return main oscillation power amplification pulse laser
CN113922199A