A laser coherent synthesis method based on a double-flow network and a reinforcement learning framework

By adopting a laser coherent synthesis method based on a two-stream network and reinforcement learning framework, the problems of beam quality degradation and insufficient phase control bandwidth are solved, achieving efficient multi-channel laser coherent synthesis and improving the real-time performance and robustness of the system.

CN119168881BActive Publication Date: 2026-02-13GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411189228.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2026-02-13
Estimated Expiration
2044-08-28

AI Technical Summary

Technical Problem

In existing technologies, increasing the power of a single fiber laser leads to a decrease in beam quality. Furthermore, in large-scale coherent combining systems, there are issues such as insufficient phase control bandwidth and difficulties in obtaining labels and long training times in deep learning methods.

Method used

A coherent laser synthesis method based on a two-stream network and reinforcement learning framework is adopted. By splitting the beam, phase modulation, photoelectric detection and reinforcement learning training of the neural network, coherent synthesis of multiple lasers is achieved. The robustness of the model and the AQC module are improved by using a two-stream network and AQC module.

Benefits of technology

This approach achieves improved overall output power while maintaining beam quality, solves the problems of insufficient phase control bandwidth and difficulty in acquiring deep learning labels, reduces training time to the linear level, and ensures the real-time performance and efficiency of the coherent synthesis system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119168881B_ABST
    Figure CN119168881B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of laser coherent synthesis, and discloses a laser coherent synthesis method based on a double-flow network and a reinforcement learning framework, which is used for solving the problem of low robustness of a machine learning phase control method. The method improves the efficiency and robustness of laser coherent synthesis by constructing a double-flow network structure and utilizing interference pattern information of a non-focal plane and focal plane pattern information interaction fusion. In the training process, the double-flow network is updated and optimized by using an Actor_Critic framework in reinforcement learning, so that the Actor network can adaptively adjust the phases of the laser beams, thereby realizing efficient laser coherent synthesis. The method can effectively realize phase control of the laser, realize high-power output of the coherent synthesis system, and has high robustness and real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of laser coherent synthesis, and particularly relates to a laser coherent synthesis method based on a double-flow network and a reinforcement learning framework. BACKGROUND

[0002] Optical fiber lasers are widely used in military, medical and other fields due to their compact structure, high efficiency, portability and excellent beam quality. However, as the power of a single optical fiber laser continues to increase, physical limitations may lead to a decrease in beam quality. To solve this problem, coherent beam synthesis technology has become an effective method. By synchronizing the phases of multiple optical fiber laser sub-beams, coherent maximum output is achieved, thereby maximizing the overall output power while ensuring beam quality.

[0003] However, in actual engineering, the phase synchronization and stability of the sub-beams face challenges due to various noises and disturbances, so a phase control technology must be used to ensure the optimization of the combination efficiency. Active phase-locked method is one of the commonly used phase control technologies. It compensates for dynamic phase errors through phase detection and feedback servo control systems to ensure that each sub-beam achieves in-phase coherent output and achieves the best effect.

[0004] Active phase-locked methods include heterodyne method, multi-dithering method, stochastic parallel gradient descent (SPGD) algorithm, etc. In particular, the SPGD algorithm, as a model-free optimization algorithm based on artificial intelligence, has broad application prospects. The SPGD algorithm does not require the establishment of a complex mathematical model or strict restrictions on variables. It only needs to obtain the far-field light intensity characteristics of the synthesis through a photoelectric detector without measuring the phase of each unit beam. It is simple to implement, low in cost, highly scalable and real-time, and has become a mainstream method in the field of phase locking. However, the control bandwidth of the classic active method decreases significantly with the increase in the number of sub-beams, especially in large-scale coherent beam synthesis systems (CBC), the bandwidth of phase control remains a challenge.

[0005] In recent years, with the continuous development of deep learning technology, research shows that the application of deep learning in active phase-locked technology is feasible, especially supervised learning. However, supervised learning relies on a large number of labels to train a neural network with perfect performance, which is difficult and tedious in practical applications. Environmental changes can also affect the performance of deep learning models, requiring frequent retraining of neural networks, consuming a lot of time and resources. SUMMARY

[0006] The purpose of the present application is to overcome the deficiencies of the prior art, provide a laser coherent synthesis method based on a double-flow network and a reinforcement learning framework, which can realize the coherent synthesis of multiple laser beams, can simultaneously solve the problems of small control bandwidth of traditional methods and difficult acquisition of labels of deep learning methods and long training time of reinforcement learning, and can guarantee the real-time performance of the coherent synthesis system.

[0007] The technical solution of the present application to solve the above technical problems is:

[0008] A laser coherent synthesis method based on a double-flow network and a reinforcement learning framework, comprising the following steps:

[0009] (S1), a seed laser source is divided into N beams by a beam splitter and input into a phase modulator to control the phase;

[0010] (S2), the N beams passing through the phase modulator are power amplified and output to a specific multi-aperture combiner;

[0011] (S3), the N beams of laser emitted by the specific combiner are reflected by a half-mirror, with 99% of the energy as system output, and the other part is reflected to a lens for focusing, and a photoelectric detector CCD1 is placed on the focal point to obtain the diffraction pattern I of the focal plane after the coherent synthesis of the N beams of laser 1, A photoelectric detector CCD2 is placed on the non-focal point to obtain the diffraction pattern I2 of the non-focal plane after the coherent synthesis of the N beams of laser;

[0012] (S4), the obtained diffraction pattern is input into the trained intelligent agent, and N phase control signals are output to the phase controller to correct the phases of the laser beams and improve the coherent synthesis efficiency of the N beams of laser.

[0013] Preferably, in step (S4), the training of the intelligent agent comprises the following steps:

[0014] (S4-1), an Actor and a Critic network are built, the Actor network is used to output the probability density function of N corrected phases, and the Critic network is used to evaluate the current state

[0015] (S4-2), the two diffraction images obtained by the photoelectric detector are stacked into a state input, and the feature vector obtained after inputting the state input into the double-flow network is input into the Actor network and the Critic network, the Actor network outputs the probability density function of N corrected phases, and then the N probability density functions are sampled to obtain N phase correction signals, the signals are fed back to the N phase modulators in the optical system to obtain a new state and obtain the corresponding reward;

[0016] (S4-3), the Actor network and the Critic network perform backpropagation on the gradient value of the loss function according to the reinforcement learning proximal policy optimization algorithm to update the weight and bias parameters of the neural network, output the phase correction signal again, feed the signal into the phase modulator, repeat steps (S4-2) and (S4-3), update the neural network parameters multiple times, and stop updating when the coherent synthesis evaluation function value is greater than a preset value.

[0017] (S4-4) when the number of synthesis channels is increased, the weights of the N-1 channels that have been trained are loaded as pre-trained weights into the model, the last output layer does not use the pre-trained weights, steps (S4-1) to (S4-3) are repeated, and the N-channel coherent synthesis trained agent is obtained.

[0018] Preferably, in step (S4-2), the two diffraction images obtained by the photodetector are stacked into a state input, including but not limited to a focal plane-non-focal plane image pair and a pre-modulation and post-modulation image pair, the state shape is 2*64*64, and it is input into the dual-stream network. After passing through the dual-stream network, 256 feature vectors are obtained and input into respective decoupling heads. The decoupling head of the Actor network is a fully connected layer and two parallel fully connected layers that map the output to the parameters of the probability density function of the action. The decoupling head of the Critic network is two fully connected layers, and the last layer maps the output to a specific value as an estimate of the state value.

[0019] Preferably, in step (S4-2), the dual-stream network is composed of two parameter-shared input streams and a feature exchange module AQC. The image pair is input from two input streams, and then output as 256 feature vectors after passing through the AQC module. The structure of the input stream is composed of three stages: the first part is the Stem layer, composed of two 3*3 convolutional neural networks and one SPPF module. The input is down-sampled to one-fourth of the original, and the channel number is changed to 128. The second part is the Stage1 layer, composed of one 3*3 convolutional neural network and one SPPF module. The input is down-sampled to one-half of the original, and the channel number is changed to 256. The third part is the Stage2 layer, composed of one 3*3 convolutional neural network and one SPPF module. The input is down-sampled to one-half of the original, and the channel number is changed to 256.

[0020] Preferably, in step (S4-2), the AQC module input is two 16th down-sampled feature maps, which generate Q, K, and V vectors after passing through two 3*3 convolutional neural networks. After point multiplication of the K and Q vectors, they pass through a Sigmoid function and then multiply element by element with the V vector. After global average pooling, 256 feature vectors are obtained.

[0021] Preferably, in step (S4-2), the reward expression is as follows:

[0022] reward=αPIB t +β(PIB t -PIB t-1 ) (1)

[0023] wherein, alpha and beta are adjustable parameters, PIB t is the normalized bucket power value at the current moment, PIB t-1 is the normalized bucket power value at the previous moment.

[0024] Preferably, in step (S4-3), the coherent synthesis evaluation function includes but is not limited to the bucket power detected by the photodetector, the maximum output power, the main lobe power, the synthesized beam quality factor, and the combination of the above physical quantities.

[0025] Compared with the prior art, the present application has the following beneficial effects:

[0026] 1. The present application adopts a network structure with a depth of 6 layers, and introduces a double-flow network and an AQC module, and realizes rapid fitting through parameter sharing techniques, thereby adapting to the rapid convergence of reinforcement learning algorithms. The double-flow network structure not only improves the robustness of the model under environmental changes, but also improves the stability when the non-focal plane detector position changes slightly.

[0027] 2. The present application also solves the problem of being unable to obtain labels in practical applications through a reinforcement learning update method, which provides a possibility for effective training of neural networks. In the training process, there is no need for excessive human intervention, and a large number of labels do not need to be prepared in advance. The algorithm can be directly deployed in the actual environment of the coherent synthesis system, and through multiple random phase inputs, the neural network continuously updates the parameters. After training is completed, only a part of the Actor network is extracted as an agent, and efficient phase prediction and correction can be realized. In addition, the application of the double-flow network further improves the robustness of the algorithm.

[0028] 3. Compared with other machine learning methods, the training time of the present application is not exponentially increased, and through the use of pre-training weights, the training time is reduced to a linear level. At the same time, through the use of different diffraction image input methods, the uniqueness and convergence of network training are ensured, which provides higher feasibility and efficiency for the laser coherent synthesis method based on reinforcement learning. BRIEF DESCRIPTION OF DRAWINGS

[0029] Figure 1 is a system block diagram of a laser coherent synthesis method based on a double-flow network and a reinforcement learning framework of the present application.

[0030] Figure 2 is a structural schematic diagram of the double-flow network of the present application.

[0031] Figure 3 Structure diagram of each module of the present application.

[0032] Figure 4 Network structure composition (Table 1). DETAILED DESCRIPTION

[0033] The present application will be further described in conjunction with the embodiments and the accompanying drawings, but the embodiments of the present application are not limited thereto.

[0034] Referring to Figure 1 , the implementation of the laser coherent synthesis method based on a double-flow network and a reinforcement learning framework of the present application includes the following steps:

[0035] (S1), the seed laser is divided into 7 beams by a beam splitter and input into a phase modulator to control the phase;

[0036] (S2), the 7 beams of light passing through the phase modulator are power amplified and output to a specific split-aperture combiner;

[0037] (S3), the 7 beams of laser emitted by the specific combiner are reflected by a half-transmission half-reflection mirror, with 99% of the energy as system output, and the other part is reflected to a lens for focusing. A photodetector CCD1 is placed at the focal point to obtain the diffraction pattern I1 of the focal plane after the coherent synthesis of the 7 beams of laser 1, A photodetector CCD2 is placed at a non-focal point to obtain the diffraction pattern I2 of the non-focal plane after the coherent synthesis of the 7 beams of laser;

[0038] (S4), the obtained diffraction patterns are input into the trained Actor network, and 7 phase control signals are output to the phase controller to correct the phases of each laser beam, thereby improving the coherent synthesis efficiency of the 7 beams of laser.

[0039] In step (S4), the training of the Actor network includes the following steps:

[0040] (S4-1), two networks of Actor and Critic are built, the Actor network is used to output the probability density function of N correction phases, and the Critic network is used to evaluate the current state

[0041] (S4-2), the two diffraction images obtained by the photodetector are stacked into a state input, and the feature vector obtained after inputting the state input into the double-flow network is input into the Actor network and the Critic network. After the Actor network outputs the probability density function of N correction phases, the N probability density functions are respectively sampled to obtain N beam phase correction signals, which are fed back to the N phase modulators in the optical system to obtain a new state and the corresponding reward.

[0042] (S4-3), the Actor network and the Critic network perform backpropagation on the gradient value of the loss function according to the reinforcement learning proximal policy optimization algorithm to update the weight and bias parameters of the neural network, output the phase correction signal again, feed the signal into the phase modulator, repeat steps (S4-2) and (S4-3), update the neural network parameters multiple times, and stop updating when the coherent synthesis evaluation function value is greater than a preset value.

[0043] (S4-4) when the number of synthesis paths is increased, the weights of N-1 paths that have been trained are loaded as pre-trained weights on the model, the last output layer does not use the pre-trained weights, steps (S4-1) to (S4-3) are repeated, and an N-path coherent synthesis trained agent is obtained.

[0044] Referring to Figure 2 and Table 1, the dual-stream network is composed of 2 parameter-shared input streams and a feature exchange module AQC. The image pairs enter from the 2 input streams and then pass through the AQC module to output 256 feature vectors. The structure of the input stream is composed of 3 stages: the first part is the Stem layer, which is composed of 2 3*3 convolutional neural networks and 1 SPPF module. The input is down-sampled to 1 / 4 of the original, and the channel number is changed to 128. The second part is the Stage1 layer, which is composed of a 3*3 convolutional neural network and 1 SPPF module. The input is down-sampled to 1 / 2 of the original, and the channel number is changed to 256. The third part is the Stage2 layer, which is composed of a 3*3 convolutional neural network and 1 SPPF module. The input is down-sampled to 1 / 2 of the original, and the channel number is changed to 256. The decoupling head of the Actor network is a fully connected layer and 2 parallel fully connected layers that map the output to the parameters of the probability density function of the action; the decoupling head of the Critic network is 2 fully connected layers, and the last layer maps the output to a specific value as an estimate of the state value.

[0045] Referring to Figure 3 , the input of the AQC module is 2 16th down-sampled feature maps, which are respectively generated into Q, K, and V vectors through 2 3*3 convolutional neural networks. After the K and Q vectors are multiplied and then passed through a Sigmoid function, they are multiplied with the V vector element by element, and then global average pooling is performed to obtain 256 feature vectors. The input of the SPPF module is a feature map, which is reduced in channel dimension or fused through a 1*1 convolutional neural network, thereby reducing the amount of calculation and improving the compactness of feature expression. Then, the feature map passes through three consecutive max-pooling layers, each of which can have a different kernel size. Next, the outputs of the three max-pooling operations are concatenated with the output of the original 1x1 convolution in the channel dimension. Finally, the concatenated features pass through a 1x1 convolution layer for feature fusion or dimension increasing processing to output the final feature map.

[0046] In addition, in step (S4-2), the state input is two interference images obtained by the focal plane photodetector CCD1 and the non-focal plane photodetector CCD2 respectively, and the two images are given to the network as a state input. The next action input is waited for, and the process is repeated.

[0047] In addition, in step (S4-2), the reward expression is as follows:

[0048] reward = aPIB t + b(PIB t - PIB t-1 ) (1)

[0049] Wherein, a and b are adjustable parameters, PIB t is the normalized bucket power value at the current moment, and PIB t-1 is the normalized bucket power value at the last moment.

[0050] In addition, in step (S4-3), the coherent synthesis evaluation function includes but is not limited to the bucket power detected by the photodetector, the highest output power, the main lobe power, the synthesized beam quality factor, and the combination of the above physical quantities.

[0051] The above is the preferred embodiment of the present application, but the embodiments of the present application are not limited by the above, and any change, modification, substitution, combination, simplification made without departing from the spirit and principles of the present application shall be an equivalent replacement method, which is included in the protection scope of the present application.

Claims

1. A laser coherent synthesis method based on a two-stream network and a reinforcement learning framework, characterized in that, Includes the following steps: (S1) The seed laser source is split into N beams by a beam splitter and input into a phase modulator to control the phase; (S2) Amplify the power of N beams that have passed through the phase modulator and output them to a specific aperture combiner; (S3) The N laser beams emitted in the specific aperture beam combiner are reflected by a semi-transparent mirror, with 99% of the energy being used as the system output. The other part is reflected onto the lens for focusing. The photodetector CCD1 is placed at the focal point to obtain the diffraction image I of the focal plane after the N laser beams are coherently combined. 1, The photodetector CCD2 is placed on the non-focal point to obtain the diffraction image I2 of the non-focal plane after coherent synthesis of N laser beams; (S4) Input the obtained diffraction image into the trained agent and output N phase control signals to the phase modulator to correct the phase of each laser, thereby improving the coherent synthesis efficiency of N laser beams. This involves stacking two diffraction images acquired by a photodetector into a single state input, namely a focal plane-off-focal plane image pair, and then inputting it into a two-stream network. After passing through the two-stream network, 256 feature vectors are obtained and input into the decoupling heads of the Actor and Critic networks. The two-stream network consists of two parameter-shared input streams and a feature exchange module (AQC). The image pairs enter from the two input streams respectively, and then the AQC module outputs 256 feature vectors. The input stream structure consists of three stages: the first part is the Stem layer, composed of two 3x3 convolutional neural networks and one SPPF module, where the input is downsampled to one-quarter of its original size, and the number of channels is changed. The first part is Stage 1, consisting of a 3x3 convolutional neural network and an SPPF module. The input is downsampled to half its original size, and the number of channels becomes 256. The second part is Stage 2, consisting of a 3x3 convolutional neural network and an SPPF module. The input is downsampled to half its original size, and the number of channels becomes 256. The AQC module takes two 1 / 16 downsampled feature maps as input, which are processed by two 3x3 convolutional neural networks to generate Q, K, and V vectors. The K and Q vectors are multiplied by a dot product, then multiplied by the element-wise of the V vector using the Sigmoid function, and finally multiplied by global average pooling to obtain 256 feature vectors.

2. The laser coherent synthesis method based on a two-stream network and reinforcement learning framework according to claim 1, characterized in that, The decoupling head of the Actor network consists of a fully connected layer and two parallel fully connected layers that map the output to the parameters of the probability density function of the action. The decoupling head of the Critic network consists of two fully connected layers, with the last layer mapping the output to a specific value as an estimate of the state value.

3. The laser coherent synthesis method based on a two-stream network and reinforcement learning framework according to claim 1, characterized in that, In step (S4), obtaining the trained agent includes the following steps: (S4-1) Construct two networks, Actor and Critic. The Actor network is used to output the probability density function of N corrected phases, and the Critic network is used to evaluate the current state. (S4-2) Stack the two diffraction images acquired by the photodetector into a state input, and input it into the feature vector obtained by the two-stream network. Then input the Actor network and the Critic network. After the Actor network outputs the probability density functions of N corrected phases, sample the N probability density functions respectively to obtain N beam phase correction signals. Feed the signals back to the N phase modulators in the optical system to obtain a new state and obtain the corresponding reward. (S4-3) The Actor network and Critic network backpropagate the gradient values ​​of the loss function of the reinforcement learning proximal policy optimization algorithm to update the weights and bias parameters of the neural network, output the phase correction signal again, feed the signal back to the phase modulator, repeat steps (S4-2) and (S4-3), update the neural network parameters multiple times, and stop updating when the coherent synthesis evaluation function value is greater than the preset value. (S4-4) When increasing the number of synthesis paths, load the weights of the N-1 paths that have already been trained as pre-trained weights onto the model. The final output layer does not use the pre-trained weights. Repeat steps (S4-1) to (S4-3) to obtain an agent that has been trained with N coherent synthesis.

4. The laser coherent synthesis method based on a two-stream network and reinforcement learning framework according to claim 3, characterized in that, The reward expression is as follows: reward=αGDP t +β(GDP t -GDP t-1 ) (1) Where α and β are adjustable parameters, PIB t The normalized power value in the bucket at the current moment, PIB t-1 This is the normalized power value in the bucket at the previous moment.

5. The laser coherent synthesis method based on a two-stream network and reinforcement learning framework according to claim 3, characterized in that, The coherent combining evaluation function includes the power in the barrel detected by the photodetector, the maximum output power, the main lobe power, the combined beam quality factor, and a combination of the above physical quantities.

Citation Information

Patent Citations

  • Laser coherent combination phase control method based on reinforcement learning

    CN116316023A

  • Intelligent phase control method for laser coherent combination

    CN117199984A