A control method and device for air laser modulation
By optimizing the liquid crystal molecule arrangement of the spatial light modulator using reinforcement learning algorithms, the problems of femtosecond laser propagation distortion and air laser signal optimization under air pressure conditions were solved, resulting in a significant enhancement of the air laser signal intensity.
Patent Information
- Application Number
- CN202411165328.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-23
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-08-23
AI Technical Summary
In existing technologies, femtosecond lasers cause beam distortion when propagating in a medium, making it difficult to optimize N2+ air laser signals under different air pressure conditions. Furthermore, manual adjustment based on experience is inefficient and cannot meet the requirements of high-intensity air lasers.
The DDPG network, employing reinforcement learning algorithms, optimizes the liquid crystal molecule arrangement of a spatial light modulator. By adjusting the laser wavefront phase through a phase grayscale driving matrix and combining the reward value from the spectrometer, the optimal actor-critic network is trained, achieving precise modulation of the laser.
The intensity of the N2+ air laser signal was significantly enhanced under different nitrogen pressure conditions, with an increase of up to one order of magnitude, thereby improving the focused light intensity of the air laser.
Smart Images

Figure CN119045226B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of laser modulation, and in particular to a control method and device for air laser modulation. Background Technology
[0002] The interaction between near-infrared intense femtosecond laser and air / nitrogen can induce the generation of N2 at 391 nm and 428 nm. + Air lasers. This phenomenon has significant applications in long-range atmospheric detection, studying molecular rotational coherence, and generating supercontinuum spectra. However, the laser output from a femtosecond laser amplifier system is not strictly a perfect Gaussian beam, and when a strong femtosecond laser pulse propagates in a medium, nonlinear effects cause distortion of the pulse wavefront, resulting in a decrease in intensity at the focal point after the ultrafast laser is focused. This is unfavorable for generating high-intensity air lasers. Furthermore, current research focuses on optimizing N2 under different atmospheric pressure conditions. + The air laser signal still relies on manual adjustment based on experience, thus significantly enhancing N2 under different air pressure conditions. + Air laser signals present certain challenges. Summary of the Invention
[0003] The purpose of this application is to provide a control method and device for air laser modulation, which can more accurately and quickly modulate the laser using a spatial light modulator, thereby enhancing the intensity of the air laser signal.
[0004] To achieve the above objectives, this application provides the following solution:
[0005] In a first aspect, this application provides a method for controlling air laser modulation, comprising:
[0006] The phase grayscale driving matrix of the laser modulated by the spatial light modulator at the current moment is obtained and used as the current state data; the phase grayscale driving matrix is used to control the grayscale value of the pixel matrix on the phase plate in the spatial light modulator, thereby determining the liquid crystal arrangement state; the liquid crystal arrangement state is used to modulate the wavefront phase distribution of the laser.
[0007] The current state data is input into the actor network of the DDPG algorithm to obtain the corresponding current action; the current action is a matrix adjustment parameter used to adjust the grayscale value of the pixel matrix on the phase plate.
[0008] The phase grayscale driving matrix is updated based on the current action to determine the next state data. Then, the next state data is loaded into the spatial light modulator to adjust the liquid crystal alignment state, and the spectrum of the corresponding air laser collected by the spectrometer is obtained.
[0009] The current reward value is calculated based on the spectrum of air laser collected by the spectrometer; the current state data, the corresponding current action, the current reward value, and the next state data constitute experience data and are stored in the experience replay pool;
[0010] A training sample set is randomly sampled from the experience replay pool. The current state data and corresponding current action of any training sample are input into the critic network of the DDPG algorithm to obtain the corresponding action value. The next state data of the training sample is input into the target actor network of the DDPG algorithm to obtain the corresponding next action. Then, the next state data and the corresponding next action are input into the target critic network of the DDPG algorithm to obtain the next action value. The target action value is calculated based on the next action value and the current reward value in the training sample. The actor-critic network is trained using the action value corresponding to the current state data and the target action value to obtain the optimal actor-critic network. The action value is used to characterize the quality of the current action. The optimal actor-critic network is used to determine the corresponding optimal matrix adjustment parameters according to the phase grayscale driving matrix of the spatial light modulator when receiving laser light.
[0011] Secondly, this application provides an air laser modulation device, including a laser, a beam splitter, a first transmission optical path component, a second transmission optical path component, a spatial light modulator, a controller, a reflector, a dichroic mirror, an air laser generating component, and a spectrometer.
[0012] The laser is used to output femtosecond laser light to the beam splitter; the beam splitter is used to split the received femtosecond laser light into two beams, and output them to the first transmission optical path component and the second transmission optical path component respectively; wherein, the single pulse energy of the laser in the first transmission optical path component is higher than the single pulse energy of the laser in the second transmission optical path component.
[0013] The first transmission optical path component is used to perform energy and polarization adjustment on the received laser light before transmitting it to the spatial light modulator; the spatial light modulator is used to determine the phase grayscale driving matrix when receiving the laser light; the controller is used to: determine the optimal actor-critic network determined by the control method of air laser modulation, determine the corresponding optimal matrix adjustment parameters according to the phase grayscale driving matrix, and send them to the spatial light modulator; the spatial light modulator is also used to load the optimal matrix adjustment parameters to change the liquid crystal arrangement state, thereby realizing phase modulation of the received laser light; the phase-modulated laser light is reflected by the reflector to the dichroic mirror.
[0014] The second transmission optical path component is used to perform frequency doubling, filtering, and delay on the received laser before sending it to the dichroic mirror; the dichroic mirror is used to combine the laser from the first transmission optical path component and the laser from the second transmission optical path component to obtain a combined laser beam and send it to the air laser generating component;
[0015] The air laser generating component is used to react the combined laser beam with nitrogen gas to generate an air laser and focus it onto the spectrometer; the spectrometer is used to collect the spectrum of the air laser and send it to the controller.
[0016] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a control method and device for air laser modulation. The phase grayscale driving matrix of the laser modulated by the spatial light modulator at the current moment is used as state data. Adjustment parameters used to adjust the grayscale values of the pixel matrix on the phase plate are used as actions. The current reward value is calculated based on the spectrum of the air laser collected by the spectrometer. Based on the above parameters, an actor-critic network using the DDPG algorithm is trained and applied. This allows for the determination of the optimal matrix adjustment parameters given the phase grayscale driving matrix of the spatial light modulator when receiving laser light. This more accurately and quickly changes the arrangement of liquid crystal molecules on the spatial light modulator, modulating the wavefront phase of the laser beam reflected by the spatial light modulator, thereby increasing the light intensity at the focal point and ultimately enhancing the air laser signal strength. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating a control method for air laser modulation in one embodiment of this application.
[0019] Figure 2 This is a schematic diagram illustrating the process of using a reinforcement learning-based DDPG algorithm to control a spatial light modulator and thereby optimize an air laser signal, as provided in one embodiment of this application.
[0020] Figure 3 This is a schematic diagram of the structure of the DDPG algorithm provided in an embodiment of this application.
[0021] Figure 4 This is a schematic diagram of the structure of an air laser modulation device in another embodiment of this application.
[0022] Figure 5 is a schematic diagram comparing the signal strength obtained by different modulation methods provided in this application.
[0023] Figure label:
[0024] 1-Laser, 2-Beam splitter, 3-First half-wave plate, 4-Polarizer, 5-Second half-wave plate, 6-Spatial light modulator, 7-Controller, 8-Reflector, 9-Frequency doubling crystal, 10-Filter, 11-First reflector, 12-Second reflector, 13-Third reflector, 14-Dichroic mirror, 15-First lens, 16-Vacuum cavity, 17-Second lens, 18-Third lens, 19-Filter, 20-Spectrometer. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0026] Spatial light modulators can modulate the phase of a light beam in arbitrary shapes, thereby compensating for the phase distortion of ultrafast and ultra-intense femtosecond laser pulses, ultimately increasing the intensity of the focused light and enhancing the intensity of the air laser signal. Machine learning is an important approach to artificial intelligence. Machine learning can be divided into three categories: supervised learning (SL), unsupervised learning (UL), and reinforcement learning (RL). Among them, reinforcement learning, which uses trial and error and rewards to train the agent's behavior, is quite different from the other two branches. The core idea of reinforcement learning is that the agent interacts with its environment to determine a policy. First, the agent performs an action based on its state. The environment rewards this action and changes the agent's state. Then, the agent decides its next action based on the magnitude of the reward. If the reward is positive, the agent's tendency to perform the action will be strengthened.
[0027] Based on this, this application utilizes reinforcement learning algorithms to automatically optimize and adjust the arrangement of liquid crystal molecules on the spatial light modulator, thereby modulating the wavefront phase of the femtosecond laser beam reflected by the spatial light modulator, ultimately increasing the light intensity at the focal point and significantly enhancing N2 under different nitrogen gas pressure conditions (30 mbar and atmospheric pressure). + Air laser signals can reach up to an order of magnitude.
[0028] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0029] In one exemplary embodiment, a control method for air laser modulation is provided. This method is executed by a computer device, specifically a terminal or server, or both. Specifically, this application uses the reinforcement learning-based DDPG algorithm to control a spatial light modulator and optimize the air laser signal. Before processing the algorithm, the relevant parameters of the DDPG algorithm are defined, including the following six parameters: state s, action a, reward value r, action value Q, policy network actor, and value network critic.
[0030] State Data: It should be noted that in this application, the wavefront phase distribution of the pump light (or laser) is altered by a spatial light modulator. The working principle of the spatial light modulator is to independently control the grayscale value of each pixel by changing the addressing voltage. That is, by changing the grayscale value of a two-dimensional grayscale image of size "1024×1272" (the grayscale value range is 0-255), the voltage is controlled, thereby changing the liquid crystal alignment state, thus changing the refractive index of each pixel, and ultimately changing the wavefront phase of the femtosecond laser reflected by the pixel. Therefore, the state data s can be represented as the pixel matrix Mask controlling the spatial light modulator, or it can be expressed as a phase grayscale driving matrix.
[0031] Action: Since adjusting pixel parameters one by one increases the difficulty of data processing, this application uses a circularly distributed two-dimensional pixel group to control the phase plate, thereby reducing the amount of data processing and efficiently adjusting pixel parameters. Specifically, this application sets the action as three adjustment parameters: outer circle radius a1, inner circle radius a2, and grayscale value a3. By adjusting these three parameters, the size and brightness are represented and adjusted in the form of a circular pixel matrix. At the same time, in order to prevent the algorithm from converging to a local optimum too early, this application adds a noise B to the deterministic policy to increase the agent's detection in the action space. That is, the action selected by the agent is jointly determined by the policy function and the noise, expressed as: action a = (a1, a2, a3) + B.
[0032] Reward value: In this application, N2 + The intensity of the lasing signal (Airlaser) is collected by a fiber optic spectrometer (other spectrometers can be set as needed) and integrated and summed using a program to obtain the spectral signal intensity; then, the N2 collected at the previous moment is calculated. +The difference rate between the lasing signal intensity and the intensity at the next moment is used as the current reward value r. Specifically, the function for calculating the current reward value is: r = (I new -I old ) / I old Among them, I new I represents the current spectral signal intensity. old The previous spectral signal intensity represents the signal intensity value of the air laser spectrum collected by the spectrometer corresponding to the current moment; the previous spectral signal intensity represents the signal intensity value of the air laser spectrum collected by the spectrometer corresponding to the previous moment.
[0033] Action value Q: This represents the value given by the critic network based on the current input state data and the action, used to measure the quality of the current action 'a'.
[0034] Policy network: also known as actor network, contains convolutional layers and fully connected layers, takes the current state data as input, and outputs actions.
[0035] Value network: also known as critic network, contains convolutional layers and fully connected layers. It takes the current state data and action as input and outputs action value Q. The action value Q is used to estimate the expected reward for a given state and action under the current policy.
[0036] In the embodiments of this application, such as Figure 1 and Figure 2 As shown, the method specifically includes the following steps 100-500.
[0037] Step 100: Obtain the phase grayscale driving matrix of the spatial light modulator modulating the laser at the current moment, and use it as the current state data; the phase grayscale driving matrix is used to control the grayscale value of the pixel matrix on the phase plate in the spatial light modulator, thereby determining the liquid crystal arrangement state; the liquid crystal arrangement state is used to modulate the wavefront phase distribution of the laser.
[0038] In a specific application instance, the system executing the method needs to be initialized before the method is executed to avoid previous data affecting subsequent processing.
[0039] Step 200: Input the current state data into the actor network of the DDPG algorithm to obtain the corresponding current action; the current action is a matrix adjustment parameter used to adjust the grayscale value of the pixel matrix on the phase plate.
[0040] Step 300: Update the phase grayscale driving matrix based on the current action to determine the next state data, then load the next state data into the spatial light modulator to adjust the liquid crystal alignment state, and obtain the spectrum of the corresponding air laser collected by the spectrometer.
[0041] In a specific application example, the next state data is loaded into the spatial light modulator to adjust the liquid crystal alignment state. Based on the adjusted liquid crystal alignment state, the wavefront phase distribution of the femtosecond laser beam reflected by the spatial light modulator is modulated. Then, laser transmission is performed according to the air laser modulation device described below in this application until air laser is generated and the corresponding spectrum is collected by a spectrometer.
[0042] Step 400: Calculate the current reward value based on the spectrum of the air laser collected by the spectrometer. The current state data, along with the corresponding current action, current reward value, and next state data, constitute experience data and are stored in the experience replay pool.
[0043] It should be noted that throughout the entire loop, the next state data determined in the current step will be used as the current state data for the next moment for subsequent processing.
[0044] Step 500: Randomly sample a training sample set from the experience replay pool. Input the current state data and corresponding current action of any training sample into the critic network of the DDPG algorithm to obtain the corresponding action value. Input the next state data of the training sample into the target actor network of the DDPG algorithm to obtain the corresponding next action. Then, input the next state data and corresponding next action into the target critic network of the DDPG algorithm to obtain the next action value. Calculate the target action value based on the next action value and the current reward value in the training sample. Train the actor-critic network using the action value corresponding to the current state data and the target action value to obtain the optimal actor-critic network. The action value is used to characterize the quality of the current action. The optimal actor-critic network is used to determine the corresponding optimal matrix adjustment parameters based on the phase grayscale driving matrix of the spatial light modulator when receiving laser light.
[0045] In one application example, during network training, a training cycle (episode) is defined as 32 time points. In the first training cycle, the agent randomly samples actions, also known as the sampling cycle. The learned experience data during this process is stored in the experience replay pool R as initial experience. The total number of training cycles can be set as needed, such as 300 training cycles. After the sampling cycle ends, the training cycle begins. The agent calculates actions based on the decisions of the actor network, and Gaussian distributed random noise B is added to the action output value for exploring new actions. The reward value r for the corresponding action is calculated, and the new experience data is stored in R. During the training cycle, once sufficient experience data has accumulated, the agent enters the training process each time it interacts with the environment. During training, a portion of the experience is randomly sampled from the experience replay pool to update the neural network parameters (actor and critic) in the DDPG algorithm. Furthermore, after a certain number of training cycles (e.g., 200 training cycles), the random noise B added to the action output value is removed.
[0046] It should be noted that DDPG is an actor-critic algorithm, and the goal of the DDPG algorithm is to learn a deterministic policy μ(s|θ). μ This strategy can directly output an action a that maximizes the action value Q based on the state data s. The DDPG algorithm also uses a computer neural network instead of the actor network μ(s|θ) μ ) and the critic network Q(s,a|θ Q The DDPG algorithm utilizes experience replay technology to improve the efficiency of using experience data and reduce the correlation between data points. Additionally, it sets up actor and critic target networks to enhance training stability. Therefore, the DDPG algorithm requires four neural networks, as shown in the algorithm structure diagram below. Figure 3 As shown.
[0047] Combination Figure 3 During the training of the actor-critic network using the training sample set, the critic network is updated by minimizing the first loss function through gradient descent; wherein, the first loss function is:
[0048]
[0049] The actor network is updated by maximizing the second loss function using gradient ascent; where the second loss function is:
[0050]
[0051] Where L represents the value of the first loss function, N represents the total number of training samples, and y iLet Q represent the target action value, and s represent the action value. i Represents the state data at time i, a i Represents the action at time i; Q(s) i ,a i |θ Q ) represents the parameter θ in the critic network. Q Based on the state data and actions at time i, the action value is given. J(θ) represents the expectation; μ ) represents the value of the second loss function; μ(s) i |θ μ ) represents the parameter θ in the actor network. μ Based on the state data s i The strategy for outputting actions; Q(s) i ,μ(s i |θ μ )|θ Q ) represents the parameter θ in the critic network. Q Based on the state data at time i and μ(s) i |θ μ The value of the action given.
[0052] During the training of the actor-critic network, the parameters of the target actor network and the target critic network are softly updated, as expressed by the formula:
[0053] θ Q' ←τθ Q +(1-τ)θ Q' θ μ' ←τθ μ +(1-τ)θ μ' .
[0054] Where, θ Q' θ represents the parameters of the target critic network. μ' τ represents the parameters of the target actor network; τ is a hyperparameter that ranges from 0 to 1, and is often set to a small value, such as τ = 0.005, which controls the target network to update its parameters at a slow rate. This soft update method can improve the stability of training.
[0055] After training based on the aforementioned loss function, the optimal actor-critic network can be obtained. It should be noted that an optimal actor-critic network corresponds to a specific gas pressure (the pressure inside the vacuum chamber); for different gas pressure requirements, steps 100-500 above need to be executed separately to obtain the optimal actor-critic network corresponding to different gas pressures. The optimal actor-critic network obtained by training under different gas pressures can significantly enhance N2 under different nitrogen gas pressure conditions (30 mbar and atmospheric pressure). + Air laser signals can achieve an effect of up to an order of magnitude.
[0056] Based on the same inventive concept, this application also provides an air laser modulation device employing the above-described method. The solution provided by this device is similar to the solution described in the above-described method; therefore, the specific limitations in one or more device embodiments provided below can be found in the limitations of the method described above, and will not be repeated here.
[0057] In one exemplary embodiment, such as Figure 4 As shown, an air laser modulation device is provided, including a laser 1, a beam splitter 2, a first transmission optical path component, a second transmission optical path component, a spatial light modulator 6, a controller 7, a reflector 8, a dichroic mirror 14, an air laser generating component, and a spectrometer 20.
[0058] The laser 1 is used to output femtosecond laser to the beam splitter 2; in one application example, this application uses a Ti:Sapphire femtosecond laser amplifier system, which outputs a femtosecond laser pulse with a wavelength of 800nm, a pulse width of about 50fs, a single pulse energy of 1mJ, and a repetition frequency of 800Hz.
[0059] The beam splitter 2 is used to split the received femtosecond laser into two beams, which are then output to the first transmission optical path component and the second transmission optical path component, respectively. The single-pulse energy of the laser in the first transmission optical path component is higher than that in the second transmission optical path component. Furthermore, the reflectivity to transmittance ratio of the beam splitter 2 is 9:1.
[0060] For the first transmission optical path component, the single-pulse energy of its laser is relatively high, and it can be used as pump light for subsequent processing. The first transmission optical path component is used to adjust the energy and polarization of the received laser before transmitting it to the spatial light modulator 6. Specifically, the first transmission optical path component includes a first half-wave plate 3, a polarizer 4, and a second half-wave plate 5 arranged sequentially; the first half-wave plate 3 and the polarizer 4 are used to adjust the single-pulse energy of the laser by changing the optical axis angle and transmit the resulting linearly polarized light to the second half-wave plate 5; the second half-wave plate 5 is used to adjust the polarization direction of the received linearly polarized light to horizontal polarization and transmit it to the spatial light modulator 6, specifically, the horizontally polarized laser hits the liquid crystal panel of the spatial light modulator 6 (LCOS-SLM).
[0061] In one application example, a computer is used as controller 7. A pre-written reinforcement learning artificial intelligence program on the computer automatically controls the liquid crystal arrangement state of the spatial light modulator, thereby changing the wavefront phase distribution of the input laser pulse (without changing the intensity and polarization state of the laser) to achieve phase modulation of the pump laser pulse. The specific algorithm process is detailed in the embodiment of the air laser modulation control method described above.
[0062] The spatial light modulator 6 is connected to a computer and is used to determine the phase grayscale driving matrix when receiving laser light. The controller 7 is used to: determine the optimal actor-critic network determined by the control method of air laser modulation, determine the corresponding optimal matrix adjustment parameters according to the phase grayscale driving matrix, and send them to the spatial light modulator 6. The spatial light modulator 6 is also used to load the optimal matrix adjustment parameters to change the liquid crystal arrangement state and realize the phase modulation of the received laser light. The phase-modulated laser light is reflected by the reflector 8 to the dichroic mirror 14.
[0063] In one application example, mirror 8 is an 800 high-reflection mirror (800HR).
[0064] The second transmission optical path component has a low single-pulse energy laser. This component is used to frequency-double, filter, and delay the received laser light before transmitting it to the dichroic mirror 14. The second transmission optical path component includes a frequency-doubled crystal 9, a filter 10, and a mirror group arranged sequentially. The frequency-doubled crystal 9 generates a second harmonic wave based on the received laser light and sends this second harmonic wave as seed light to the filter 10. The frequency-doubled crystal 9 has a thickness of 200 μm, and after passing through it, a second harmonic wave with a center wavelength of approximately 400 nm is generated. The filter 10 filters and selects the frequency of the seed light to remove noise; specifically, the filter 10 filters out light with a wavelength of 800 nm. The mirror group delays the filtered seed light and transmits it to the dichroic mirror 14. In one application example, the reflector group includes three reflectors arranged in sequence, namely a first reflector 11, a second reflector 12 and a third reflector 13, and all three reflectors are 400 HR. By changing the relative position of the three reflectors to delay the time, the laser from the first transmission optical path component and the laser from the second transmission optical path component can reach the dichroic mirror 14 at the same time.
[0065] The dichroic mirror 14 is used to combine the laser beams from the first transmission optical path component and the second transmission optical path component to obtain a combined laser beam, which is then sent to the air laser generating component. In one application example, the dichroic mirror 14 exhibits high reflectivity under light with a wavelength of 800 nm and high transmittance under light with a wavelength of 400 nm.
[0066] The air laser generating assembly is used to react the combined laser beam with nitrogen gas to generate an air laser and focus it onto the spectrometer 20. In one application example, the air laser generating assembly includes a first lens 15, a vacuum cavity 16, a second lens 17, and a third lens 18 arranged sequentially. The first lens 15 is used to focus the received combined laser beam onto the vacuum cavity 16 to generate an air laser within the vacuum cavity 16. Specifically, the first lens 15 is a quartz lens with a focal length of 7.5 cm, and the vacuum cavity 16 is filled with nitrogen gas and is a metal cavity. In practical applications, when generating N2... + In the signal experiment of the 391nm air laser, the air pressure in vacuum cavity 16 was 30mbar; when generating N2... + In the 428nm air laser experiment, the air pressure in the vacuum chamber 16 was 1000 mbar. The second lens 17 was used to collimate the air laser output from the vacuum chamber 16; the second lens 17 was a quartz lens with a focal length of 7.5 cm. The third lens 18 was used to focus the collimated air laser onto the spectrometer 20; the third lens 18 was a quartz collecting lens with a focal length of 6 cm to collect the generated N2. +An air laser is focused into the entrance slit of a spectrometer 20 (specifically, a fiber optic spectrometer).
[0067] The air laser generating assembly also includes a filter 19; the filter 19 is disposed between the third lens 18 and the spectrometer 20, specifically between the third lens 18 and the fiber optic probe of the spectrometer 20. The filter 19 is used to filter out stray light in the collimated air laser, such as 800nm fundamental frequency light and other stray light, so that the collected light consists only of seed light and N2. + Air laser.
[0068] The spectrometer 20 is used to collect the spectrum of the air laser and send it to the controller 7. Specifically, N2 + The air laser spectrum is collected by a fiber optic spectrometer and transmitted to controller 7. An artificial intelligence program performs intensity analysis, and through reinforcement learning algorithms, automatically adjusts the arrangement of liquid crystal molecules on the spatial light modulator based on changes in signal intensity, thereby altering the wavefront phase of the pump light. This ultimately increases the light intensity at the focal point, significantly enhancing the N2 intensity under different nitrogen pressure conditions. + Air laser signals can achieve an effect of up to an order of magnitude.
[0069] Figure 5 shows a comparison of signal intensities obtained by different modulation methods. The solid line (Mask) represents the air laser intensity spectrum after modulation, the dashed line (None) represents the air laser intensity spectrum adjusted to the optimal level manually without modulation, the dotted line (Seed) represents the corresponding seed spectrum intensity, and the inset to the left of the spectrum is the modulated phase plate. Figure 5(a) shows the air laser modulation effect at 391 nm under 30 mbar and the corresponding modulated phase plate. According to Figure 5(a), the signal intensity of the modulated air laser at the solid line (Mask) is an order of magnitude higher than the manually adjusted optimal dashed line (None). Figure 5(b) shows the air laser modulation effect at 428 nm under atmospheric pressure and the corresponding modulated phase plate. According to Figure 5(b), the signal intensity of the modulated air laser at the solid line (Mask) is 3-4 times higher than the manually adjusted optimal black dashed line (None).
[0070] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0071] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for controlling air laser modulation, characterized in that, The control method for air laser modulation includes: The phase grayscale driving matrix of the laser modulated by the spatial light modulator at the current moment is obtained and used as the current state data; the phase grayscale driving matrix is used to control the grayscale value of the pixel matrix on the phase plate in the spatial light modulator, thereby determining the liquid crystal arrangement state; the liquid crystal arrangement state is used to modulate the wavefront phase distribution of the laser. The current state data is input into the actor network of the DDPG algorithm to obtain the corresponding current action; the current action is a matrix adjustment parameter used to adjust the grayscale value of the pixel matrix on the phase plate. The phase grayscale driving matrix is updated based on the current action to determine the next state data. Then, the next state data is loaded into the spatial light modulator to adjust the liquid crystal alignment state, and the spectrum of the corresponding air laser collected by the spectrometer is obtained. The current reward value is calculated based on the spectrum of air laser collected by the spectrometer; the current state data, the corresponding current action, the current reward value, and the next state data constitute experience data and are stored in the experience replay pool; A training sample set is randomly sampled from the experience replay pool. The current state data and corresponding current action of any training sample are input into the critic network of the DDPG algorithm to obtain the corresponding action value. The next state data of the training sample is input into the target actor network of the DDPG algorithm to obtain the corresponding next action. Then, the next state data and the corresponding next action are input into the target critic network of the DDPG algorithm to obtain the next action value. The target action value is calculated based on the next action value and the current reward value in the training sample. The actor-critic network is trained using the action value corresponding to the current state data and the target action value to obtain the optimal actor-critic network. The action value is used to characterize the quality of the current action. The optimal actor-critic network is used to determine the corresponding optimal matrix adjustment parameters according to the phase grayscale driving matrix of the spatial light modulator when receiving laser light.
2. The control method for air laser modulation according to claim 1, characterized in that, The adjustment parameters used to adjust the grayscale values of the pixel matrix on the phase plate include the outer radius, inner radius, grayscale value, and noise, so that the corresponding action is characterized in terms of both size and brightness in the form of a circular pixel matrix.
3. The control method for air laser modulation according to claim 1, characterized in that, The function for calculating the current reward value is: r = (I new -I old ) / I old Where r represents the current reward value, I new I represents the current spectral signal intensity. old The previous spectral signal intensity is represented by the signal intensity value of the air laser spectrum collected by the spectrometer corresponding to the current moment; the previous spectral signal intensity is the signal intensity value of the air laser spectrum collected by the spectrometer corresponding to the previous moment.
4. The control method for air laser modulation according to claim 1, characterized in that, During the training of the actor-critic network, the critic network is updated by minimizing the first loss function through gradient descent; where the first loss function is: The actor network is updated by maximizing the second loss function using gradient ascent; where the second loss function is: Where L represents the value of the first loss function, N represents the total number of training samples, and y i Let Q represent the target action value, and s represent the action value. i Represents the state data at time i, a i Represents the action at time i; Q(s) i ,a i |θ Q ) represents the parameter θ in the critic network. Q Based on the state data and actions at time i, the action value is given. J(θ) represents the expectation; μ ) represents the value of the second loss function; μ(s) i |θ μ ) represents the parameter θ in the actor network. μ Based on the state data s i The strategy for outputting actions; Q(s) i ,μ(s i |θ μ )|θ Q ) represents the parameter θ in the critic network. Q Based on the state data at time i and μ(s) i |θ μ The value of the action given.
5. The control method for air laser modulation according to claim 1, characterized in that, During the training of the actor-critic network, the parameters of the target actor network and the target critic network are softly updated, using the following formula: i Q′ ←tth Q +(1-τ)θ Q′ ;θ μ′ ←tth μ +(1-τ)θ μ′ ; Where, θ Q The parameters of the critic network are represented by θ. Q′ θ represents the parameters of the target critic network. μ θ represents the parameters of the actor network. μ′ τ represents the parameters of the target actor network; τ is a hyperparameter.
6. An air laser modulation device, characterized in that, The air laser modulation device includes a laser, a beam splitter, a first transmission optical path assembly, a second transmission optical path assembly, a spatial light modulator, a controller, a reflector, a dichroic mirror, an air laser generating assembly, and a spectrometer. The laser is used to output femtosecond laser light to the beam splitter; the beam splitter is used to split the received femtosecond laser light into two beams, and output them to the first transmission optical path component and the second transmission optical path component respectively; wherein, the single pulse energy of the laser in the first transmission optical path component is higher than the single pulse energy of the laser in the second transmission optical path component. The first transmission optical path component is used to perform energy and polarization adjustment on the received laser light before transmitting it to the spatial light modulator; the spatial light modulator is used to determine the phase grayscale driving matrix when receiving the laser light; the controller is used to: determine the optimal actor-critic network determined by the control method of air laser modulation according to any one of claims 1-5, determine the corresponding optimal matrix adjustment parameters according to the phase grayscale driving matrix, and send them to the spatial light modulator; the spatial light modulator is also used to load the optimal matrix adjustment parameters to change the liquid crystal arrangement state, thereby realizing phase modulation of the received laser light; the phase-modulated laser light is reflected to the dichroic mirror via the reflector. The second transmission optical path component is used to perform frequency doubling, filtering, and delay on the received laser before sending it to the dichroic mirror; the dichroic mirror is used to combine the laser from the first transmission optical path component and the laser from the second transmission optical path component to obtain a combined laser beam and send it to the air laser generating component; The air laser generating component is used to react the combined laser beam with nitrogen gas to generate an air laser and focus it onto the spectrometer; the spectrometer is used to collect the spectrum of the air laser and send it to the controller.
7. The air laser modulation device according to claim 6, characterized in that, The first transmission optical path component includes a first half-wave plate, a polarizer, and a second half-wave plate arranged sequentially. The first half-wave plate and the polarizer are used to adjust the single-pulse energy of the laser by changing the optical axis angle and to transmit the resulting linearly polarized light to the second half-wave plate. The second half-wave plate is used to adjust the polarization direction of the received linearly polarized light to horizontal polarization and transmit it to the spatial light modulator.
8. The air laser modulation device according to claim 6, characterized in that, The second transmission optical path assembly includes a frequency doubling crystal, a filter, and a mirror group arranged sequentially. The frequency doubling crystal is used to generate a second harmonic based on the received laser light and send the second harmonic as seed light to the filter; the filter is used to filter and select the frequency of the seed light; the mirror group is used to delay the filtered seed light and emit it to the dichroic mirror.
9. The air laser modulation device according to claim 6, characterized in that, The air laser generating assembly includes a first lens, a vacuum cavity, a second lens, and a third lens arranged in sequence. The first lens is used to focus the received combined laser beam onto the vacuum cavity to generate an air laser within the vacuum cavity; the second lens is used to collimate the air laser output from the vacuum cavity; and the third lens is used to focus the collimated air laser onto the spectrometer.
10. The air laser modulation device according to claim 9, characterized in that, The air laser generating component also includes a filter; the filter is disposed between the third lens and the spectrometer, and the filter is used to filter out stray light in the collimated air laser.
Citation Information
Patent Citations
Test method and test system for acquiring phase modulation curve of spatial light modulator
CN112802154A
Light source device
JP1999101944A