Intracavity laser system performance optimization method and device based on reinforcement learning and storage medium
By applying reinforcement learning-based methods in the intracavity laser system, and using analytical expressions and reinforcement learning algorithms to optimize key parameters, the problems of high computing resource consumption, high parameter adjustment complexity and insufficient real-time adaptability during the optimization process of the intracavity laser system are solved, and efficient and real-time performance optimization effects are achieved.
Patent Information
- Application Number
- CN202510207797.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-23
AI Technical Summary
In the performance optimization process, intracavity laser systems have problems such as high computing resource consumption, high parameter adjustment complexity and insufficient real-time adaptability.
A reinforcement learning-based method is adopted to describe the relationship between the output power of the laser system in the cavity and the key parameters through analytical expressions, and a preset reinforcement learning algorithm is used to optimize the key parameters. This method includes offline initialization training, online real-time optimization and continuous strategy optimization, and uses near-end strategy optimization algorithm and soft update mechanism to achieve efficient optimization of laser system parameters.
It significantly reduces computing resource consumption and optimization time, improves optimization efficiency and real-time, can achieve stable output in a dynamic environment, and improves the reliability and applicability of the laser system.
Smart Images

Figure CN120030907A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the intersection of laser technology and artificial intelligence, and in particular to a method, device and storage medium for optimizing the performance of an intracavity laser system based on reinforcement learning. Background Art
[0002] Intracavity laser systems are widely used in high-precision scenarios such as precision manufacturing, remote sensing mapping, and optical communications. Their performance directly affects the system's output power, field of view (FoV), and three-dimensional positioning accuracy. However, traditional laser system optimization methods mainly rely on empirical debugging or optimization methods based on simple physical models, which have the following problems:
[0003] (1) Parameter space complexity: The performance of intracavity laser systems is affected by multiple parameters (such as pump power, mirror angle, gain medium characteristics, etc.). The coupling relationship between these parameters is complex and difficult to be fully optimized by traditional means.
[0004] (2) Low optimization efficiency: Due to the dynamic behavior and nonlinear characteristics of the laser system, traditional optimization methods often require a large amount of experimental or simulation data, resulting in a time-consuming and resource-intensive optimization process.
[0005] (3) Lack of adaptability: In existing technologies, optimization methods based on fixed models lack flexibility in practical applications and are difficult to cope with environmental changes (such as temperature and vibration) or real-time adjustments to task requirements.
[0006] The potential of artificial intelligence, especially reinforcement learning, in the optimization of complex systems is gradually emerging. Reinforcement learning can use intelligent agents to learn optimal strategies through interaction with the environment, thereby achieving efficient exploration and global optimization of complex parameter spaces. However, its current application in laser systems is still relatively limited, mainly facing the following challenges:
[0007] (1) Sample efficiency problem: High-precision simulation of laser systems is time-consuming, which limits the direct application of reinforcement learning algorithms;
[0008] (2) Real-time requirements: How to control computing resource consumption while meeting real-time optimization requirements. Summary of the invention
[0009] The purpose of the present invention is to overcome the defects of the above-mentioned prior art in the process of intracavity laser system optimization, such as high computing resource consumption, large complexity of parameter adjustment and insufficient real-time adaptability, and to provide a method, device and storage medium for optimizing the performance of the intracavity laser system based on reinforcement learning, which are used to improve the performance of the intracavity laser system, including power output, field of view angle expansion and three-dimensional positioning accuracy.
[0010] The purpose of the present invention can be achieved by the following technical solutions:
[0011] According to a first aspect of the present invention, a method for optimizing the performance of an intracavity laser system based on reinforcement learning is provided, comprising the following steps: optimizing key parameters of the intracavity laser system based on an analytical expression of the output power of the intracavity laser system using a preset reinforcement learning algorithm; wherein the analytical expression is preset, and in the reinforcement learning algorithm, the state is represented by the output power and key input parameter characteristics of the current system, the action is defined as the adjustment of the key input parameters, and the reward function takes the improvement of the output power as the core goal.
[0012] As a preferred technical solution, the key input parameters include pump power and output mirror reflectivity.
[0013] As a preferred technical solution, a preset reinforcement learning algorithm is used to optimize the key parameters of the intracavity laser system. The specific process includes: offline initialization training: using the analytical expression to generate initial training data, and using the preset reinforcement learning algorithm to train the intelligent agent offline, the initial training data includes the output power under different pump power and reflectivity combinations; online real-time optimization: in the actual operating environment, the intelligent agent receives the state feedback of the intracavity laser system in real time, and predicts the optimal action based on the current strategy; continuous strategy optimization: using the experience replay mechanism to store the interaction records of the intelligent agent, including the state, action, reward and new state quadruple for batch training through random sampling of experience data.
[0014] As a preferred technical solution, the reinforcement learning algorithm adopts a proximal strategy optimization algorithm.
[0015] As a preferred technical solution, during the online real-time optimization process, a soft update mechanism is used to smoothly adjust the target network parameters.
[0016] As a preferred technical solution, the analytical expression is used to describe the various parameters of the intracavity laser system and the output power P out The relationship between , the power transfer model is expressed as:
[0017]
[0018] Where P out is the output power, η b is the overlap efficiency of the laser beam and the gain medium, I s is the saturation intensity of the gain medium, R c is the output mirror reflectivity, V is the one-way transmission factor, V s is the loss factor of the gain medium, η e is the excitation efficiency of the gain medium, P in is the pump power, Ag is the cross-sectional area of the gain medium.
[0019] As a preferred technical solution, the one-way transmission factor V is defined as:
[0020] δ=1-|V| 2
[0021] Where δ is the diffraction loss, expressed as:
[0022]
[0023] Where N F is the Fresnel number of the resonant cavity.
[0024] As a preferred technical solution, the reward function is expressed as:
[0025] R=ΔP out -λ E ·P in
[0026] Where R is the reward value, ΔP out is the contribution of the current action to the output power improvement, λ E is the weight of the energy consumption regularization term.
[0027] According to a second aspect of the present invention, there is provided an intracavity laser system performance optimization device based on reinforcement learning, comprising a memory, a processor, and a program stored in the memory, wherein the processor implements the method when executing the program.
[0028] According to a third aspect of the present invention, there is provided a storage medium having a program stored thereon, wherein the program implements the method described above when executed.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] 1. The present invention adopts an analytical expression of the output power of the intracavity laser system to replace the traditional complex simulation model, which can significantly reduce the consumption of computing resources and optimization time. The influence of parameter adjustment on the output power is directly calculated by the analytical expression, which can ensure the real-time and accuracy of the optimization process and greatly improve the optimization efficiency.
[0031] 2. The present invention introduces a deep reinforcement learning algorithm to efficiently optimize key parameters of the laser system (such as pump power, output mirror reflectivity, etc.), which can solve the multi-parameter coupling and nonlinear optimization problems in the laser system. The reinforcement learning agent interacts with the environment and autonomously learns the optimization strategy without relying on manual debugging or prior knowledge, and can achieve efficient adaptive optimization capabilities;
[0032] 3. The present invention adds a regularization term to the optimization target, that is, the reward function. In addition to maximizing the output power, it can also effectively balance energy consumption and system stability, avoiding the adverse consequences caused by over-optimization. By reasonably setting the reward function, the reinforcement learning agent can quickly converge to the global optimal solution while taking into account the physical constraints of the laser system;
[0033] 4. By strengthening the online interactive optimization capability of the learning agent, the present invention can adapt to changes in the dynamic environment (such as vibration, temperature fluctuation, etc.) in real time, ensure the stable output of the laser system under complex working conditions, and thus significantly improve the reliability and applicability of the laser system in high-demand scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A schematic diagram of the structure of a distributed single-cavity laser system in an embodiment of the present invention;
[0035] Figure 2 This is a diagram of a reinforcement learning optimization framework in an embodiment of the present invention;
[0036] Figure 3 This is a flow chart of reinforcement learning training and optimization in an embodiment of the present invention. DETAILED DESCRIPTION
[0037] The present invention provides a method for optimizing the performance of an intracavity laser system based on reinforcement learning. The method is based on an analytical expression of the output power of the intracavity laser system and uses a preset reinforcement learning algorithm to optimize the key parameters of the intracavity laser system; wherein the analytical expression is preset, and in the reinforcement learning algorithm, the state is represented by the output power of the current system and the characteristics of key input parameters, the action is defined as the adjustment of the key input parameters, and the reward function takes the improvement of the output power as the core goal. The method uses a reinforcement learning algorithm to autonomously learn the optimal adjustment strategy through the dynamic interaction between the intelligent agent and the environment, and can effectively balance energy consumption and system stability while greatly improving the output power. In addition, the method has the characteristics of high efficiency and flexibility, and can adapt to dynamic changes in complex working conditions in real time, providing an intelligent and high-performance laser system optimization solution for fields such as precision manufacturing, remote sensing mapping, and optical communications.
[0038] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0039] Example
[0040] like Figure 1As shown, the distributed single-cavity laser system of this embodiment is mainly composed of gain medium, pump light source, output mirror, reflector and feedback sensor. The power of the pump light source is adjustable; the reflectivity of the output mirror and the reflector is adjustable. The feedback sensor monitors the output power and related status information of the intracavity laser system in real time, and uses these data as the basis for optimization control.
[0041] The method for optimizing the performance of an intracavity laser system based on reinforcement learning provided in this embodiment mainly includes two parts: one is to obtain an analytical expression for the output power of the intracavity laser system, and the other is to optimize the key parameters of the intracavity laser system using a preset reinforcement learning algorithm. The specific process is as follows:
[0042] 1. System modeling based on analytical expressions
[0043] In order to reduce the computational overhead in the simulation process, analytical expressions are used to approximate the various parameters and output power (P out ). The power transmission model of the single-cavity system is established based on the following formula:
[0044]
[0045] Where P out is the output power, η b is the overlap efficiency of the laser beam and the gain medium, I s is the saturation intensity of the gain medium, R c is the reflectivity of the output mirror; V is the one-way transmission factor, which is used to describe the diffraction loss in the cavity; V s is the loss factor of the gain medium, η e is the excitation efficiency of the gain medium, P in is the pump power, A g is the cross-sectional area of the gain medium.
[0046] 2. Accurate modeling and fitting of the transmission factor V
[0047] The transmission factor V is affected by the cavity geometry parameters and is defined as:
[0048] δ=1-|V| 2
[0049] Where δ is the diffraction loss, which can be expressed as the Fresnel number N of the resonant cavity. F The relationship between them is generally an empirical formula obtained by theoretical derivation or experimental measurement. In this embodiment, a correction factor is introduced to obtain an accurate analytical expression, and the correction factor can be obtained by fitting simulation data. The specific process is as follows:
[0050] (1) Calculate the Fresnel number N F
[0051] N F The expression is:
[0052]
[0053] Where r is the radius of the aperture (i.e., finite-size element) in the resonant cavity, λ is the wavelength of the laser beam, and L is the cavity length of the resonant cavity.
[0054] (2) Fitting diffraction loss δ
[0055] The formula used in this embodiment is:
[0056]
[0057] According to the simulation results, δ~N F Data, this embodiment adopts nonlinear regression fitting method, and obtains the accurate expression as follows:
[0058]
[0059] 3. Reinforcement Learning Optimization Methods
[0060] This embodiment uses the reinforcement learning algorithm to optimize the key parameters of the laser system based on the analytical expression of the output power of the intracavity laser system. Through the interaction between the intelligent agent and the environment, the input parameters are dynamically adjusted to maximize the output power. Specifically:
[0061] (1) Strengthening learning environment design
[0062] State space:
[0063] The status is represented by the current system output power and key parameter characteristics, including:
[0064] Current output power P out ; Pump power P in ; Output mirror reflectivity R c .
[0065] Action space:
[0066] Actions are defined as adjustments to key input parameters, including:
[0067] Pump power P in Increase or decrease of output mirror reflectivity R c Adjustment
[0068] Reward function:
[0069] The reward function is based on the output power Pout The core goal is to improve the
[0070] P=ΔP out
[0071] Always, ΔP out It is the contribution value of the current action to the output power improvement.
[0072] However, if the goal is only to maximize the output power without considering other physical constraints or system dynamics, it may lead to problems such as excessive energy consumption of the laser system and reduced safety factor. To solve these problems, this embodiment introduces relevant regularization terms into the optimization objective, so that it can take into account energy consumption, stability and convergence while improving performance. When only considering the input power, the reward function can be redefined as:
[0073] R=ΔP out -λ E ·P in
[0074] In the formula, R is the reward value, λ E is the weight of the energy consumption regularization term, which is used to balance output power and energy consumption.
[0075] (2) Algorithm selection
[0076] To achieve efficient optimization, this embodiment uses the Proximal Policy Optimization (PPO) algorithm. The PPO algorithm has the following advantages: it can handle continuous action space and adapt to the continuous adjustment requirements of laser system parameters; it ensures the stability and efficiency of the optimization process through the constraints in the strategy update.
[0077] (3) Training and Optimization Process
[0078] Offline initialization training:
[0079] Initial training data is generated using analytical expressions based on the intracavity laser system, covering the output power at different pump power and reflectivity combinations. The efficient performance of the analytical expression significantly reduces data acquisition time.
[0080] The agent is trained offline using a preset reinforcement learning algorithm (such as PPO). The agent learns the initial strategy through multiple rounds of interaction with the environment simulation, and quickly grasps the influence of parameter adjustment on output power. The agent maximizes the cumulative reward in the offline stage, and initially builds an optimization strategy framework, which provides a basis for subsequent online optimization.
[0081] Online real-time optimization:
[0082] In the actual operation environment, the intelligent agent receives real-time status feedback of the intracavity laser system (such as output power Pout and current parameter values), and predicts the optimal action based on the current strategy. Through real-time interaction with the environment, the agent dynamically adjusts key parameters such as pump power and output mirror reflectivity to ensure that the laser system always maintains the optimal output state.
[0083] Continuous optimization of strategies:
[0084] The experience replay mechanism (Replay Buffer) is used to store the interaction records of the intelligent agent, including the state, action, reward and new state quadruple (s, a, r, s′). Batch training is performed by randomly sampling experience data to avoid time correlation between samples and improve training stability.
[0085] During the online optimization process, a soft update mechanism is used to smoothly adjust the target network parameters to ensure the stability and convergence of the strategy. By continuously optimizing the strategy, the agent gradually approaches the global optimal solution.
[0086] In summary, in order to achieve dynamic optimization of laser system performance, this embodiment proposes an intelligent optimization method based on reinforcement learning. Specifically, the reinforcement learning agent forms a closed-loop control with the laser system through three stages: offline training, online real-time optimization, and continuous strategy optimization, and gradually adjusts key parameters such as pump power and reflectivity to maximize the output power of the system.
[0087] like Figure 2 As shown in the figure, the reinforcement learning optimization framework includes state input module, policy network, action output module and reward function. The current state information of the laser system (such as pump power, reflectivity and output power) is fed back to the state input module through sensors. The policy network predicts the optimal action based on the input state, and the action output module adjusts the laser system parameters accordingly. The reward function comprehensively considers the output power increase and energy consumption changes, and feeds back the optimization results to the intelligent agent to ensure the balance and efficiency of the optimization process.
[0088] like Figure 3 As shown in the figure, the reinforcement learning optimization process is divided into three stages: offline training, online real-time optimization, and continuous strategy optimization. In the offline training stage, multiple sets of simulation data are generated based on the analytical expression of the laser system. These data cover different state combinations such as pump power and reflectivity and their corresponding output power and energy consumption. The policy network and value network are trained using the PPO algorithm, and the initial policy network is output for use in the online stage. In the online real-time optimization stage, the agent predicts the optimal action in real time based on the current state and adjusts the system parameters. Through the feedback of the laser system, the agent continuously corrects the optimization strategy to achieve efficient adjustment in a dynamic environment. In the continuous strategy optimization stage, the agent stores the online interaction data in the experience replay pool, and regularly uses this data to update the policy network and value network to enhance long-term optimization performance.
[0089] Through the above-mentioned method, the intracavity laser system can operate efficiently in a dynamic environment, and its output power is significantly improved. At the same time, the energy consumption of the optimization process is low, and it has strong robustness and practical value.
[0090] Further, this embodiment also provides an intracavity laser system performance optimization device based on reinforcement learning, including a memory, a processor, and a program stored in the memory, and the processor implements the aforementioned method when executing the program. The device processor includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit to a random access memory (RAM). In the RAM, various programs and data required for the operation of the device can also be stored. The CPU, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus. Multiple components in the device are connected to the I / O interface, including: input units, such as keyboards, mice, etc.; output units, such as various types of displays, speakers, etc.; storage units, such as disks, optical disks, etc.; and communication units, such as network cards, modems, wireless communication transceivers, etc. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunication networks. The processing unit performs the various methods and processes described above, such as one or more steps of the aforementioned method. For example, in some embodiments, one or more steps of the foregoing method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via a ROM and / or a communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of the foregoing method may be executed. Alternatively, in other embodiments, the CPU may be configured to execute one or more steps of the foregoing method in any other appropriate manner (e.g., by means of firmware). The functions described above may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0091] Further, the present embodiment also has a storage medium on which a program is stored, and the aforementioned method is implemented when the program is executed. The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or the block diagram to be implemented. The program code can be executed completely on the machine, partially on the machine, partially on the machine as an independent software package and partially on a remote machine or completely on a remote machine or server. In the context of the present invention, a computer-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the above. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk-read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0092] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A method for optimizing the performance of an intracavity laser system based on reinforcement learning, characterized in that: The following steps are involved: Based on the analytical expression of the output power of the intracavity laser system, the key parameters of the intracavity laser system are optimized using the preset reinforcement learning algorithm; Among them, the analytical expression is pre-set, and in the reinforcement learning algorithm, the state is represented by the output power and key input parameter characteristics of the current system, the action is defined as the adjustment of the key input parameters, and the reward function takes the improvement of output power as the core goal.
2. The method for optimizing the performance of an intracavity laser system based on reinforcement learning according to claim 1, characterized in that: The key input parameters include pump power and output mirror reflectivity.
3. The method for optimizing the performance of an intracavity laser system based on reinforcement learning according to claim 1, characterized in that: The key parameters of the intracavity laser system are optimized using a preset reinforcement learning algorithm. The specific process includes: Offline initialization training: using the analytical expression to generate initial training data, and using a preset reinforcement learning algorithm to perform offline training on the intelligent agent, the initial training data includes output power under different pump power and reflectivity combinations; Online real-time optimization: In the actual operating environment, the agent receives real-time feedback on the status of the intracavity laser system and predicts the optimal action based on the current strategy; Continuous strategy optimization: Use the experience replay mechanism to store the agent's interaction records, including the state, action, reward, and new state quadruple, and perform batch training by randomly sampling experience data.
4. The method for optimizing the performance of an intracavity laser system based on reinforcement learning according to claim 3, characterized in that: The reinforcement learning algorithm adopts a proximal strategy optimization algorithm.
5. The method for optimizing the performance of an intracavity laser system based on reinforcement learning according to claim 3, characterized in that: During the online real-time optimization process, a soft update mechanism is used to smoothly adjust the target network parameters.
6. The method for optimizing the performance of an intracavity laser system based on reinforcement learning according to claim 1, characterized in that: The analytical expressions are used to describe the various parameters of the intracavity laser system and the output power P out The relationship between , the power transfer model is expressed as: Where P out is the output power, η b is the overlap efficiency of the laser beam and the gain medium, I s is the saturation intensity of the gain medium, R c is the output mirror reflectivity, V is the one-way transmission factor, V s is the loss factor of the gain medium, η e is the excitation efficiency of the gain medium, P in is the pump power, A g is the cross-sectional area of the gain medium.
7. The method for optimizing the performance of an intracavity laser system based on reinforcement learning according to claim 6, characterized in that: The one-way transmission factor V is defined as: δ=1-|V| 2 Where δ is the diffraction loss, expressed as: Where N F is the Fresnel number of the resonant cavity.
8. The method for optimizing the performance of an intracavity laser system based on reinforcement learning according to claim 7, characterized in that: The reward function is expressed as: R=ΔP out -l E ·P in Where R is the reward value, ΔP out is the contribution of the current action to the output power improvement, λ E is the weight of the energy consumption regularization term.
9. A device for optimizing the performance of an intracavity laser system based on reinforcement learning, comprising a memory, a processor, and a program stored in the memory, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
10. A storage medium having a program stored thereon, characterized in that: When the program is executed, the method according to any one of claims 1 to 8 is implemented.