Spectrally programmable optical frequency comb generation method based on deep reinforcement learning

The spectrally programmable optical frequency comb generation method based on deep reinforcement learning solves the frequency loss and limited tuning range problems of optical frequency comb generation in existing technologies, realizes the spectral programmability and flexible application of the optical frequency comb, and improves the robustness of the system and the accuracy of spectrum generation.

CN117192865BActive Publication Date: 2025-09-05SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311082437.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-28
Publication Date
2025-09-05
Estimated Expiration
2043-08-28

AI Technical Summary

Technical Problem

Existing optical frequency comb generation technology has problems such as frequency loss, phase jitter and limited tuning range, which affect its application in high-precision measurement and optical communications, and cannot achieve accurate control and shaping of the optical frequency comb spectrum.

Method used

A spectrally programmable optical frequency comb generation method based on deep reinforcement learning is adopted. By constructing an interaction model between a deep reinforcement learning agent and the experimental environment, the Actor-Critic architecture is used to design a strategy algorithm and construct a reward function to achieve spectral programmability and spectrum control, and generate a customized spectrum.

Benefits of technology

The spectrum programmability of the optical frequency comb is achieved, its application range is expanded, the robustness and flexibility of the system are improved, the labor cost is reduced, and the accurate generation of the spectrum and the reliability in complex environments are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117192865B_ABST
    Figure CN117192865B_ABST
Patent Text Reader

Abstract

A spectrally programmable optical frequency comb generation system based on deep reinforcement learning is proposed. Based on the experimental physical structure of a broadband optical frequency comb, a deep reinforcement learning agent is deployed and an interaction model between the agent and the experimental environment is constructed. A strategy algorithm framework based on the deep reinforcement learning Actor-Critic architecture is constructed, and the interaction content and rules between the agent and the experimental environment are designed accordingly. A reward function is constructed using the root mean square error between the target spectrum and the experimental spectrum as a parameter, and the action execution and reward feedback between the agent and the experimental environment are designed. The optimal phase modulation decision is obtained through the training strategy of the deep reinforcement learning algorithm, thereby achieving spectrally programmable generation of optical frequency combs. The present invention uses deep reinforcement learning technology to train neural networks, select the optimal phase modulation strategy, and realize spectral programming and control of optical frequency combs. This expands the application range of optical frequency combs and provides greater flexibility for their use in optical communications, precision measurement and other fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of optical signal processing, specifically a method for generating spectrally programmable optical frequency combs based on deep reinforcement learning. Background Art

[0002] Existing optical frequency comb generation technologies include new technologies such as optical frequency combs based on nonlinear optical effects and optical frequency combs based on microwave optical mixing. However, these technologies suffer from problems such as frequency loss, phase jitter, and limited tuning range, which affect their application in high-precision measurement, optical communication and other fields.

[0003] After searching the prior art, Ting Yang et al., in their paper "Comparison Analysis of Optical Frequency Comb Generation with Nonlinear Effects in Highly Nonlinear Fibers," published in Optics Express, Vol. 21, No. 7, 2013, proposed a broadband optical frequency comb generation scheme based on cascaded four-wave mixing and self-phase modulation. This scheme utilizes two cascaded highly nonlinear fibers with different zero-dispersion wavelengths to achieve a 259-line optical frequency comb with a repetition rate of 10 GHz and a flatness within 5 dB. While this scheme enables independent tuning of the repetition rate and center frequency of the optical frequency comb, it does not support control or shaping of the frequency comb's spectrum.

[0004] In the paper "Enhanced nonlinear spectral broadening and sub-picosecond pulse generation by adaptive spectral phase optimization of electro-optic frequency combs" published in Optics Express, Vol. 28, No. 8, 2020, BS Vikram et al. proposed a broadband optical frequency comb generation method based on an adaptive optimization algorithm. This scheme uses a Fourier pulse shaper to adaptively adjust the spectral phase of the electro-optic frequency comb to increase the stimulated Brillouin scattering threshold, thereby increasing the bandwidth of the optical frequency comb by more than 13 times. However, optical frequency comb broadening is a complex nonlinear process. This scheme can only handle relatively vague optimization goals and cannot accurately control the spectrum of the optical frequency comb. Summary of the Invention

[0005] This paper addresses the limitations of existing technologies, such as the difficulty in fully fabricating target optical frequency combs based on simulation data and the inability of optical frequency combs fabricated using microring resonator platforms to dynamically control the spectrum in real time. This paper proposes a spectrally programmable optical frequency comb generation system based on deep reinforcement learning. This system uses deep reinforcement learning to train a neural network, select the optimal phase modulation strategy, and implement spectral programming and control of the optical frequency comb. This system expands the application range of optical frequency combs and provides greater flexibility for their use in optical communications, precision measurement, and other fields.

[0006] The present invention is achieved through the following technical solutions:

[0007] The present invention relates to a spectrally programmable optical frequency comb generation method based on deep reinforcement learning. The method deploys a deep reinforcement learning agent based on the experimental physical structure of a broadband optical frequency comb and constructs an interaction model between the agent and the experimental environment. A strategy algorithm framework based on the deep reinforcement learning actor-critic architecture is constructed, and the interaction content and rules between the agent and the experimental environment are designed accordingly. A reward function is constructed using the root mean square error between the target spectrum and the experimental spectrum as a parameter, and the action execution and reward feedback between the agent and the experimental environment are designed. The optimal phase modulation decision is obtained through the training strategy of the deep reinforcement learning algorithm, thereby realizing spectrally programmable generation of an optical frequency comb.

[0008] Technical Effects

[0009] The present invention uses deep reinforcement learning to construct an interaction model between the intelligent agent and the optical frequency comb generation system environment, broadens the initial optical frequency comb through the nonlinear effect of highly nonlinear optical fiber, and designs a training strategy for the deep reinforcement learning algorithm to construct a reward function, thereby outputting the optimal phase modulation decision. The present invention applies deep reinforcement learning to the control and shaping of the optical frequency comb in a closed-loop optical system, thereby realizing a spectrum-programmable broadband optical frequency comb. Compared with the existing technology, the technical effect of the present invention is significantly superior in many aspects: First, the spectral programmability enables the system to generate customized spectra according to specific needs, further expanding the application field of broadband optical frequency combs. Second, through the training of the intelligent agent based on the experimental environment, the system's optimization decisions are robust and not affected by environmental noise, thereby ensuring the reliability of the system in complex environments. In addition, through the optimization decision of the intelligent agent, the system can achieve spectrum shaping and control of the optical frequency comb under the premise that the experimental structure remains unchanged, significantly reducing labor costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 System block diagram of the present invention;

[0011] Figure 2 Flowchart of the embodiment;

[0012] Figure 3A diagram of the deep reinforcement learning network structure of an embodiment. DETAILED DESCRIPTION

[0013] like Figure 1 As shown, this embodiment relates to a spectrum-designable optical frequency comb generation system based on machine learning, including: a laser, a polarization controller, an intensity modulator, a phase modulator, a programmable optical processor, a fiber amplifier, a nonlinear optical fiber and a spectrometer connected in sequence, wherein: the intensity modulator and the phase modulator control the modulation parameters through the radio frequency source and the phase shifter, the spectrometer generates the spectrum initial state and the spectrum execution state according to the collected broadened spectrum information, and the phase of the programmable optical processor is controlled by the intelligent agent.

[0014] like Figure 2 As shown, this embodiment relates to a method for generating an optical frequency comb based on the above system, including:

[0015] Step A: Construct a stable electro-optical frequency comb and control the spectral phase, using nonlinear effects to broaden the optical frequency comb spectrum. Specifically, it includes:

[0016] A1. Generate a stable initial electro-optical frequency comb using a single-wavelength laser, a cascaded intensity modulator, and two phase modulators. The single-wavelength laser outputs input light with a wavelength of 1550 nm and a power of 10 dBm. The generated initial optical frequency comb has a center frequency of 193.548 THz and a repetition rate of 10 GHz.

[0017] Preferably, a DC bias voltage is applied to the intensity modulator to make it operate at a quadrature point, thereby generating an initial optical frequency comb with a flat spectrum and more comb teeth.

[0018] A2. Modulate the phase of the initial optical frequency comb using a programmable optical processor: The phase within an 8nm bandwidth near the center frequency of the initial optical frequency comb, i.e., from 1546nm to 1554nm, is selected as the modulation object. The phase modulation curve is represented by the weighted sum of 20th-order Chebyshev polynomials with random weights, specifically: the Chebyshev polynomial function T with an nth power n (x) = cos(ncos -1 (x)), Where: w k represents the Chebyshev polynomial T k The weight of (x), W 20 (x) represents the result of adding the Chebyshev polynomials according to the weights.

[0019] A3. The modulated electro-optical frequency comb is input into an erbium-doped fiber amplifier and injected into a highly nonlinear optical fiber, and the nonlinear effect is used to broaden the optical frequency comb spectrum.

[0020] Preferably, the modulated initial optical frequency comb is amplified to 23 dBm and then injected into a 2 km long high nonlinear optical fiber for nonlinear broadening, and its nonlinear parameter is 0.03 ps / (nm 2 km).

[0021] A4. Use a spectrometer to collect broadened spectrum information and analyze it to obtain the initial state and execution state of the spectrum.

[0022] Step B: Based on the physical structure of the broadband optical frequency comb obtained in the experiment, i.e., the initial state of the spectrum and the execution state of the spectrum obtained in step A, a deep reinforcement learning agent is deployed to establish an interaction model between the agent and the experimental environment. Figure 3 As shown in the figure, a policy algorithm framework based on the deep reinforcement learning actor-critic architecture is established, and the interaction content and rules between the deep reinforcement learning agent and the experimental environment module are designed, including:

[0023] B1. Construct a deep reinforcement learning agent module based on the Actor-Critic architecture, in which: the Actor network generates action strategies, such as output phase modulation decisions, and updates the action strategies based on the value function Q feedback provided by the Critic network to improve the agent's performance in the environment; the Critic network evaluates the quality of the action strategies and calculates the value function Q based on the actions taken by the agent to evaluate the quality of the current action strategies, and stores the state transition process after the Actor network interacts with the environment in the experience replay memory.

[0024] B2. Train the agent through random sampling of network experience replay memory, specifically including:

[0025] ① Considering that the root mean square error (RMSE) between the target spectrum and the experimental spectrum is a key indicator for evaluating the control and shaping effect of the optical frequency comb, the reward function R is set t =-RMSE(S target , S exp ), where: target spectrum S target , experimental spectrum S exp , the negative sign is to transform RMSE into a maximization problem, so that the task of the agent becomes to minimize the difference between the target spectrum and the experimental spectrum;

[0026] ②Set the agent's goal to learn the optimal policy function π(a t |s t ,θ), where: a t is the action generated by the agent through the Actor network, s t is the state of the experimental environment, θ is the parameter of the Actor network; the policy gradient in the Actor network Where: A(s t,a t ) is the advantage function of the Critic network; calculate the loss function Where: E represents the expected value, KL represents the KL divergence, θ old Represents the old Actor network parameters, and λ is a hyperparameter used to control the balance between policy update and policy stability;

[0027] ③ Evaluate the quality of the action strategy generated by the Actor network through the Critic network: Set the value function Q(s) of the Critic network t ,a t ), that is, in state s t Take action t The expected return is updated using the TD error-based method, and the TD error with the advantage function is: δ t =r t +γQ(s t+1 ,a t+1 )-Q(s t ,a t ), advantage function A(s t ,a t )=Q(s t ,a t )-V(s t ), where: V(s t ) is the state value function of the Critic network, and the advantage function is used to reduce the variance and improve the convergence speed of the algorithm.

[0028] B3. At each time step, a small Gaussian noise is added to the output of the Actor network to encourage the agent to explore new action strategies and avoid over-reliance on past experience. Specifically, for the generated action a t , the added Gaussian noise follows a Gaussian distribution with a mean of zero and a standard deviation of σ: a' t =a t +N(0,σ), where: a t represents the original action generated by the agent at time step t, that is, the phase modulation decision, which is the output of the Actor network; a' t Represents the action after adding Gaussian noise, which will be used to execute in the actual environment; N(0,σ) represents a Gaussian distribution with mean zero and standard deviation σ.

[0029] The agent model is trained to obtain the optimal strategy function π(a t |s t ,θ). Based on the optimal phase control decision output by the intelligent agent, the broadened optical frequency comb obtained by passing the modulated initial optical frequency comb through the highly nonlinear fiber will have the target spectrum defined by the user.

[0030] Through specific practical experiments, this embodiment of the spectrally programmable optical frequency comb generation system achieves user-defined target spectra. The achievable spectrum types can be described by Gaussian functions and Gaussian mixture models with different characteristic parameters: Where: S(f) represents the target spectrum to be generated, N represents the number of Gaussian functions, A i 、f i and σ i represent the amplitude, center frequency, and standard deviation of the i-th Gaussian function respectively.

[0031] This invention utilizes a deep reinforcement learning (DRL) interaction model between an intelligent agent and an experimental environment, combining the nonlinear effects of highly nonlinear optical fibers with intelligent phase control to achieve precise control of the optical frequency comb's broadened spectrum. Compared to existing methods, this invention not only handles complex nonlinear processes but is also unaffected by process errors in the optical frequency comb's manufacturing process, ensuring accurate generation of the target spectrum. Secondly, this invention employs a policy algorithm framework based on a deep reinforcement learning (DRL) actor-critic architecture. By designing the interaction content and rules between the agent and the experimental environment, the system enables real-time dynamic spectrum control during closed-loop experiments. Compared to existing technologies, this invention can flexibly shape the spectrum according to demand, achieving customized spectrum generation, thereby expanding the application scope of optical frequency combs. Furthermore, by constructing a reward function and designing the action execution and reward feedback between the agent and the experimental environment, this invention effectively optimizes the training process of phase modulation decisions and realizes spectral programming and control of the optical frequency comb. Compared to existing technologies, the DRL algorithm employed in this invention can more precisely adjust spectral parameters, ensuring that the performance parameters of the optical frequency comb remain stable even in complex environments, improving the robustness and reliability of the system.

[0032] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principles and purpose of the present invention. The scope of protection of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. All implementation schemes within its scope shall be subject to the constraints of the present invention.

Claims

1. A spectrally programmable optical frequency comb generation method based on deep reinforcement learning, characterized in that: Based on the experimental physical structure of a broadband optical frequency comb, a deep reinforcement learning agent is deployed and an interaction model between the agent and the experimental environment is constructed. A policy algorithm framework based on the deep reinforcement learning actor-critic architecture is constructed, and the interaction content and rules between the agent and the experimental environment are designed accordingly. A reward function is constructed using the root mean square error between the target spectrum and the experimental spectrum as a parameter, and the action execution and reward feedback between the agent and the experimental environment are designed. The optimal phase modulation decision is obtained through the training strategy of the deep reinforcement learning algorithm, thereby realizing spectrally programmable generation of optical frequency combs. Specifically, Step A: Construct a stable electro-optical frequency comb and control the spectral phase, using nonlinear effects to broaden the optical frequency comb spectrum. Specifically, it includes: A1. Generate a stable initial electro-optical frequency comb using a single-wavelength laser, a cascaded intensity modulator, and two phase modulators. The single-wavelength laser outputs input light at a wavelength of 1550 nm and a power of 10 dBm. The generated initial optical frequency comb has a center frequency of 193.548 THz and a repetition rate of 10 GHz. A2. Modulate the phase of the initial optical frequency comb using a programmable optical processor: The phase within an 8 nm bandwidth near the center frequency of the initial optical frequency comb, i.e., from 1546 nm to 1554 nm, is selected as the modulation object. The phase modulation curve is represented by the weighted sum of 20th-order Chebyshev polynomials with random weights, specifically: a Chebyshev polynomial function with an nth power , , where: w k represents the Chebyshev polynomial T k The weight of (x), W 20 (x) represents the result of adding Chebyshev polynomials according to weights; A3. Input the modulated electro-optical frequency comb into an erbium-doped fiber amplifier and inject it into a highly nonlinear fiber, utilizing the nonlinear effect to broaden the optical frequency comb spectrum. A4. Using a spectrometer to collect broadened spectrum information, analyze and obtain the spectrum initial state and spectrum execution state; Step B: Deploy a deep reinforcement learning agent based on the physical structure of the broadband optical frequency comb obtained experimentally, namely the spectral initial state and spectral execution state obtained in Step A. Establish a policy algorithm framework based on the deep reinforcement learning actor-critic architecture, and design the interaction content and rules between the deep reinforcement learning agent and the experimental environment module. Specifically, B1. Construct a deep reinforcement learning agent module based on the actor-critic architecture, in which: the actor network generates action strategies, such as output phase modulation decisions, and updates the action strategies based on the value function Q feedback provided by the critic network to improve the agent's performance in the environment; the critic network evaluates the quality of the action strategies and calculates the value function Q based on the actions taken by the agent to evaluate the quality of the current action strategy, and stores the state transition process after the actor network interacts with the environment in the experience replay memory; B2. Train the agent through random sampling of network experience replay memory, specifically including: ① Considering that the root mean square error (RMSE) between the target spectrum and the experimental spectrum is a key indicator for evaluating the control and shaping effect of the optical frequency comb, the reward function R is set t = - RMSE(S target , S exp ), where: target spectrum S target , experimental spectrum S exp , the negative sign is to transform RMSE into a maximization problem, so that the agent's task becomes to minimize the difference between the target spectrum and the experimental spectrum; ②Set the agent's goal to learn the optimal policy function π(a t |s t , θ), where: a t is the action generated by the agent through the Actor network, s t is the state of the experimental environment, θ is the parameter of the Actor network; the policy gradient in the Actor network , where: A(s t , a t ) is the advantage function of the Critic network; calculate the loss function , where: E represents the expected value, KL represents the KL divergence, θ old Represents the old Actor network parameters, and λ is a hyperparameter used to control the balance between policy update and policy stability; ③ Evaluate the quality of the action strategy generated by the Actor network through the Critic network: Set the value function Q(s) of the Critic network t , a t ), that is, in state s t Take action t The expected return is updated using the TD error-based method. The TD error with the advantage function is: , advantage function , where: V(s t ) is the state value function of the Critic network, which uses the advantage function to reduce the variance and improve the convergence speed of the algorithm; B3. At each time step, a small Gaussian noise is added to the output of the Actor network to encourage the agent to explore new action strategies and avoid over-reliance on past experience; specifically, for the generated action a t , the added Gaussian noise follows a Gaussian distribution with a mean of zero and a standard deviation of σ: , where: a t represents the original action generated by the agent at time step t, that is, the phase modulation decision, which is the output of the Actor network; a' t Represents the action after adding Gaussian noise, which will be used to execute in the actual environment; N(0, σ) represents a Gaussian distribution with mean zero and standard deviation σ.

2. The method for generating spectrally programmable optical frequency combs based on deep reinforcement learning according to claim 1, wherein: By loading a DC bias voltage on the intensity modulator to make it work at the quadrature point, an initial optical frequency comb with a flat spectrum and more comb teeth is generated.

3. The method for generating spectrally programmable optical frequency combs based on deep reinforcement learning according to claim 1, wherein: The modulated initial optical frequency comb is amplified to 23 dBm and then injected into a 2 km long high nonlinear fiber for nonlinear broadening. Its nonlinear parameter is 0.03 ps / (nm 2 km).

4. The method for generating spectrally programmable optical frequency combs based on deep reinforcement learning according to claim 1, wherein: The update of the Critic network uses mean square error as the loss function: .

5. The method for generating spectrally programmable optical frequency combs based on deep reinforcement learning according to claim 1, wherein: The agent model is trained to obtain the optimal strategy function π(a t |s t , θ), according to the optimal phase control decision output by the intelligent agent, the broadened optical frequency comb obtained by modulating the initial optical frequency comb through the highly nonlinear optical fiber will have the target spectrum defined by the user programming.

6. A spectrum-programmable optical frequency comb generation system for implementing the spectrum-programmable optical frequency comb generation method based on deep reinforcement learning as described in any one of claims 1 to 5, characterized in that: include: The laser, polarization controller, intensity modulator, phase modulator, programmable optical processor, erbium-doped fiber amplifier, nonlinear optical fiber and spectrometer are connected in sequence, wherein: the intensity modulator and phase modulator control the modulation parameters through the radio frequency source and phase shifter, the spectrometer generates the spectrum initial state and spectrum execution state according to the collected broadened spectrum information, and the phase of the programmable optical processor is controlled by the intelligent agent.

Citation Information

Patent Citations

  • Optical frequency comb performance analysis method and system based on machine learning

    CN114218834A

  • All-optical convolver based on double combs

    CN115906977A