Nonlinear optical pulse shaping parameter optimization method and device based on neural network
By establishing a pulse shaping parameter inversion model based on neural network, the complex parameter optimization problem in nonlinear optical pulse shaping is solved, efficient and accurate pulse shaping is achieved, and the cycle and cost of the shaping process is reduced.
Patent Information
- Application Number
- CN202411900590.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively optimize parameters in nonlinear optical pulse shaping, resulting in complex shaping process, high cost, difficult to control operation, and low power efficiency.
Using a neural network-based method, a pulse shaping parameter inversion model is established, and the optimal input parameters are inverted by training the model using the target pulse waveform to achieve multi-parameter optimization of nonlinear pulse shaping.
The nonlinear plastic surgery process is simplified, transformed into numerical optimization problems, reduced the cycle and cost of the plastic surgery process, improved the generalization ability and reliability of the model, and ensured the efficiency and accuracy of pulse plastic surgery.
Smart Images

Figure CN120068574A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of laser shaping, and more specifically, relates to a method and device for optimizing non-linear optical pulse shaping parameters based on a neural network. Background Art
[0002] Currently, for time-domain pulse shaping, well-known Gaussian and sech-shaped pulses are generally shaped into more peculiar flat-top, parabolic, or triangular pulses. Among them, the most commonly used is the simple linear filtering technique. Limited by electronic components, so far, there is no time-domain filter on an ultrafast time scale and it is impossible to shape ultrafast time-domain pulses. Therefore, indirect pulse shaping in the frequency domain has become an effective means.
[0003] Based on the 4F system, laser pulse shaping is achieved through Fourier transform. Traditional spatial masks, liquid crystal spatial light modulators, liquid crystal on silicon spatial light modulators, acousto-optic spatial light modulators, digital micromirror devices (DMDs), etc. have been used as spatial filters. The latter three methods allow programmability and adaptability of the pulse shaping response. The disadvantages of this method are obvious, such as high cost, large volume, complex structure, difficult to operate, and unable to be integrated with waveguides. These problems have prompted people to seek other technical methods, such as arrayed waveguide gratings, fiber delay line arrays, temporal coherent synthesis, FWM time lens technology, dispersion Fourier transform (DFT), superstructure fiber Bragg gratings (SSFBG), etc. It has been confirmed that by using a multi-arm interferometer and combining temporal coherent synthesis (TCS) technology, various transform-limited pulse shapes, such as flat-top, parabolic, and triangular pulses, can be generated. Using the superstructure fiber Bragg grating technology, flat-top, parabolic, and sawtooth (asymmetric triangular) pulses can be generated. Among them, the superstructure fiber Bragg grating depends on the design and fabrication of the grating; temporal coherent synthesis and FWM time lens also have complex structures and are difficult to operate. The methods mentioned above all belong to linear pulse shaping, and their inherent disadvantage is that the bandwidth of the output spectrum is determined by the bandwidth of the input spectrum, and the power efficiency is low during the shaping process. Some people have also mentioned that pulse shaping is achieved through the interaction between pulse pre-chirping and non-linear transmission in a positive dispersion fiber. This method that combines the third-order non-linear process and dispersion provides an effective new solution to overcome the disadvantages of linear pulse shaping.
[0004] However, determining the optimal parameters of a non-linear fiber system to achieve desired pulse characteristics is more complex than linear pulse shaping, which only requires the input waveform and the target waveform. In fact, the non-linear pulse shaping effect depends on the input pulse conditions and fiber characteristics. Summary of the Invention
[0005] In view of the above deficiencies or improvement requirements of the prior art, the present invention provides a method and device for optimizing non - linear optical pulse shaping parameters based on a neural network. A method combining the third - order non - linear process and dispersion of optical fibers is used for non - linear pulse shaping. Taking the target pulse waveform as the input and the input pulse conditions and fiber characteristics as the output, an inverse model of pulse shaping parameters based on a neural network is established and trained. After training, the target pulse waveform is used as the input to inversely deduce the optimal values of the input parameters for pulse shaping, realizing the multi - parameter optimization of non - linear pulse shaping.
[0006] To achieve the above object, according to the first aspect of the embodiments of the present invention, a method for optimizing non - linear optical pulse shaping parameters based on a neural network is provided, including the following steps:
[0007] S100. Perform time - domain waveform shaping on the pulsed laser output by the laser;
[0008] S200. Measure the shaped time - domain waveform;
[0009] S300. Collect the time - domain waveform data after shaping;
[0010] S400. Based on the inverse model of pulse shaping parameters based on a neural network, use sample data to train the model;
[0011] S500. Take the target pulse waveform as the input and substitute it into the inverse model of pulse shaping parameters to inversely deduce the optimal values of the pulse shaping parameters;
[0012] In step S400, it specifically includes the following steps:
[0013] S410. Collect the variables in the input parameters during shaping in step S100 as feature data, and collect the time - domain distributions after shaping under different feature data as label data, and organize them into a data set;
[0014] S420. Build an inverse model of pulse shaping parameters based on a neural network, denoted as the NNs model;
[0015] S430. Divide the data in the data set into two parts, a large part and a small part, and use most of the data as sample data to train the NNs model;
[0016] S440. Select several groups of feature data from the remaining small part of the data and substitute them into the NNs model for verification.
[0017] Further, in step S410, it specifically further includes the following steps:
[0018] S411. Systematically collect input parameters including the output power of the laser under different conditions and the settings of the non - linear shaping module, as well as the corresponding shaped time - domain waveform data, and record the influencing factors including experimental environmental conditions at the same time;
[0019] S412. Identify and quantify the key factors affecting the pulse shaping effect, construct features using physical prior knowledge, and apply feature selection algorithms to screen the most influential features, reducing redundancy and improving the model efficiency;
[0020] S413. Use data augmentation techniques to generate more training samples. For scarce but important cases, adopt oversampling or synthetic minority over - sampling techniques to increase their representativeness.
[0021] Further, in step S411, considering that the laser output power and environmental conditions may change over time, these variables can be regarded as random processes. For the laser output power P(t), it is described by a geometric Brownian motion with drift and diffusion terms, specifically:
[0022] dP(t) = μ(P, t)dt + σ(P, t)dWt,
[0023] where μ(P, t) is the drift term,
[0024] σ(P, t) is the diffusion coefficient,
[0025] W t is a Wiener process.
[0026] Further, in step S412, use the Bayesian framework for feature selection, and evaluate the importance of features by calculating the posterior probability. Given the dataset D, the posterior probability of feature F i is obtained through Bayes' theorem, and its posterior distribution P(F i |D) is:
[0027]
[0028] where P(D|F i ) is the likelihood function,
[0029] P(F i ) is the prior distribution,
[0030] P(D) is the evidence.
[0031] Further, in step S420, add a layer based on physical equations at the front end and an interpretation layer at the end of the grid to combine the advantages of traditional physical models and deep learning;
[0032] Accordingly, the neural network architecture in the NNs model sequentially includes: an input layer, a physics-aware layer, a deep learning layer, an interpretation layer, and an output layer.
[0033] Further, the physics-aware layer directly receives the input feature F i , and calculates the intermediate variable Y phy , and the intermediate variable Y phy includes spectral width and peak power, specifically:
[0034] Y phy = f phy (F i ; θ phy ),
[0035] where θ phy is the relevant physical parameter,
[0036] f phy (·) is the NLS equation;
[0037] The NLS equation f phy (·) is specifically:
[0038]
[0039] where A(z, t) is the normalized electric field amplitude,
[0040] z is the propagation distance,
[0041] t is the time;
[0042] The deep learning layer includes a hidden layer and a residual connection layer. The hidden layer includes multiple fully connected layers or convolutional layers for capturing complex non-linear mapping relationships. Each layer is expressed as:
[0043] Z (k) = σ(W (k) Z (k-1) + b (k) ),
[0044] where k represents the k-th layer,
[0045] W (k) is the weight matrix,
[0046] b (k) is the bias vector,
[0047] σ(·) is the activation function,
[0048] Z (0) = [Y phy , F], and F is the original input feature;
[0049] The residual connection layer helps with gradient transmission and prevents the problem of vanishing gradients in deep networks, specifically as follows:
[0050] Z (k) =Z (k-1) +σ(W (k) Z (k-1) +b (k) );
[0051] The interpretation layer quantifies the importance score S of each input feature through SHAP values j , specifically as follows:
[0052] S j =φ j (Z (K) ),
[0053] where S j is the importance score of the j-th feature,
[0054] is the interpretation function,
[0055] K is the index of the last layer.
[0056] Furthermore, in step S430, it specifically further includes the following steps:
[0057] S431: Divide the dataset into a training set and a validation set to ensure that each part is independent and has a consistent distribution;
[0058] S432: Use the data in the training set to implement batch gradient descent for model training, use the data in the validation set for cross-validation to further stabilize the performance of the NNs model, introduce an early stopping mechanism, and terminate training early when the performance on the validation set no longer improves to prevent overfitting;
[0059] S433: Use Bayesian optimization to systematically adjust the optimal hyperparameter configuration, and use automated machine learning tools to simplify the process of adjusting the optimal hyperparameter configuration;
[0060] S434: Apply regularization techniques to ensure the generalization ability of the model;
[0061] S435: Use transfer learning to preload some weights to accelerate the convergence speed of the model in the new environment.
[0062] Furthermore, in step S440, it specifically further includes the following steps:
[0063] S441: Randomly select unused data from the validation set as the test set, use the test set to evaluate the true performance of the NNs model, and pay attention to the evaluation criteria specific to the optical pulse shaping task;
[0064] S442. Apply XAI technology to analyze the decision-making path of the NNs model, provide intuitive understanding, and improve the experimental data based on the obtained information.
[0065] Further, in step S100, it specifically further includes the following steps:
[0066] S110. Use a Gaussian filter to make the laser pulse waveform emitted by the laser be Gaussian-shaped;
[0067] S120. Use a pre-chirp generator to generate a Gaussian-shaped chirped pulse;
[0068] S130. Adjust the output average power of the pre-chirp generator through an adjustable attenuator, and then adjust the peak power of the pulsed laser;
[0069] S140. Provide strong nonlinearity through a zero-dispersion fiber, and use the SPM effect to stretch and shape the spectrum;
[0070] S150. Provide a large dispersion amount through a large-dispersion fiber, and obtain a pulse waveform consistent with the spectrum shape through the DFT effect.
[0071] According to the second aspect of the embodiments of the present invention, there is provided a device for optimizing nonlinear optical pulse shaping parameters based on a neural network, including: a laser, a nonlinear shaping module provided at the rear end of the laser, a photodetector provided at the rear end of the nonlinear shaping module, an oscilloscope provided at the rear end of the photodetector, and a computer;
[0072] The laser has a pigtail and outputs pulsed laser light;
[0073] The nonlinear shaping module is used for time-domain waveform shaping of the pulses output by the laser, and includes a Gaussian filter, a pre-chirp generator, an adjustable attenuator, a zero-dispersion fiber, and a large-dispersion fiber;
[0074] The photodetector and the oscilloscope are used to measure the shaped time-domain waveform;
[0075] The computer is connected to the signal output end of the oscilloscope and is signal-connected between the pre-chirp generator, the adjustable attenuator, the zero-dispersion fiber, and the large-dispersion fiber.
[0076] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0077] 1. The pulse shaping parameter optimization method of the present invention uses a method combining the fiber third-order nonlinear process and dispersion for nonlinear pulse shaping. Taking the target pulse waveform as the input and the input pulse conditions and fiber characteristics as the output, an inverse model of pulse shaping parameters based on a neural network is established, and this model is trained. After the training is completed, the target pulse waveform is used as the input to inversely deduce the optimal values of the input parameters for pulse shaping, realizing the multi-parameter optimization of nonlinear pulse shaping.
[0078] 2. The pulse shaping parameter optimization method of the present invention simplifies the nonlinear shaping process into a numerical optimization problem in three-dimensional space, effectively solving problems such as long period, high cost, and complex operation in multi-parameter optimization of pulse shaping.
[0079] 3. The pulse shaping parameter optimization method of the present invention uses an evaluation criterion specific to the optical pulse shaping task, ensuring the true performance of the model, providing valuable information about the model's generalization ability, effectively reducing the risk of overfitting, improving the prediction accuracy and stability of the model on new data, and making the model more reliable in practical applications.
[0080] 4. The pulse shaping parameter optimization method of the present invention not only improves the generalization ability and reliability of the model, but also enhances the interpretability and transparency of the model, optimizes the quality of experimental data, and accelerates the iteration and innovation process. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 It is a schematic flowchart of a method for optimizing nonlinear optical pulse shaping parameters based on a neural network according to an embodiment of the present invention;
[0082] Figure 2 It is a schematic diagram of the steps of a method for optimizing nonlinear optical pulse shaping parameters based on a neural network according to an embodiment of the present invention;
[0083] Figure 3 It is a schematic diagram of the specific steps of step S100 in a method for optimizing nonlinear optical pulse shaping parameters based on a neural network according to an embodiment of the present invention;
[0084] Figure 4 It is a schematic diagram of the specific steps of step S400 in a method for optimizing nonlinear optical pulse shaping parameters based on a neural network according to an embodiment of the present invention;
[0085] Figure 5 It is a schematic diagram of the specific steps of step S410 in a method for optimizing nonlinear optical pulse shaping parameters based on a neural network according to an embodiment of the present invention;
[0086] Figure 6 It is a schematic diagram of the specific steps of step S430 in a method for optimizing nonlinear optical pulse shaping parameters based on a neural network according to an embodiment of the present invention;
[0087] Figure 7 This is a schematic diagram of the specific steps of step S440 in an optimization method for non - linear optical pulse shaping parameters based on a neural network according to an embodiment of the present invention;
[0088] Figure 8 This is a schematic structural diagram of an optimization device for non - linear optical pulse shaping parameters based on a neural network according to an embodiment of the present invention.
[0089] In all the drawings, the same reference numerals represent the same technical features. Specifically: 1 - laser, 2 - non - linear shaping module, 21 - Gaussian filter, 22 - pre - chirp generator, 23 - adjustable attenuator, 24 - zero - dispersion fiber, 25 - large - dispersion fiber, 3 - photodetector, 4 - oscilloscope, 5 - computer. Detailed implementation manners
[0090] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0091] Embodiment 1
[0092] As Figure 1 、 2 shown, an embodiment of the present invention provides an optimization method for non - linear optical pulse shaping parameters based on a neural network, including the following steps:
[0093] S100. Use the non - linear shaping module 2 to perform time - domain waveform shaping on the pulsed laser output by the laser 1;
[0094] S200. Use the photodetector 3 and the oscilloscope 4 to measure the shaped time - domain waveform;
[0095] S300. Collect the data of the shaped time - domain waveform through the computer 5;
[0096] S400. In the computer 5, based on the pulse shaping parameter inversion model of the neural network, use the sample data to train the model;
[0097] S500. Take the target pulse waveform as the input and bring it into the pulse shaping parameter inversion model to invert the optimal value of the pulse shaping parameters.
[0098] In step S100, in the non-linear shaping module 2, the output pulse time domain distribution is such that the second-order dispersion coefficient, non-linear coefficient of the three-section optical fiber, and the bandwidth of the Gaussian filter 21 are all fixed values. There are a total of 5 parameter variables in the non-linear shaping process, namely the output peak power after the adjustable attenuator 23, the pulse width at 1 / e, the optical fiber length L1 of the adjustable attenuator 23, the optical fiber length L2 of the zero-dispersion optical fiber 24, and the optical fiber length L3 of the large-dispersion optical fiber 25.
[0099] As Figure 3 shown, in step S100, it specifically further includes the following steps:
[0100] S110. Use the Gaussian filter 21 to make the laser pulse waveform emitted by the laser 1 be Gaussian-shaped;
[0101] S120. Use the pre-chirp generator 22 to generate Gaussian-shaped chirped pulses;
[0102] S130. Adjust the output average power of the pre-chirp generator 22 through the adjustable attenuator 23, and thus adjust the peak power of the pulsed laser;
[0103] S140. Provide strong non-linearity through the zero-dispersion optical fiber 24, and use the SPM effect to stretch and shape the spectrum;
[0104] S150. Provide a large dispersion amount through the large-dispersion optical fiber 25, and obtain a pulse waveform consistent with the spectrum shape through the DFT effect.
[0105] As Figure 4 shown, in step S400, it specifically includes the following steps:
[0106] S410. Collect the variables in the input parameters during shaping in step S100 as feature data, collect the time domain distributions after shaping under different feature data as label data, and organize and generate a data set;
[0107] S420. Build an inverse model of pulse shaping parameters based on a neural network, denoted as the NNs model;
[0108] S430. Use 80% of the data in the data set as sample data to train the NNs model;
[0109] S440. Select several groups of feature data from the remaining 20% of the data, and substitute them into the NNs model for verification.
[0110] As Figure 5 shown, in step S410, it specifically further includes the following steps:
[0111] S411. Systematically collect input parameters including the output power of the laser 1 and the settings of the non - linear shaping module 2 under different conditions, as well as the corresponding shaped time - domain waveform data, and simultaneously record influencing factors including experimental environmental conditions (such as temperature, humidity);
[0112] S412. Identify and quantify the key factors affecting the pulse shaping effect, construct features using physical prior knowledge, and apply feature selection algorithms to screen the most influential features, reducing redundancy and improving the model efficiency;
[0113] S413. Use data augmentation techniques to generate more training samples. For rare but important cases, adopt oversampling or synthetic minority over - sampling techniques to increase their representativeness.
[0114] In step S411, considering that the laser output power and environmental conditions may vary with time, these variables can be regarded as random processes. For the laser output power P(t), it is described by a geometric Brownian motion with drift and diffusion terms, specifically:
[0115] dP(t) = μ(P, t)dt + σ(P, t)dW t ,
[0116] where μ(P, t) is the drift term,
[0117] σ(P, t) is the diffusion coefficient,
[0118] W t is a Wiener process.
[0119] The drift term μ(P, t) represents the trend change of the laser output power P(t) with time t, depending on the internal physical mechanism of the laser, external environmental conditions, and operating parameters, specifically:
[0120] μ(P, t) = αP(t) + βf env (t) + g op (P, t),
[0121] where α is the inherent attenuation or gain coefficient inside the laser,
[0122] β is the environmental coefficient,
[0123] f env (t) is the influence of environmental factors,
[0124] g op (P, t) is the influence of operating parameters on the output power.
[0125] The diffusion coefficient σ(P, t) describes the intensity of random fluctuations and is related to the noise source, specifically:
[0126]
[0127] Among them, δ is the noise level proportional to the power,
[0128] η is the noise intensity coefficient,
[0129] f noise (t) is an additional time-varying noise source.
[0130] In step S411, it is also necessary to establish a multivariate autoregressive integrated moving average model to predict the influence of environmental conditions on the experiment, so as to predict the values at future moments and adjust the laser parameters accordingly. Specifically:
[0131] Φ(B)(T t , Ht) = Θ(B) ∈ t
[0132] Among them, B is the lag operator,
[0133] Φ(B) is the autoregressive average polynomial,
[0134] T t is the time series data of temperature,
[0135] H t is the time series data of humidity,
[0136] Θ(B) is the moving average polynomial,
[0137] ∈ t is white noise.
[0138] In step S412, the Bayesian framework is used for feature selection, and the importance of features is evaluated by calculating the posterior probability. Given the dataset D, the posterior probability of feature F i is obtained through Bayes' theorem, and its posterior distribution P(F i |D) is:
[0139]
[0140] Among them, P(D|F i ) is the likelihood function,
[0141] P(F i ) is the prior distribution,
[0142] P(D) is the evidence.
[0143] The likelihood function P(D|F i ) represents the probability of observing the data D under the condition of the given feature F i , which reflects the fitting degree of the model to the data. Specifically:
[0144]
[0145] where N is the number of observations,
[0146] D j is the j-th observation,
[0147] f(F i ) is the laser output power predicted by the model,
[0148] σ 2 is the variance of the observation noise.
[0149] The prior distribution P(F i ) is specifically:
[0150]
[0151] where μ 0 is the mean vector,
[0152] ∑ 0 is the covariance matrix.
[0153] In step S412, a sparsity constraint also needs to be introduced to achieve automatic feature extraction. The goal is to minimize the reconstruction error while maintaining the sparsity degree, specifically:
[0154]
[0155] where X is the original data matrix,
[0156] W is the dictionary matrix,
[0157] H is the sparse coding matrix,
[0158] ||·|| F is the Frobenius norm,
[0159] ||·|| 1 is the L1 norm,
[0160] λ is the regularization parameter.
[0161] In step S413, the use of data augmentation techniques to generate more training samples is specifically to simulate different experimental conditions by means including adding noise and transforming the time scale. The adding of noise includes: for the original time-domain waveform x(t), generating new samples by adding random noise to it. The newly generated sample x noise (t) is:
[0162] x noise (t) = x(t) + ηn(t),
[0163] Among them, η is the noise intensity coefficient, which is used to control the amplitude of the added noise.
[0164] σ 2 is the variance of the observation noise.
[0165] The transformed time scale includes: stretching or compressing the pulse width. For the original time-domain waveform x(t), by applying a time scaling factor γ, the new sample after transformation is obtained:
[0166] x scale (t) = x(γt),
[0167] where γ > 1 represents time compression, making the pulse narrower;
[0168] 0 < γ < 1 represents time stretching, making the pulse wider.
[0169] To maintain energy conservation, it is usually necessary to normalize the transformed signal. The normalized sample x norm (t) is:
[0170]
[0171] When adding noise and transforming the time scale simultaneously, the new sample obtained is:
[0172]
[0173] In step S420, a layer based on physical equations is added at the front end, and an interpretation layer is added at the end of the grid to combine the advantages of traditional physical models and deep learning. Accordingly, the neural network architecture in the NNs model sequentially includes: an input layer, a physics-aware layer, a deep learning layer, an interpretation layer, and an output layer.
[0174] The features input by the input layer include the output power of the laser, the settings of the nonlinear shaping module, and the environmental conditions.
[0175] The physics-aware layer directly receives the input feature F i , and calculates the intermediate variable Y phy , and the intermediate variable Y phy includes the spectral width and peak power, specifically:
[0176] Y phy = f phy (F i ; θ phy ),
[0177] where θ phy is the relevant physical parameter,
[0178] fphy (·) is the NLS equation;
[0179] The NLS equation f phy is specifically:
[0180]
[0181] where A(z, t) is the normalized electric field amplitude,
[0182] z is the propagation distance,
[0183] t is the time.
[0184] The deep learning layer includes a hidden layer and a residual connection layer. The hidden layer includes multiple fully connected layers or convolutional layers for capturing complex non-linear mapping relationships. Each layer is expressed as:
[0185] Z (k) = σ(W (k) Z (k-1) + b (k) ),
[0186] where k represents the k-th layer,
[0187] W (k) is the weight matrix,
[0188] b (k) is the bias vector,
[0189] σ(·) is the activation function,
[0190] Z (0) = [Y phy , F], and F is the original input feature;
[0191] The residual connection layer helps with gradient transmission and prevents the vanishing gradient problem in deep networks. Specifically:
[0192] Z (k) = Z (k-1) + σ(W (k) Z (k-1) + b (k) ).
[0193] The interpretation layer quantifies the importance score S of each input feature through SHAP values j , specifically:
[0194] S j = φ j (Z (K) ),
[0195] where S j is the importance score of the j-th feature,
[0196] φ j (·) is an explanatory function,
[0197] K is the index of the last layer.
[0198] The output layer outputs the optimized pulse shaping parameters or the predicted time-domain waveform data.
[0199] As Figure 6 shown, in step S430, it specifically further includes the following steps:
[0200] S431. Divide the dataset into a training set (80%) and a validation set (20%), ensuring that each part is independent and has a consistent distribution;
[0201] S432. Use the data in the training set to implement batch gradient descent for model training, use the data in the validation set for cross-validation to further stabilize the performance of the NNs model, introduce an early stopping mechanism, and terminate the training in advance when the performance on the validation set no longer improves to prevent overfitting;
[0202] S433. Use Bayesian optimization to systematically adjust the optimal hyperparameter configuration, and use automated machine learning tools to simplify the process of adjusting the optimal hyperparameter configuration;
[0203] S434. Apply regularization techniques to ensure the generalization ability of the model;
[0204] S435. Use transfer learning to preload some weights to accelerate the convergence speed of the model in the new environment.
[0205] In step S431, the training set is D train , and the validation set is D val .
[0206] In step S432, the batch gradient descent is specifically:
[0207]
[0208] where W is the model parameter,
[0209] q is the learning rate,
[0210] L(·) is the loss function,
[0211] X batch , Y batch are the batch data extracted from the training set D train .
[0212] The early stopping mechanism is specifically:
[0213]
[0214] Among them, L val (·) is the loss function on the validation set D val and
[0215] t stop is the time step to stop training.
[0216] In step S433, the Bayesian optimization is specifically as follows:
[0217]
[0218] Among them, is the expectation.
[0219] In step S434, the L2 regularization is specifically as follows:
[0220]
[0221] Among them, λ reg is the regularization strength coefficient,
[0222] is the L2 norm.
[0223] As Figure 7 shown, in step S440, it specifically further includes the following steps:
[0224] S441. Randomly select unused partial data in the validation set as the test set, use the test set to evaluate the true performance of the NNs model, and pay attention to the evaluation criteria specific to the optical pulse shaping task;
[0225] S442. Apply XAI technology to analyze the decision path of the NNs model, provide an intuitive understanding, and improve the experimental data according to the obtained information.
[0226] Embodiment 2
[0227] As Figure 8 shown, the embodiment of the present invention provides a device for optimizing non - linear optical pulse shaping parameters based on a neural network, including: a laser 1, a non - linear shaping module 2 arranged at the rear end of the laser 1, a photodetector 3 arranged at the rear end of the non - linear shaping module 2, an oscilloscope 4 arranged at the rear end of the photodetector 3, and a computer 5.
[0228] The laser 1 has a tail fiber and outputs pulsed laser light.
[0229] The non - linear shaping module 2 is used for shaping the time - domain waveform of the pulses output by the laser 1, and includes a Gaussian filter 21, a pre - chirp generator 22, an adjustable attenuator 23, a zero - dispersion fiber 24, and a large - dispersion fiber 25. The Gaussian filter 21 makes the pulse waveform output by the laser 1 be Gaussian - shaped. The pre - chirp generator 22 is used to generate Gaussian - shaped chirped pulses. The zero - dispersion fiber 24 is a zero - dispersion highly non - linear fiber, which provides strong non - linearity, stretches and shapes the spectrum by using the SPM effect. The large - dispersion fiber 25 is a large - dispersion fiber, the dispersion length LD is much larger than the non - linear length LNL, and its length is sufficient to generate the DFT effect, providing a large amount of dispersion. Through the dispersion Fourier transform (DFT) effect, a pulse waveform more consistent with the spectral shape is obtained.
[0230] The photodetector 3 and the oscilloscope 4 are used to measure the shaped time - domain waveform.
[0231] The computer 5 is connected to the signal output end of the oscilloscope 4, and is also signal - connected between the pre - chirp generator 22, the adjustable attenuator 23, the zero - dispersion fiber 24, and the large - dispersion fiber 25.
[0232] Preferably, the laser 1 outputs a center wavelength of 1030 nm, a repetition frequency of 5 MHz, a 3 - dB spectral width of about 12 nm, and a 3 - dB pulse width of about 20 ps. The bandwidth of the Gaussian filter 21 is 10 nm. The pre - chirp generator 22 uses PM980 fiber, with a mode field diameter of 6.5 um, a second - order dispersion coefficient of 25 ps² / km, and a non - linear coefficient of 3.9 W⁻¹ / km. The zero - dispersion fiber 24 uses a non - linear photonic crystal fiber SC - 5.0 - 1040 - PM, with a mode field diameter of 4 um, a second - order dispersion coefficient of 0 ps² / km, and a non - linear coefficient of 11 W⁻¹ / km. The large - dispersion fiber 25 uses a hollow photonic crystal fiber, HC - 1060 - 02, with a mode field diameter of 7.5 um, a second - order dispersion coefficient of 67.5 ps² / km, and a non - linear coefficient of 3 W⁻¹ / km. The time - domain waveform is detected by the photodetector 3 (Newports, 1544 - B, wavelength 500 - 1630 nm, bandwidth 12 GHz) and the oscilloscope 4 (Zhongke Siyi, MSO64B, bandwidth 8 GHz, sampling rate 50 GSa / s).
[0233] It is easy for those skilled in the art to understand that the above - mentioned is only a preferred embodiment of the present invention, and is not used to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for optimizing nonlinear optical pulse shaping parameters based on a neural network, characterized in that: The following steps are involved: S100, performing time domain waveform shaping on the pulsed laser output by the laser (1); S200, measuring the shaped time domain waveform; S300, collecting the shaped time domain waveform data; S400, a pulse shaping parameter inversion model based on a neural network, and training the model using sample data; S500, taking the target pulse waveform as input and bringing it into the pulse shaping parameter inversion model to invert the optimal value of the pulse shaping parameter; Step S400 specifically includes the following steps: S410, collecting variables in the input parameters during shaping in step S100 as feature data, collecting the time domain distribution after shaping under different feature data as label data, and arranging and generating a data set; S420, building a pulse shaping parameter inversion model based on a neural network, denoted as a NNs model; S430, dividing the data in the data set into two parts, large and small, and using most of the data as sample data to train the NNs model; S440, select several sets of feature data from the remaining small part of the data and bring them into the NNs model for verification.
2. The method for optimizing nonlinear optical pulse shaping parameters based on a neural network according to claim 1, characterized in that: In step S410, the following steps are specifically included: S411, systematically collecting input parameters including the output power of the laser (1) under different conditions, the settings of the nonlinear shaping module (2), and the corresponding time domain waveform data after shaping, and recording influencing factors including experimental environment conditions; S412, identify and quantify the key factors that affect the pulse shaping effect, use physical prior knowledge to construct features, apply feature selection algorithms to select the most influential features, reduce redundancy, and improve model efficiency; S413. Use data augmentation techniques to generate more training samples. For scarce but important cases, use oversampling or synthetic minority class oversampling techniques to increase their representativeness.
3. The method for optimizing nonlinear optical pulse shaping parameters based on a neural network according to claim 2, characterized in that: In step S411, the laser output power and environmental conditions are regarded as random processes. The laser output power P(t) is described by a geometric Brownian motion with drift and diffusion terms, specifically: dP(t)=μ(P,t)dt+σ(P,t)dW t , Among them, μ(P, t) is the drift term, σ(P, t) is the diffusion coefficient, W t It is the Wiener process.
4. The method for optimizing nonlinear optical pulse shaping parameters based on a neural network according to claim 2, characterized in that: In step S412, feature selection is performed using the Bayesian framework, and the importance of features is evaluated by calculating the posterior probability. Given a data set D, feature F is obtained by using the Bayesian theorem. i The posterior probability of i |D) is: Among them, P(D|F i ) is the likelihood function, P(F i ) is the prior distribution, P(D) is the evidence.
5. A method for optimizing nonlinear optical pulse shaping parameters based on a neural network according to any one of claims 1 to 4, characterized in that: In step S420, a layer based on physical equations is added at the front end, and an interpretation layer is added at the end of the grid to combine the advantages of traditional physical models and deep learning; Accordingly, the neural network architecture in the NNs model includes: input layer, physical perception layer, deep learning layer, interpretation layer and output layer in sequence.
6. The method for optimizing nonlinear optical pulse shaping parameters based on a neural network according to claim 5, characterized in that: The physical perception layer directly receives the input feature F i , and calculate the intermediate variable Y phy , the intermediate variable Y phy Including spectrum width and peak power, specifically: Y phy =f phy (F i ;θ phy ), Among them, θ phy are the relevant physical parameters, f phy (·) is the NLS equation; The NLS equation f phy (·) Specifically: Where A(z, t) is the normalized electric field amplitude, z is the propagation distance, t is time; The deep learning layer includes a hidden layer and a residual connection layer. The hidden layer includes multiple fully connected layers or convolutional layers to capture complex nonlinear mapping relationships. Each layer is represented as: WITH (k) =σ(W (k) WITH (k-1) +b (k) ), Among them, k represents the kth layer, W (k) is the weight matrix, b (k) is the bias vector, σ(·) is the activation function, Z (0) =[Y phy , F], and F is the original input feature; The residual connection layer helps the gradient transfer and prevents the gradient vanishing problem in the deep network. Specifically: WITH (k) =Z (k-1) +σ(W (k) WITH (k-1) +b (k) ); The explanation layer quantifies the importance score S of each input feature through the SHAP value j , specifically: S j =φ j (Z (K) ), Among them, S j is the importance score of the jth feature, φ j (·) is the explanation function, K is the index of the last layer.
7. A method for optimizing nonlinear optical pulse shaping parameters based on a neural network according to any one of claims 1 to 4, characterized in that: In step S430, the following steps are specifically included: S431. Divide the data set into a training set and a validation set to ensure that each part is independent and has a consistent distribution; S432. Use the data in the training set to implement the batch gradient descent method for model training, use the data in the validation set for cross-validation to further stabilize the performance of the NNs model, introduce an early stopping mechanism, and terminate the training early when the performance on the validation set no longer improves to prevent overfitting; S433. Use Bayesian optimization to systematically adjust the optimal hyperparameter configuration and use automated machine learning tools to simplify the process of adjusting the optimal hyperparameter configuration. S434, Apply regularization techniques to ensure model generalization ability; S435. Use transfer learning to preload some weights to speed up the convergence of the model in the new environment.
8. A method for optimizing nonlinear optical pulse shaping parameters based on a neural network according to any one of claims 1 to 4, characterized in that: In step S440, the following steps are specifically included: S441, randomly selecting unused data from the validation set as a test set, using the test set to evaluate the actual performance of the NNs model, focusing on evaluation criteria specific to the optical pulse shaping task; S442. Apply XAI technology to analyze the decision path of NNs models, provide intuitive understanding, and improve experimental data based on the information obtained.
9. A method for optimizing nonlinear optical pulse shaping parameters based on a neural network according to any one of claims 1 to 4, characterized in that: In step S100, the following steps are specifically included: S110, using a Gaussian filter (21) to make the laser pulse waveform emitted by the laser (1) Gaussian; S120, using the pre-chirp generator (22) to generate a Gaussian chirped pulse; S130, adjusting the output average power of the pre-chirp generator (22) through an adjustable attenuator (23), thereby adjusting the peak power of the pulsed laser; S140, providing strong nonlinearity through zero-dispersion optical fiber (24), and utilizing the SPM effect to stretch and shape the spectrum; S150, providing a large dispersion amount through a large dispersion optical fiber (25), and obtaining a pulse waveform consistent with the spectrum shape through the DFT effect.
10. A nonlinear optical pulse shaping parameter optimization device based on a neural network, used to implement a nonlinear optical pulse shaping parameter optimization method based on a neural network as described in any one of claims 1 to 9, characterized in that: include: A laser (1), a nonlinear shaping module (2) arranged at the rear end of the laser (1), a photodetector (3) arranged at the rear end of the nonlinear shaping module (2), an oscilloscope (4) arranged at the rear end of the photodetector (3), and a computer (5); The laser (1) has a pigtail and outputs pulsed laser; The nonlinear shaping module (2) is used for time domain waveform shaping of the output pulse of the laser (1), and comprises a Gaussian filter (21), a pre-chirp generator (22), an adjustable attenuator (23), a zero-dispersion optical fiber (24) and a large-dispersion optical fiber (25); The photodetector (3) and the oscilloscope (4) are used to measure the shaped time domain waveform; The computer (5) is connected to the signal output end of the oscilloscope (4), and is signal-connected between the pre-chirp generator (22), the adjustable attenuator (23), the zero-dispersion optical fiber (24) and the large-dispersion optical fiber (25).
Citation Information
Cited By
Optical probability diffusion type image generation method and system
CN121386188A