An Optimization Method and Device for Multi-Dither Phase-Locked Control Based on Integral Points

By introducing the reinforcement learning model of the Q-learning algorithm, the integral points are automatically optimized, and the complexity and error problems of manual adjustment of integral points in multi-jitter phase lock control are solved, thereby achieving more efficient phase lock control.

CN119960313BActive Publication Date: 2025-07-11LASER FUSION RES CENT CHINA ACAD OF ENG PHYSICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510429285.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The selection of integral points in existing multi-jitter phase lock control depends on manual adjustment, the process is complex and has great subjectivity and error, which affects the automation and accuracy of system optimization.

Method used

The reinforcement learning strategy optimization model based on the Q-learning algorithm is adopted to automatically adjust the integral point number, and the integral point number setting is optimized through the reinforcement learning controller to achieve dynamic adjustment.

Benefits of technology

It improves the automation degree and response speed of phase lock control, reduces errors caused by human factors, and improves the stability and efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960313B_ABST
    Figure CN119960313B_ABST
Patent Text Reader

Abstract

The present application discloses an optimization method and device for multi-dither phase-locked control based on integral points, relating to the field of phase-locked control. The method includes: performing phase-locked control on a fiber laser based on the multi-dither method and obtaining the current output voltage signal; inputting the current output voltage signal into a preset optimization model to obtain the corresponding optimized integral points; the preset optimization model is a policy optimization model based on the Q-learning algorithm; multiplying the current output voltage signal by the demodulation signal of the corresponding frequency, integrating according to the optimized integral points to obtain an error signal, and then performing phase correction on the fiber laser according to the error signal. The present application can dynamically adjust the integral points to achieve efficient and effective correction of the phase error.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of phase-locked control, and particularly to an optimization method and device for multi-dither phase-locked control based on the number of integration points. Background Art

[0002] The multi-dither method is a phase-locked technique for coherent synthesis. By slightly dithering the phases of the outputs of multiple lasers and monitoring the changes in the output interference signals, real-time correction of the phase error is achieved. The implementation steps of the multi-dither method on the phase-locked control board are as follows: applying sine modulation signals with different frequencies to each phase modulator; sampling and filtering the signals of the photodetector; multiplying the filtered signals of the photodetector by the demodulation signals corresponding to the frequencies, and integrating to obtain the error signal; multiplying the error signal by a proportionality coefficient and outputting it to the phase modulator for correction. Through this closed-loop control process, real-time correction of the phase error is achieved, thereby maintaining the stability of the coherent synthesis system.

[0003] In the current implementation steps of multi-dither phase-locked control, when demodulating the error signal by integration, a shorter integration time may result in a faster control response, but it may also make the system more sensitive to noise; a longer integration time can smooth short-term fluctuations and improve the stability of the system, but it may also cause the system response to slow down, which may not be desirable in applications that require a fast response. Therefore, selecting an appropriate integration time is a key system design decision. The current method for selecting an appropriate number of integration points (the number of integration points has the same meaning as the integration time) mainly relies on manual adjustment. Generally, 5 times or more of the modulation period is taken, and then the peak-to-peak change of the output voltage is observed with an oscilloscope to judge the rationality of the number of integration points. However, this method has a complex process and relies on manual experience. It is not only inefficient but also has large subjectivity and errors, seriously restricting the automation and accuracy of system optimization. Summary of the Invention

[0004] The purpose of the present application is to provide an optimization method and device for multi-dither phase-locked control based on the number of integration points, which can dynamically adjust the number of integration points and achieve efficient and effective correction of the phase error.

[0005] To achieve the above purpose, the present application provides the following solutions:

[0006] In the first aspect, the present application provides an optimization method for multi-dither phase-locked control based on the number of integration points, including:

[0007] Performing phase-locked control on a fiber laser based on the multi-dither method and obtaining the current output voltage signal;

[0008] Input the current output voltage signal into a preset optimization model to obtain corresponding optimized integration points; the preset optimization model is a policy optimization model based on the Q-learning algorithm;

[0009] Multiply the current output voltage signal by a demodulation signal of a corresponding frequency, integrate according to the optimized integration points to obtain an error signal, and then perform phase correction on the fiber laser according to the error signal.

[0010] In a second aspect, the present application provides a multi-dither phase-locked control optimization device based on integration points, including:

[0011] A voltage signal output module for: performing phase-locked control on a fiber laser based on the multi-dither method and obtaining a current output voltage signal;

[0012] A reinforcement learning controller for: inputting the current output voltage signal into a preset optimization model to obtain corresponding optimized integration points; the preset optimization model is a policy optimization model based on the Q-learning algorithm;

[0013] A phase-locked control board for: multiplying the current output voltage signal by a demodulation signal of a corresponding frequency, integrating according to the optimized integration points to obtain an error signal, and then performing phase correction on the fiber laser according to the error signal.

[0014] According to the specific embodiments provided by the present application, the present application has the following technical effects: The present application provides a multi-dither phase-locked control optimization method and device based on integration points, which adopts a preset optimization model, which is a policy optimization model based on the Q-learning algorithm. By introducing the Q-learning algorithm of reinforcement learning, the integration point setting is automatically optimized, avoiding the cumbersome process of traditional manual adjustment of integration points. Compared with the manual adjustment method in the prior art, the present application realizes more efficient and automated integration point optimization, greatly improving the automation degree and response speed of phase-locked control, and reducing the error caused by human factors. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0016] Figure 1 It is a schematic flowchart of a multi-dither phase-locked control optimization method based on integration points in an embodiment of the present application.

[0017] Figure 2 It is a principle block diagram of multi-dither phase-locking.

[0018] Figure 3 It is a schematic diagram of the principle of reinforcement learning.

[0019] Figure 4 It is a schematic diagram of the simulation result provided in an embodiment of the present application.

[0020] Figure 5 It is a schematic diagram of an optimization device for multi-dither phase-locking control based on integral points in an embodiment of the present application. Specific implementation manners

[0021] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0022] To make the objectives, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0023] In an exemplary embodiment, as Figure 1 shown, an optimization method for multi-dither phase-locking control based on integral points is provided, including the following steps 101 to 103.

[0024] Step 101, perform phase-locking control on the fiber laser based on the multi-dither method, and obtain the current output voltage signal. As Figure 2 shown, the process of performing phase-locking control on the fiber laser based on the multi-dither method includes: applying sine modulation signals with different frequencies to each phase modulator; sampling and filtering the signals of the photodetector; multiplying the filtered signals of the photodetector with the demodulation signals of the corresponding frequencies in the PD feedback controller, integrating to obtain the error signal, and then multiplying the error signal with the proportional coefficient and outputting it to the phase modulator for correction. Through this closed-loop control process, real-time correction of the phase error is achieved, thereby maintaining the stability of the coherent synthesis system. Corresponding to step 101 of the present application, the obtained output voltage signal can be the signal directly sampled by the photodetector or the filtered signal of the photodetector. Generally, in order to ensure the accuracy of subsequent data, the filtered photodetector signal is selected.

[0025] Step 102: Input the current output voltage signal into a preset optimization model to obtain the corresponding optimized integral points; the preset optimization model is a policy optimization model based on the Q-learning algorithm. Combined with Figure 3 , in an application example, the training iteration process of the preset optimization model in Step 102 includes the following steps (21)-(25).

[0026] (21) In one iteration, use the multi-dithering method to perform phase-locked control on the fiber laser to obtain the state at time t; the state at time t represents the peak-to-peak value of the output voltage. Specifically, let represent the output voltage value, and the state at time t The calculation formula is . min() is the minimum value function, is the maximum value function.

[0027] In order to make the Q-learning algorithm in reinforcement learning easier to process continuous peak-to-peak data, the peak-to-peak value is discretized into a finite state set, and then the reinforcement learning algorithm can make decisions for the discrete states. The calculation formula for the state at time t is:

[0028] .

[0029] Among them, is the state at time t, is the output voltage value at time t, is the total number of peak-to-peak discrete states (such as 100 peak-to-peak discrete states). is the minimum output voltage, is the maximum output voltage, which can usually be set as an empirical value or estimated by observing data. Through the minimum value function and the maximum value function, ensure that the state is between 1 and . The peak-to-peak discrete state is obtained by discretizing the peak-to-peak value of the output voltage.

[0030] (22) Based on the -greedy policy, determine the action at time t according to the state at time t and the Q-table; the Q-table is used to store the value estimates of the state at time t and the action at time t; the action at time t represents the optimized integral points.

[0031] Action is the action taken by the agent intelligent body in the state In this application, the action is to select an integral point n value, and the action space can be set as a finite action set: . Among them, is the first integration point number, is the second integration point number, is the k-th integration point number. At the initial stage of iteration, an action is randomly selected from the set of optional actions ; while in the subsequent iteration stages, in order to prevent getting stuck in local optimal solutions, a commonly used strategy is -greedy strategy. This strategy selects the action with the current largest Q value in the Q-table most of the time, but will with a certain probability randomly explore other actions.

[0032] The formula for determining the action at time t is as follows:

[0033] .

[0034] Among them, is the action at time t, is the Q value when selecting action in state , is the probability determined based on the -greedy strategy, () is the maximum value function.

[0035] As the training progresses, gradually decrease , so that the agent gradually selects the optimal action more and more, rather than exploring random actions.

[0036] (23) Based on the action at time t, perform phase correction on the fiber laser to obtain the state at time t + 1 and the corresponding reward at time t; specifically, multiply the state at time t by the demodulation signal of the corresponding frequency, and perform integration according to the action at time t to obtain an error signal, and then perform phase correction on the fiber laser according to the error signal. That is, after the action acts on the environment, the environment changes accordingly and feeds back a new state and the corresponding reward . The reward is the feedback obtained by the agent after executing the action at each time step, and is used to guide the agent to learn the optimal strategy.

[0037] The reward target of the reward at time t is: to minimize the peak-to-peak output voltage, so the reward is designed as: , and, , the above two formulas indicate that the larger the peak-to-peak value, the smaller the reward.

[0038] (24) Based on the state at time t, the action at time t, the state at time t + 1, and the reward at time t, use the Bellman equation to update the Q-table.

[0039] The core objective of reinforcement learning is to let the agent maximize the cumulative reward and tend to select the n value that can minimize the peak-to-peak value. In Q-learning, the Q-table is a table that stores the value estimates of each state-action pair and is used to help the agent select the optimal action in each state. That is, the Q-table is used to store each state and the value estimate of each integral point number n (i.e., the action ), so as to guide the agent to select the integral point number that minimizes the voltage peak-to-peak value. The Q-table is a two-dimensional matrix, with each row corresponding to a state , and each column corresponding to an action . Each element in the Q-table represents the cumulative reward (future return) expected to be obtained when performing the action in the state .

[0040] .

[0041] At the initial stage of Q-learning training, the values in the Q-table are usually initialized to 0 or random values. After each time step t, the Q-table is updated according to the current state , the current action , the reward and the next state . Specifically, when using the Bellman equation to update the Q-table, the following update formula is used:

[0042] .

[0043] Among them, is the state at time t, is the action at time t, is the Q value when selecting the action in the state ; is the learning rate, which controls the step size of each update; is the reward at time t, is the discount factor, is the maximum Q value among all possible actions in the state at time t + 1 . The meaning of the above update formula is: at each time step t, Q-learning updates the Q value of the current state-action pair , considering the immediate reward obtained currently and the future in the state The maximum return that can be obtained. Through this loop, the agent continuously learns and optimizes its policy to select the optimal action that can obtain the maximum reward in different states.

[0044] (25) If the preset iteration stop condition is not met, based on the updated Q-table, the next iteration is performed; if the preset iteration stop condition is met, the iteration is stopped, and the updated Q-table is marked as the preset optimization model. After multiple iterations, the Q-values in the Q-table will gradually approach the actual state-action value function. Finally, the agent will select the action that maximizes the long-term return . This means that the agent has learned to select the optimal n value in each state, thereby minimizing the peak-to-peak voltage.

[0045] Step 103: Multiply the current output voltage signal by the demodulation signal corresponding to the frequency, integrate according to the optimized integration points to obtain an error signal, and then perform phase correction on the fiber laser according to the error signal.

[0046] In a specific application example, the present application also performed simulations. The simulation conditions were as follows: the sampling frequency was 250 MHz, the integration points were initially set to one sampling period, a sinusoidal wave phase noise with a frequency of 20 kHz was added at the 20th microsecond, and a reinforcement learning controller was added at the 40th microsecond to optimize the integration points. The obtained simulation results are as Figure 4 shown. It can be seen from this figure that after using the method of the present application to optimize the integration points, the peak-to-peak value of the output voltage is significantly reduced.

[0047] In summary, by introducing reinforcement learning, the present application automatically optimizes the integration point setting, avoiding the cumbersome process of traditional manual adjustment of integration points. Compared with the manual adjustment method in the prior art, the present application greatly improves the automation degree and response speed of the phase-locked control optimization, and reduces the error caused by human factors.

[0048] Based on the same inventive concept, the embodiment of the present application also provides a multi-dither phase-locked control optimization device based on integration points. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more device embodiments provided below can refer to the limitations on the method in the above text, and will not be repeated here. As Figure 5 shown, the multi-dither phase-locked control optimization device based on integration points of the present application includes a voltage signal output module, a reinforcement learning controller, and a phase-locked control board, and forms a control closed-loop. This control closed-loop can dynamically adjust the integration points to effectively correct the phase error.

[0049] The voltage signal output module is used to: perform phase-locked control on the fiber laser based on the multi-dither method, obtain the current output voltage signal, and then input it to the reinforcement learning controller. Wherein, the current output voltage signal is the peak-to-peak value of the output voltage.

[0050] The reinforcement learning controller is used to: input the current output voltage signal into a preset optimization model to obtain the corresponding optimized integral points, and then input the optimized integral points to the phase-locked control board; the preset optimization model is a policy optimization model based on the Q-learning algorithm.

[0051] The phase-locked control board is used to: multiply the current output voltage signal by the demodulation signal of the corresponding frequency, perform integration according to the optimized integral points to obtain an error signal, and then perform phase correction on the fiber laser according to the error signal, and further generate the next output voltage signal and transmit it to the voltage signal output module.

[0052] In a specific practical application, the reinforcement learning controller includes a reinforcement learning training module and an optimization model application module; wherein, the reinforcement learning training module is used to execute the following training iteration process to realize multiple iterative calculations based on the input value:

[0053] In one iteration, perform phase-locked control on the fiber laser using the multi-dither method to obtain the state at time t; the state at time t represents the peak-to-peak value of the output voltage; based on - The greedy policy, determine the action at time t according to the state at time t and the Q-table; the Q-table is used to store the value estimates of the state at time t and the action at time t; the action at time t represents the optimized integral points; based on the action at time t, perform phase correction on the fiber laser to obtain the state at time t + 1 and the corresponding reward at time t; based on the state at time t, the action at time t, the state at time t + 1 and the reward at time t, use the Bellman equation to update the Q-table; if the preset iteration stop condition is not satisfied, then based on the updated Q-table, perform the next iteration; if the preset iteration stop condition is satisfied, then stop the iteration and mark the updated Q-table as the preset optimization model.

[0054] In summary, under the monitoring of the peak-to-peak value of the output voltage, the reinforcement learning controller can effectively ensure the system stability and the optimal phase-locked effect by adaptively adjusting the integral points. In addition, the method of automatically adjusting the integral points greatly reduces manual intervention and improves the overall reliability and efficiency of the system. Obviously, the present application proposes to use the reinforcement learning algorithm to automatically optimize the integral points, which can avoid the influence of inaccurate integral points on the phase-locked effect.

[0055] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0056] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. An optimization method for multi-dither phase-locked control based on integral points, characterized in that The multi-dither phase-locked control optimization method based on integral points includes: Performing phase-locked control on the fiber laser based on the multi-dither method and obtaining the current output voltage signal; Inputting the current output voltage signal into a preset optimization model to obtain the corresponding optimized integral points; the preset optimization model is a policy optimization model based on the Q-learning algorithm; Multiplying the current output voltage signal by the demodulation signal of the corresponding frequency, integrating according to the optimized integral points to obtain an error signal, and then performing phase correction on the fiber laser according to the error signal; The training iteration process of the preset optimization model includes: In one iteration, performing phase-locked control on the fiber laser using the multi-dither method to obtain the state at time t; the state at time t represents the peak-to-peak value of the output voltage; Based on - a greedy strategy to determine the action at time t according to the state at time t and the Q-table; the Q-table is used to store the value estimation of the state at time t and the action at time t; the action at time t represents the optimized integral points. Based on the action at time t, performing phase correction on the fiber laser to obtain the state at time t + 1 and the corresponding reward at time t; the reward target of the reward at time t is to minimize the peak-to-peak value of the output voltage; Updating the Q-table based on the state at time t, the action at time t, the state at time t + 1, and the reward at time t using the Bellman equation; If the preset iteration stop condition is not satisfied, perform the next iteration based on the updated Q-table; if the preset iteration stop condition is satisfied, stop the iteration and mark the updated Q-table as the preset optimization model.

2. The multi-dither phase-locked control optimization method based on integral points according to claim 1, wherein The calculation formula for the state at time t is: ; Among them, is the state at time t, is the output voltage value at time t, is the total number of peak-to-peak discrete states, is the minimum value of the output voltage, is the maximum value of the output voltage, min () is the minimum value function, is the maximum value function; the peak-to-peak discrete state is obtained by discretizing the output voltage peak-to-peak value.

3. The multi-dither phase-locked control optimization method based on integral points according to claim 1, characterized in that When updating the Q-table using the Bellman equation, the following update formula is used: ; Among them, is the state at time t, is the action at time t, is in the state select the action when the Q value, is the learning rate, is the reward at time t, is the discount factor, is the state at time t + 1 all possible actions the largest Q value among them, is the maximum value function.

4. The multi-dither phase-locked control optimization method based on integral points according to claim 1, characterized in that The determination formula for the action at time t is as follows: ; Among them, is the action at time t, is the Q-value when selecting the action in state , is the probability determined based on -greedy strategy, () is the maximum value function.

5. A multi-dither phase-locked control optimization device based on integral points, characterized in that, The multi-dither phase-locked control optimization device based on integral points includes: A voltage signal output module for performing phase-locked control on the fiber laser based on the multi-dither method and obtaining the current output voltage signal; A reinforcement learning controller for inputting the current output voltage signal into a preset optimization model to obtain the corresponding optimized integral points; the preset optimization model is a policy optimization model based on the Q-learning algorithm; A phase-locked control board for multiplying the current output voltage signal by the demodulation signal of the corresponding frequency, integrating according to the optimized integral points to obtain an error signal, and then performing phase correction on the fiber laser according to the error signal; The reinforcement learning controller includes a reinforcement learning training module and an optimization model application module; Among them, the reinforcement learning training module is used to execute the following training iteration process: In one iteration, performing phase-locked control on the fiber laser using the multi-dither method to obtain the state at time t; the state at time t represents the peak-to-peak value of the output voltage; Based on - a greedy strategy to determine the action at time t according to the state at time t and the Q-table; the Q-table is used to store the value estimate of the state at time t and the action at time t; the action at time t represents the optimized integral points; Based on the action at time t, performing phase correction on the fiber laser to obtain the state at time t + 1 and the corresponding reward at time t; the reward target of the reward at time t is to minimize the peak-to-peak value of the output voltage; Updating the Q-table based on the state at time t, the action at time t, the state at time t + 1, and the reward at time t using the Bellman equation; If the preset iteration stop condition is not met, then based on the updated Q-table, the next iteration is performed; if the preset iteration stop condition is met, then the iteration is stopped, and the updated Q-table is marked as the preset optimization model.

6. The multi-dither phase-locked control optimization device based on integral points according to claim 5, characterized in that The current output voltage signal is the peak-to-peak value of the output voltage.

Citation Information

Patent Citations

  • Reinforcement learning method for coherent combination

    CN112446470A

  • Signal-to-noise ratio assisted adaptive carrier tracking loop system and method

    CN117498891A