Method for Preparing Phonon Fock State Based on Deep Reinforcement Learning
By introducing Hamiltonian and deep reinforcement learning of non-resonant excitation terms, a neural network model is constructed to obtain the optimal evolution path, and the fidelity and noise robustness problems in phonon Fock state preparation are solved, and efficient and robust phonon Fock state preparation is achieved.
Patent Information
- Application Number
- CN202510575083.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The prior art has problems such as poor evolution path, limited fidelity and poor noise-robust performance in the preparation of phonon Fock states.
Hamiltonian with non-resonant excitation terms is introduced as evolutionary dynamics, and a neural network model with deep reinforcement learning is constructed. Through training, the optimal evolution path is obtained and the laser sideband pulse drives single ions to evolve from the phonon Fock ground state to the target state.
The fidelity and efficiency of phonon Fock state preparation are improved, the adaptability of the preparation is enhanced, and the fault tolerance to noise is improved, and the robustness of the model is enhanced.
Smart Images

Figure CN120087489B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of quantum computing, and particularly relates to a method for preparing phonon Fock states based on deep reinforcement learning. Background Art
[0002] In the field of quantum computing technology, phonon Fock states, as an important quantum state, can provide good quantum resources for realizing high-precision quantum operations and quantum logic gates. The manipulation of such quantum states usually relies on sideband operations, and the most typical one is the first-order red sideband pulse, whose frequency has a red detuning of one phonon frequency compared to the energy level resonance frequency, so as to achieve and the transition between. However, simple red sideband pulses face problems of insufficient speed and fidelity.
[0003] In the prior art, methods for improving fidelity include stimulated Raman adiabatic passage, dynamical decoupling, pulse shaping, and adiabatic shortcuts. The resistance of stimulated Raman adiabatic passage and dynamical decoupling to noise is achieved at the cost of time. Pulse shaping has poor resistance to noise. Since the adiabatic shortcut is based on the Hamiltonian of the JC model and does not include non-resonant excitation terms, there will be a large deviation from the experimental results and poor robustness.
[0004] In summary, the prior art has problems such as suboptimal evolution paths, limited fidelity, and poor noise robustness during the preparation of phonon Fock states. Summary of the Invention
[0005] Based on this, the technical solution provided by the present invention aims to construct a neural network model by introducing a Hamiltonian containing non-resonant excitation terms as the evolution dynamics, and obtain an optimal evolution path through the training of deep reinforcement learning to achieve the preparation of phonon Fock states with high fidelity, high speed, and high robustness.
[0006] To achieve the above object, the present invention provides a method for preparing phonon Fock states based on deep reinforcement learning, the method comprising: S1: Cooling a single ion in an ion trap to the phonon Fock ground state, and obtaining an optimal evolution path through a neural network model of deep reinforcement learning. S2: Controlling a laser sideband pulse to evolve according to the optimal evolution path, and driving the single ion with the laser sideband pulse and a carrier wave, so that the single ion evolves from the phonon Fock ground state to a phonon Fock target state.
[0007] Further, the step of obtaining an optimal evolution path through a neural network model of deep reinforcement learning comprises:
[0008] S11: Construct the neural network model, where the neural network model uses the Hamiltonian containing a non-resonant excitation term as the evolution dynamics. Use the linear combination of the action reward function and the final fidelity as the total reward function, and its specific form is:
[0009] ,
[0010] where R is the total reward function, is the final fidelity, is the weight coefficient of the final fidelity is the weight coefficient of the action reward function, k is the weight coefficient of the action reward function, N is the total number of steps for a single model training, is the step identifier, is the action reward function at the i-th step.
[0011] S12: Initialize the network parameters of the neural network model.
[0012] S13: Perform model training in a multi-step progressive manner and output the optimal evolution path.
[0013] Preferably, the specific form of the Hamiltonian is:
[0014] ,
[0015] where, is the annihilation operator, is the creation operator, is the Hamiltonian at the i-th step, is the Rabi frequency at the i-th step, is the detuning at the i-th step, is the reduced Planck constant, is the upward transition operator of the two-level internal state of the single ion, is the vibration frequency of the single ion, is the imaginary unit, t is the time, represents the Hermitian conjugate.
[0016] Furthermore, step S11 further includes normalizing the detuning and the Rabi frequency in the Hamiltonian, and its specific form is:
[0017] ,
[0018] where, is 's normalization parameter, is 's normalization parameter, is the Lamb-Dicke parameter, is the reference Rabi frequency, is the AC Stark frequency, and and both have a value range of .
[0019] Furthermore, step S12 specifically includes: setting the initial phonon state and the final phonon state for model training, with the final phonon state differing from the initial phonon state by one phonon, setting the reference Rabi frequency, calculating the evolution time, and setting the total number of steps for a single model training.
[0020] Furthermore, step S13 specifically includes: setting the weight coefficient k of the action reward function to decrease with the number of training times, and setting the weight coefficient m of the final fidelity to increase with the number of training times. Adjusting the state parameters according to the total reward function R, using the evolution path of the action parameters as the output of a single model training until the model converges, and obtaining the optimal evolution path. Among them, the state parameters include the weight coefficient m of the final fidelity and the weight coefficient k of the action reward function, and the action parameters include the normalized parameter of the detuning and the normalized parameter of the Rabi frequency.
[0021] Even further, the phonon Fock state preparation method further includes: setting the Rabi frequency with noise as , setting the detuning with noise as , using the Rabi frequency with noise to replace the Rabi frequency, using the detuning with noise to replace the detuning, calculating the Hamiltonian, and obtaining the optimal evolution path in a noisy environment.
[0022] Among them, the random frequency noise has a value range of ±5% of the reference Rabi frequency , and the random intensity noise has a value range of ±10% of the reference Rabi frequency .
[0023] The beneficial effects achieved by the present invention through the above technical solutions are:
[0024] (1) By introducing a Hamiltonian containing a non-resonant excitation term as the evolution dynamics to construct a neural network model, and training the model based on deep reinforcement learning to obtain the optimal evolution path of the laser. Compared with the traditional fixed evolution path, such as the evolution path of a pulse, not only the fidelity of the phonon Fock state preparation is improved, but also the preparation efficiency is increased, and the self-adaptability of the phonon Fock state preparation is enhanced.
[0025] (2) By adding random noise to the environment for training, the fault tolerance of the laser sideband pulse control to actual frequency and intensity perturbations is improved, and the robustness of the model is enhanced. Brief Description of the Drawings
[0026] Figure 1 is a flowchart of the method for preparing phonon Fock states based on deep reinforcement learning according to an embodiment of the present invention.
[0027] Figure 2 is the optimal evolution path for different model trainings according to an embodiment of the present invention. Among them, Figure 2 (a) is the optimal evolution path with a total evolution time of 46.15 ; Figure 2 (b) is the optimal evolution path with a total evolution time of 23.08 ; Figure 2 (c) is the optimal evolution path with a total evolution time of 11.54 ; Figure 2 (d) is the optimal evolution path with a total evolution time of 5.77 ;
[0028] Figure 3 is a schematic diagram of the comparison of the fidelity of pulses and deep reinforcement learning in a noise environment according to an embodiment of the present invention. Among them, (a) is the fidelity distribution diagram of the pulse in the noise environment Figure 3 ; (b) is the fidelity distribution diagram of deep reinforcement learning in the noise environment. Figure 3 ;
[0029] Figure 4 is a schematic diagram of the population evolution of phonon Fock states using the optimal evolution path according to an embodiment of the present invention. Detailed Embodiments
[0030] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following further describes the specific embodiments of the present invention in detail in conjunction with the embodiments. It should be understood that the embodiments described herein are only used to explain the present invention, but not to limit the scope of the present invention.
[0031] Please refer to the appended Figure 1 , Figure 1 is a flowchart of the method for preparing phonon Fock states based on deep reinforcement learning according to an embodiment of the present invention. As can be seen from the figure, the method for preparing phonon Fock states includes: S1: Cooling a single ion in an ion trap to the phonon Fock ground state, and obtaining an optimal evolution path through the neural network model of deep reinforcement learning. S2: Controlling the laser sideband pulse to evolve according to the optimal evolution path, and driving the single ion with the laser sideband pulse and the carrier wave, so that the single ion evolves from the phonon Fock ground state to the phonon Fock target state.
[0032] Illustratively, step S1 of the above-described phonon Fock state preparation method specifically includes cooling a single ion in an ion trap to a phonon Fock ground state via Doppler cooling, EIT cooling, and sideband cooling. Ideally, the phonon number of the phonon Fock ground state is zero. In practice, in embodiments of the present invention, the experimentally obtained average phonon number of the phonon Fock ground state is 0.1. In one or other embodiments of the present invention, the average phonon number of the phonon Fock ground state is generally between 0.05 and 0.1.
[0033] Exemplarily, the steps of obtaining the optimal evolution path through a deep reinforcement learning neural network model include:
[0034] (1) Establish the environment, including setting the evolutionary dynamics and total reward function.
[0035] In the embodiment of the present invention, the neural network model uses the Hamiltonian containing non-resonant excitation terms as the evolution dynamics in the environment, which is specifically expressed as follows:
[0036] , (1)
[0037] in, is the annihilation operator, To generate the operator, is the Hamiltonian at step i, is the Rabi frequency of step i, is the detuning amount at step i, is the reduced Planck constant, is the rising transition operator of the second energy level of the single ion internal state, is the vibration frequency of the single ion, is the imaginary unit, t is the time, represents Hermitian conjugation.
[0038] The linear combination of the action reward function and the final fidelity is used as the total reward function, which is in the form of:
[0039] , (2)
[0040] Where R is the total reward function, For the final fidelity, is the weight coefficient of the final fidelity F, k is the weight coefficient of the action reward function, N is the total number of steps of a single model training, is the step number indicator, is the action reward function at step i.
[0041] In the embodiment of the present invention, based on the experience of the adiabatic shortcut, the normalized parameter of the detuning amount is The varying shape is approximately close to a straight line, while the normalized parameter of the Rabi frequency should be close to a sine function. The actual training results may not exactly match the reference adiabatic shortcut curve because, in addition to the stepwise action reward, there is also a reward based on the final fidelity. The ratio between the action reward and the final fidelity reward needs to be continuously tried and adjusted to successfully converge and obtain the best comprehensive effect.
[0042] In expression (1), the laser detuning and the Rabi frequency are the parameters to be controlled, which are determined by the actions output by the model agent. To facilitate the operation of the reinforcement learning algorithm, in the embodiments of the present invention, the two parameters of the detuning and the Rabi frequency are normalized and scaled to the range of (0, 1). The specific form is:
[0043] , (3)
[0044] , (4)
[0045] where is the normalized parameter of , is the normalized parameter of , is the Lamb-Dicke parameter, is the reference Rabi frequency, is the AC Stark frequency, and , both have a value range of .
[0046] In the embodiments of the present invention, the two parameters for parameter substitution in the evolution dynamics are the detuning and the Rabi frequency respectively. The variation range of these two parameters is determined by the magnitude of the reference Rabi frequency of the laser sideband pulse, and the magnitude of the reference Rabi frequency of the laser sideband pulse also determines the total evolution time: .
[0047] In the embodiments of the present invention, an arbitrary waveform generator is used to input an acousto-optic modulator through a radio frequency power amplifier to control the frequency and intensity of light, thereby realizing the control of the laser detuning and the Rabi frequency. In one or other embodiments of the present invention, other methods can also be used to control the two parameters of the laser detuning and the Rabi frequency. In one or other embodiments of the present invention, the evolution path of the laser can also be controlled by controlling other laser or ion parameters to realize the preparation of the phonon Fock state.
[0048] (2) Initialize the network parameters of the neural network model.
[0049] Exemplarily, the initialization step includes setting the initial phonon state and the final phonon state for model training. The final phonon state differs from the initial phonon state by one phonon.
[0050] Exemplarily, the initialization step further includes setting the reference Rabi frequency of the laser sideband pulse and the number of evolution steps for a single model training. The magnitude of the reference Rabi frequency of the laser sideband pulse determines the detuning amount and the change amplitude of the Rabi frequency of the laser sideband pulse, and also determines the total evolution time. In the embodiments of the present invention, the number of evolution steps N for a single model training is set, and the total evolution time is determined according to the reference Rabi frequency. , then the evolution time for each step is . In the embodiments of the present invention, setting the total evolution time equal to the time of the red sideband pulse is only for the convenience of comparing the pulse with the evolution result of the neural network model of the deep reinforcement learning in the embodiments of the present invention. In one or other embodiments of the present invention, there is no limitation on the total evolution time.
[0051] (3) The model training is carried out in a multi-step progressive manner to output the optimal evolution path.
[0052] In the embodiments of the present invention, in the first step of training, a larger weight coefficient of the action reward function is set to obtain a path close to the adiabatic shortcut curve, and the action parameters of this evolution are output. and . In subsequent training, the action parameters output from the previous step of training are loaded, and the weight coefficient k of the action reward function is reduced, while the weight coefficient m of the final fidelity is increased. Until the model converges, the optimal evolution path is obtained.
[0053] Please refer to the appendix Figure 2 , Figure 2 which is the optimal evolution path of different model trainings in the embodiments of the present invention. Figure 2 (a), Figure 2 (b), Figure 2 (c) and Figure 2 (d) respectively show the optimal evolution paths with the total evolution times of 46.15 , 23.08 , 11.54 and 5.77 . The abscissa in the figure represents the evolution time, the solid line represents the evolution path of the population number of the phonon Fock state being , the short dash line represents the optimal evolution path, and the dotted line represents the optimal evolution path. Comparing the results of different model trainings, it can be seen that when the total evolution time is relatively large, the evolution gradually becomes flat, while The evolution range is large. When the total evolution time is small, The maximum value of is reduced and the evolution is gentle. In particular, when the total evolution time is small, The value of the time evolution becomes larger. This is because when the total evolution time is small, the non-resonant excitation (i.e., carrier transition) has a greater impact on the training, and the output of the model training is Smaller, in order to improve the final fidelity, the model training will output a larger evolution .
[0054] In the embodiment of the present invention, the total evolution time is 46.15 hour, The fidelity of the pulse is 0.9945, and the fidelity of deep reinforcement learning is 0.9909. The total evolution time is 23.08 hour, The fidelity of the pulse is 0.9777 and the fidelity of deep reinforcement learning is 0.9990. The total evolution time is 11.54 hour, The fidelity of the pulse is 0.8958, and the fidelity of deep reinforcement learning is 0.9963. The total evolution time is 5.77 hour, The fidelity of the impulse is 0.6585 and the fidelity of deep reinforcement learning is 0.8596. As a basic pulse with constant frequency and intensity, the fidelity of the pulse will continue to decline as the total evolution time decreases. However, the final fidelity obtained by the optimal evolution path given by deep reinforcement learning decreases slowly and is better than The fidelity of the pulse. In addition to the total evolution time of 46.15 Except for the above, the final fidelity of the other three groups of deep reinforcement learning is higher than The fidelity of the pulse proves that deep reinforcement learning has a certain resistance to non-resonant excitation. However, the total evolution time is 46.15 Under the condition, the final fidelity of deep reinforcement learning is slightly lower than The fidelity of the pulse is improved because when the total evolution time is long, the non-resonant excitation has little effect on the preparation of the phonon Fock state. The technical solution provided by the present invention is relatively The advantage of the pulse is not obvious. Moreover, in the embodiment of the present invention, in order to make the action parameters and The curve approaches the adiabatic shortcut, thereby improving robustness, sacrificing a portion of high fidelity. In one or other embodiments of the present invention, for a total evolution time of 46.15 The model sets the weight coefficient k of the action reward function to zero in the total reward function, and the final fidelity obtained by training can exceed pulses, but such a setting makes the model as pulses and lack robustness.
[0055] In the embodiment of the present invention, the total evolution time is selected to be 23.08 model to train the phonon Fock state to for each step between. First, train the to model. The next-step model is trained based on the previous-step model, and the reference Rabi frequency of the laser sideband pulse is changed to , where is the number of phonons in the initial phonon state. According to the optimal evolution path of deep reinforcement learning, manipulate the phonon Fock state to with a fidelity of 0.9992, to with a fidelity of 0.9991, to with a fidelity of 0.9992, to with a fidelity of 0.9998, to with a fidelity of 0.9999. Combining the fidelity of each step, the fidelity from to is 0.9972. In the embodiment of the present invention, the fidelity of each step is continuously improved. This is because each step of the model is trained based on the previous step, indicating that more training leads to further convergence. It is also because the larger the number of phonons, the larger the Rabi frequency of the sideband transition. Under the condition of the same evolution time, the required Rabi frequency for the transition of a larger number of phonons is smaller, so the non-resonant excitation is also smaller.
[0056] In the embodiment of the present invention, in order to train an evolution path that is robust to frequency and intensity noise, an environment with added random noise is also created for further model training. In this noise environment, random frequency noise and random intensity noise are added simultaneously, and their amplitudes are ±5% and ±10% of the reference Rabi frequency respectively. Add the random intensity noise to the Rabi frequency: . Add the random frequency noise to the detuning: .
[0057] Please refer to the appendix Figure 3 , Figure 3 is a schematic diagram of the fidelity comparison between the pulse and deep reinforcement learning in the noise environment of the embodiment of the present invention. Figure 3 (a) is for the noise environment Fidelity distribution diagram of the pulse Figure 3 (b) is the fidelity distribution diagram of deep reinforcement learning in a noisy environment. In the figure, the abscissa represents the random intensity noise , and the ordinate represents the random frequency noise . In order to further compare the fidelity of the pulse and deep reinforcement learning, the embodiment of the present invention fills the position with a fidelity of 0.96 in white. In the embodiment of the present invention, the random ranges of the random frequency noise and the random intensity noise are respectively set to ±5% and ±10% of the reference Rabi frequency. In one or other embodiments of the present invention, other values can also be taken for the random range. Each training will randomly select a value within this range and add it to the evolution dynamics on average. By continuously repeating the training, the model will eventually have a certain resistance to the noise within this range. In the figure, the shade of gray represents the fidelity, and the darker the color, the higher the fidelity. As can be seen from the figure, the highest fidelity of the deep reinforcement learning path is slightly lower than the highest fidelity of the pulse, which indicates that it is difficult for the model of deep reinforcement learning to further converge near the limit value. However, the range where the final fidelity of the deep reinforcement learning path is greater than 0.96 is larger, indicating that the deep reinforcement learning path has better robustness to frequency and intensity noise. The highest fidelity of the pulse is roughly located at the center of the figure, that is, when the values of the random frequency noise and the random intensity noise are equal to the reference Rabi frequency of the laser sideband pulse. The peak of the fidelity of deep reinforcement learning deviates from the center point, indicating that the model has not fully and evenly learned the experience under different magnitudes of noise during training. In one or other embodiments of the present invention, by increasing the number of training times of deep reinforcement learning, the model evenly learns the experience under different magnitudes of noise.
[0058] Please refer to the attached Figure 4 , Figure 4 is a schematic diagram of the evolution of the phonon Fock state population using the optimal evolution path in the embodiment of the present invention. In the figure, the abscissa represents the evolution time, and the ordinate represents the population of the initial evolution phonon state . In the embodiment of the present invention, the total evolution time is set to 46.15 , the initial evolution phonon state is , and the final evolution phonon state is . At this time, the final fidelity is 0.93, which has a certain gap with the theoretical value of 0.99. This indicates that when cooling the single ion in the ion trap to the phonon Fock ground state in step S1, the single ion is not cooled to the state with zero phonon number, and the average phonon number at this time is approximately 0.1. And the average phonon number of the phonon Fock ground state being 0.1 means that it will bring at least a 0.1 loss in the final fidelity, and at the same time will cause the carrier The pulse fidelity is reduced. As shown in the figure, the phonon state population corresponding to the moment t = 0 is not equal to 1, indicating that the initial phonon state used in the experiment is not perfect .
[0059] In one or other embodiments of the present invention, only the phonon final state is used as the judgment basis, the higher the population, the greater the reward. At this time, the reward function is set as: , the significance of using the square in the function is that when the final fidelity approaches 1, the gradient of the reward will continuously increase, which can, to a certain extent, avoid the problem that it is difficult to further converge near the limit value.
[0060] Finally, it should be noted that the above specific embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for preparing phonon Fock states based on deep reinforcement learning, characterized in that, The method includes: S1: Cooling a single ion in an ion trap to the phonon Fock ground state; obtaining an optimal evolution path through a neural network model of deep reinforcement learning; S2: Controlling the laser sideband pulse to evolve according to the optimal evolution path, driving the single ion with the laser sideband pulse and the carrier wave, so that the single ion evolves from the phonon Fock ground state to the phonon Fock target state; Wherein, the Hamiltonian adopted in the evolution dynamics in the neural network model includes a non-resonant excitation term, and the specific form of the Hamiltonian is: , Among them, is an annihilation operator, is the Hamiltonian of the i-th step, is the Rabi frequency of the i-th step, is the detuning of the i-th step, is the reduced Planck constant, is the raising transition operator of the two-level internal state of the single ion, is the vibration frequency of the single ion, is the Lamb-Dicke parameter, is the step number identifier, is the imaginary unit, t is the time, represents the Hermitian conjugate.
2. The method for preparing a phonon Fock state according to claim 1, wherein The step of obtaining the optimal evolution path through the neural network model of deep reinforcement learning includes: S11: Constructing the neural network model, the neural network model takes the linear combination of the action reward function and the final fidelity as the total reward function, and the specific form is: , where R is the total reward function, is the final fidelity, is the weight coefficient of the final fidelity F, k is the weight coefficient of the action reward function, and N is the total number of steps for a single model training, is the action reward function at the i-th step; S12: Initializing the network parameters of the neural network model; S13: Training the model in a multi-step progressive manner and outputting the optimal evolution path.
3. The method for preparing a phonon Fock state according to claim 2, characterized in that, Step S11 further includes: Normalizing the detuning amount of the i-th step and the Rabi frequency of the i-th step in the Hamiltonian, and the specific form is: , Among them, is the normalization parameter of, is the normalization parameter of, is the reference Rabi frequency, is the AC Stark frequency, and , the value ranges of both are .
4. The method for preparing the phonon Fock state according to claim 3, wherein Step S12 specifically includes: Setting the initial phonon state and the final phonon state of model training, and the final phonon state differs from the initial phonon state by one phonon; Setting the reference Rabi frequency, calculating the evolution time, and setting the total number of steps for a single model training.
5. The method for preparing a phonon Fock state according to claim 2, wherein Step S13 specifically includes: Setting the weight coefficient k of the action reward function to decrease with the number of training times, and setting the weight coefficient m of the final fidelity to increase with the number of training times; Adjusting the state parameters according to the total reward function R, taking the evolution path of the action parameters as the output of a single model training until the model converges, and obtaining the optimal evolution path; Wherein, the state parameters include the weight coefficient m of the final fidelity and the weight coefficient k of the action reward function, and the action parameters include the normalized parameter of the detuning amount of the i-th step and the normalized parameter of the Rabi frequency of the i-th step.
6. The method for preparing a phonon Fock state according to claim 3, characterized in that, The method for preparing the phonon Fock state further includes: Set the noisy Rabi frequency to ; Set the detuning with noise to be ; Replacing the Rabi frequency of the i-th step with the noisy Rabi frequency, replacing the detuning amount of the i-th step with the noisy detuning amount, calculating the Hamiltonian, and obtaining the optimal evolution path in a noisy environment; Among them, the random frequency noise has a value range of ±5% of the reference Rabi frequency , and the random intensity noise has a value range of ±10% of the reference Rabi frequency .
Citation Information
Patent Citations
Control method and device for AI intelligent calculation integration, electronic equipment and storage medium
CN119690676A