Phonon Fock state preparation method based on deep reinforcement learning
By introducing Hamiltonian and deep reinforcement learning with non-resonant excitation terms in the preparation of phonon Fock states, the optimal evolution path is obtained, and the problems of poor evolution path, limited fidelity and poor noise robustness in the prior art are solved, and the preparation of phonon Fock states with high fidelity, high speed and high robustness are achieved.
Patent Information
- Application Number
- CN202510575083.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The prior art has problems such as poor evolution path, limited fidelity and poor noise-robust performance in the preparation of phonon Fock states.
By introducing Hamiltonian with non-resonant excitation terms as evolutionary dynamics, a neural network model is constructed, and a deep reinforcement learning training model is used to obtain the optimal evolution path, so as to achieve high fidelity, high speed and high robustness phonon Fock state preparation.
The fidelity and efficiency of phonon Fock state preparation is improved, the adaptability of the preparation is enhanced, and the robustness of the model is improved in noise environments.
Smart Images

Figure CN120087489A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of quantum computing, and particularly relates to a method for preparing phonon Fock states based on deep reinforcement learning. Background Art
[0002] In the field of quantum computing technology, phonon Fock states, as an important quantum state, can provide good quantum resources for realizing high-precision quantum operations and quantum logic gates. The manipulation of such quantum states usually relies on sideband operations, among which the most typical is the first-order red sideband pulse, whose frequency has a red detuning of one phonon frequency compared with the energy level resonance frequency, so as to achieve and the transition between. However, simple red sideband pulses face problems of insufficient speed and fidelity.
[0003] In the prior art, methods for improving fidelity include stimulated Raman adiabatic passage, dynamical decoupling, pulse shaping, and adiabatic shortcut. The noise resistance of stimulated Raman adiabatic passage and dynamical decoupling is achieved at the cost of time. Pulse shaping has poor noise resistance. Since the adiabatic shortcut is based on the Hamiltonian of the JC model and does not contain non-resonant excitation terms, there will be a large deviation from the experimental results and the robustness is poor.
[0004] In summary, the prior art has problems such as non-optimal evolution paths, limited fidelity, and poor noise robustness performance in the process of preparing phonon Fock states. Summary of the Invention
[0005] Based on this, the technical solution provided by the present invention aims to construct a neural network model by introducing a Hamiltonian containing non-resonant excitation terms as the evolution dynamics, and obtain the optimal evolution path through the training of deep reinforcement learning, so as to achieve the preparation of phonon Fock states with high fidelity, high speed, and high robustness.
[0006] To achieve the above object, the present invention provides a method for preparing phonon Fock states based on deep reinforcement learning, the method comprising: S1: Cooling a single ion in an ion trap to the phonon Fock ground state, and obtaining an optimal evolution path through a neural network model of deep reinforcement learning. S2: Controlling the laser sideband pulse to evolve according to the optimal evolution path, and driving the single ion with the laser sideband pulse and the carrier wave, so that the single ion evolves from the phonon Fock ground state to the phonon Fock target state.
[0007] Further, the step of obtaining the optimal evolution path through the neural network model of deep reinforcement learning includes:
[0008] S11: Construct the neural network model, where the neural network model uses the Hamiltonian containing a non-resonant excitation term as the evolution dynamics. Use the linear combination of the action reward function and the final fidelity as the total reward function, and the specific form is: , where R is the total reward function, is the final fidelity, is the weight coefficient of the final fidelity F, k is the weight coefficient of the action reward function, N is the total number of steps in a single model training, is the step identifier, is the action reward function at the i-th step.
[0009] S12: Initialize the network parameters of the neural network model.
[0010] S13: Use a multi-step progressive method for model training and output the optimal evolution path.
[0011] Preferably, the specific form of the Hamiltonian is: , where, is the annihilation operator, is the creation operator, is the Hamiltonian at the i-th step, is the Rabi frequency at the i-th step, is the detuning at the i-th step, is the reduced Planck constant, is the upward transition operator of the two-level internal state of the single ion, is the vibration frequency of the single ion, is the imaginary unit, t is the time, represents the Hermitian conjugate.
[0012] Furthermore, step S11 further includes performing parameter substitution on the detuning and the Rabi frequency in the Hamiltonian, and the specific form is: , where, is 's substitution parameter, is 's substitution parameter, is the Lamb-Dicke parameter, is the reference Rabi frequency, is the AC Stark frequency, and , 's value ranges are both .
[0013] Further, step S12 specifically includes: setting the initial phonon state and the final phonon state for model training, where the final phonon state differs from the initial phonon state by one phonon, setting the reference Rabi frequency, calculating the evolution time, and setting the total number of steps for a single model training.
[0014] Further, step S13 specifically includes: setting the weight coefficient k of the action reward function to decrease with the number of training times, and setting the weight coefficient m of the final fidelity to increase with the number of training times. Adjusting the state parameters according to the total reward function R, using the evolution path of the action parameters as the output of a single model training until the model converges, and obtaining the optimal evolution path. Among them, the state parameters include the weight coefficient m of the final fidelity and the weight coefficient k of the action reward function, and the action parameters include the replacement parameter of the detuning amount and the replacement parameter of the Rabi frequency.
[0015] Furthermore, the phonon Fock state preparation method further includes: setting the Rabi frequency with noise as and setting the detuning amount with noise as Replacing the Rabi frequency with the Rabi frequency with noise, replacing the detuning amount with the detuning amount with noise, calculating the Hamiltonian, and obtaining the optimal evolution path in a noisy environment.
[0016] Among them, the random frequency noise has a value range of ±5% of the reference Rabi frequency , and the random intensity noise has a value range of ±10% of the reference Rabi frequency .
[0017] The beneficial effects achieved by the present invention through the above technical solutions are:
[0018] (1) By introducing the Hamiltonian containing non-resonant excitation terms as the evolution dynamics to construct a neural network model, and training the model based on deep reinforcement learning to obtain the optimal evolution path of the laser. Compared with the traditional fixed evolution path, such as the evolution path of the pulse, not only the fidelity of phonon Fock state preparation is improved, but also the preparation efficiency is improved, and the adaptability of phonon Fock state preparation is enhanced.
[0019] (2) By adding random noise to the environment for training, the fault tolerance ability of the laser sideband pulse control to actual frequency and intensity perturbations is improved, and the robustness of the model is enhanced. Description of the Drawings
[0020] Figure 1 is the method flow chart of phonon Fock state preparation based on deep reinforcement learning in an embodiment of the present invention.
[0021] Figure 2 is the optimal evolution path for different model trainings in the embodiments of the present invention. Among them, Figure 2 (a) is the optimal evolution path with a total evolution time of 46.15 ; Figure 2 (b) is the optimal evolution path with a total evolution time of 23.08 ; Figure 2 (c) is the optimal evolution path with a total evolution time of 11.54 ; Figure 2 (d) is the optimal evolution path with a total evolution time of 5.77 ;
[0022] Figure 3 is a schematic diagram of the fidelity comparison between pulses and deep reinforcement learning in the noise environment of the embodiments of the present invention. Among them, ; Figure 3 (a) is the fidelity distribution diagram of the pulse in the noise environment, Figure 3 and (b) is the fidelity distribution diagram of deep reinforcement learning in the noise environment.
[0023] Figure 4 is a schematic diagram of the evolution of the phonon Fock state population using the optimal evolution path in the embodiments of the present invention. Detailed implementation manners
[0024] In order to make the purpose, technical solutions, and advantages of the present invention clearer and more understandable, the following further describes the detailed implementation manners of the present invention in combination with embodiments. It should be understood that the embodiments described herein are only used to explain the present invention, but not to limit the scope of the present invention.
[0025] Please refer to the attached Figure 1 , Figure 1 which is a flowchart of the method for preparing the phonon Fock state based on deep reinforcement learning in the embodiments of the present invention. As can be seen from the figure, the method for preparing the phonon Fock state includes: S1: Cooling a single ion in an ion trap to the phonon Fock ground state, and obtaining the optimal evolution path through the neural network model of deep reinforcement learning. S2: Controlling the laser sideband pulse to evolve according to the optimal evolution path, and driving the single ion with the laser sideband pulse and the carrier wave, so that the single ion evolves from the phonon Fock ground state to the phonon Fock target state.
[0026] Exemplarily, in step S1 of the above method for preparing the phonon Fock state, it specifically includes cooling a single ion in an ion trap to the phonon Fock ground state through Doppler cooling, EIT cooling, and sideband cooling. Ideally, the number of phonons in the phonon Fock ground state is zero. In the embodiments of the present invention, the average number of phonons in the experimentally obtained phonon Fock ground state is 0.1. In one or other embodiments of the present invention, the average number of phonons in the phonon Fock ground state is generally between 0.05 and 0.1.
[0027] Exemplarily, the steps of obtaining the optimal evolution path through the neural network model of deep reinforcement learning include:
[0028] (1) Establishing an environment, including setting the evolution dynamics and the total reward function.
[0029] In the embodiments of the present invention, the neural network model uses a Hamiltonian containing a non-resonant excitation term as the evolution dynamics in the environment, and its specific form is: , (1) where is the annihilation operator, is the creation operator, is the Hamiltonian at the i-th step, is the Rabi frequency at the i-th step, is the detuning at the i-th step, is the reduced Planck constant, is the upward transition operator of the two-level internal state of the single ion, is the vibration frequency of the single ion, is the imaginary unit, t is time, represents the Hermitian conjugate.
[0030] Taking the linear combination of the action reward function and the final fidelity as the total reward function, the specific form is: , (2) where R is the total reward function, is the final fidelity, is the weight coefficient of the final fidelity F, k is the weight coefficient of the action reward function, N is the total number of steps for a single model training, is the step identifier, is the action reward function at the i-th step.
[0031] In the embodiments of the present invention, according to the experience of the adiabatic shortcut, the replacement parameter of the detuning changes in a shape that is approximately close to a straight line, while the replacement parameter The shape should be close to the sine function. The actual training results may not fit very well with the reference adiabatic shortcut curve because, in addition to the stepwise action reward, there is also a reward based on the final fidelity. The ratio between the action reward and the final fidelity reward needs to be continuously tried and adjusted to successfully converge and obtain the best comprehensive effect.
[0032] In expression (1), the laser detuning and the Rabi frequency are parameters to be controlled, which are determined by the actions output by the model agent. To facilitate the operation of the reinforcement learning algorithm, in the embodiments of the present invention, the two parameters of the detuning and the Rabi frequency are replaced and scaled to the range of (0, 1). The specific form is: , (3) , (4) where is the replacement parameter for , is the replacement parameter for , is the Lamb-Dicke parameter, is the reference Rabi frequency, is the AC Stark frequency, and , both have a value range of .
[0033] In the embodiments of the present invention, the two parameters for parameter substitution in the evolution dynamics are the detuning and the Rabi frequency respectively. The variation range of these two parameters is determined by the magnitude of the reference Rabi frequency of the laser sideband pulse, and the magnitude of the reference Rabi frequency of the laser sideband pulse also determines the total evolution time: .
[0034] In the embodiments of the present invention, an arbitrary waveform generator is used to input an acousto-optic modulator through a radio frequency power amplifier to control the frequency and intensity of light, thereby realizing the control of the laser detuning and the Rabi frequency. In one or other embodiments of the present invention, other methods can also be used to control the two parameters of the laser detuning and the Rabi frequency. In one or other embodiments of the present invention, the evolution path of the laser can also be controlled by controlling other laser or ion parameters to realize the preparation of the phonon Fock state.
[0035] (2) Initialize the network parameters of the neural network model.
[0036] Exemplarily, the initialization step includes setting the initial phonon state and the final phonon state for model training. The final phonon state differs from the initial phonon state by one phonon.
[0037] Exemplarily, the initialization step further includes setting the reference Rabi frequency of the laser sideband pulse and the number of evolution steps for a single model training. The magnitude of the reference Rabi frequency of the laser sideband pulse determines the detuning amount of the laser sideband pulse and the variation amplitude of the Rabi frequency, and also determines the total evolution time. In the embodiment of the present invention, the number of evolution steps N for a single model training is set, and the total evolution time is determined according to the reference Rabi frequency. , then the evolution time for each step is . In the embodiment of the present invention, setting the total evolution time equal to the time of the red sideband pulse is only for the convenience of comparing the pulse with the evolution result of the neural network model of the deep reinforcement learning in the embodiment of the present invention. In one or other embodiments of the present invention, there is no limitation on the total evolution time.
[0038] (3) The model training is carried out in a multi-step progressive manner to output the optimal evolution path.
[0039] In the embodiment of the present invention, in the first step of training, a larger weight coefficient of the action reward function is set to obtain a path close to the adiabatic shortcut curve, and the action parameters of this evolution are output and . In subsequent training, the action parameters output from the previous step of training are loaded, and the weight coefficient k of the action reward function is reduced, while the weight coefficient m of the final fidelity is increased. Until the model converges, the optimal evolution path is obtained.
[0040] Please refer to the appendix Figure 2 , Figure 2 which is the optimal evolution path of different model trainings in the embodiment of the present invention. Figure 2 (a), Figure 2 (b), Figure 2 (c) and Figure 2 (d) respectively show the optimal evolution paths with the total evolution times of 46.15 , 23.08 , 11.54 and 5.77 . The abscissa in the figure represents the evolution time, the solid line represents the evolution path of the population number of the phonon Fock state as , the short dashed line represents 's optimal evolution path, and the dotted line represents 's optimal evolution path. By comparing the results of different model trainings, it can be seen that when the total evolution time is relatively large, 's evolution gradually tends to be flat, while 's evolution amplitude is relatively large. When the total evolution time is relatively small, 's maximum value decreases and the evolution is relatively gentle. In particular, when the total evolution time is relatively small, The value that evolves over time increases. This is because when the total evolution time is small, the non-resonant excitation (i.e., the carrier transition) has a greater impact on the training, and the output by the model training is smaller. To improve the final fidelity, the model training will output a with a larger evolution amplitude.
[0041] In the embodiment of the present invention, when the total evolution time is 46.15 , the fidelity of the pulse is 0.9945, and the fidelity of the deep reinforcement learning is 0.9909. When the total evolution time is 23.08 , the fidelity of the pulse is 0.9777, and the fidelity of the deep reinforcement learning is 0.9990. When the total evolution time is 11.54 , the fidelity of the pulse is 0.8958, and the fidelity of the deep reinforcement learning is 0.9963. When the total evolution time is 5.77 , the fidelity of the pulse is 0.6585, and the fidelity of the deep reinforcement learning is 0.8596. As the basic pulse with constant frequency and intensity, the fidelity of the pulse will continuously decrease as the total evolution time decreases, while the final fidelity obtained by the optimal evolution path given by the deep reinforcement learning decreases more slowly and is better than the pulse. Except when the total evolution time is 46.15 , the final fidelity of the deep reinforcement learning in the other three groups is higher than that of the pulse, proving that the deep reinforcement learning has a certain resistance to non-resonant excitation. However, under the condition that the total evolution time is 46.15 , the final fidelity of the deep reinforcement learning is slightly lower than that of the pulse. This is because when the total evolution time is large, the non-resonant excitation has little impact on the preparation of the phonon Fock state, and the advantage of the technical solution provided by the present invention over the pulse is not obvious. Moreover, in the embodiment of the present invention, in the model of the deep reinforcement learning, in order to make the action parameters and approach the curve of the adiabatic shortcut to improve the robustness, a part of the high fidelity is sacrificed. In one or other embodiments of the present invention, for the model with a total evolution time of 46.15 , if the weight coefficient k of the action reward function is set to zero in the total reward function, the final fidelity obtained by the training can exceed that of the pulse, but such a setting will make this model lack robustness like the pulse.
[0042] In the embodiment of the present invention, a model with a total evolution time of 23.08 is selected to train the model for each step between the phonon Fock state and to . First, the model from to is trained. The next model is trained based on the previous model, and the reference Rabi frequency of the laser sideband pulse is changed to , where is the number of phonons in the initial phonon state. According to the optimal evolution path of deep reinforcement learning, the fidelity of the manipulation of the phonon Fock state from to is 0.9992, to is 0.9991, to is 0.9992, to is 0.9998, to is 0.9999. By synthesizing the fidelity of each step, the fidelity from to is 0.9972. In the embodiment of the present invention, the fidelity of each step is continuously improved. This is because each step of the model is trained based on the previous step, indicating that more training leads to further convergence. It is also because the larger the number of phonons, the larger the Rabi frequency of the sideband transition. Under the condition of the same evolution time, the required Rabi frequency for the transition of a larger number of phonons is smaller, so the non-resonant excitation is also smaller.
[0043] In the embodiment of the present invention, in order to train an evolution path that is robust to frequency and intensity noise, an environment with added random noise is also created for further model training. In this noise environment, random frequency noise and random intensity noise are added simultaneously, and their amplitudes are ±5% and ±10% of the reference Rabi frequency respectively. The random intensity noise is added to the Rabi frequency: . The random frequency noise is added to the detuning: .
[0044] Please refer to Figure 3 , Figure 3 which is a schematic diagram comparing the fidelity of the pulse and deep reinforcement learning in the noise environment of the embodiment of the present invention. Figure 3 (a) is the fidelity distribution diagram of the pulse in the noise environment, Figure 3(b) is the fidelity distribution diagram of deep reinforcement learning in a noisy environment. In the figure, the abscissa represents the random intensity noise , and the ordinate represents the random frequency noise . To further compare the fidelity of the pulse and deep reinforcement learning, the embodiments of the present invention perform white filling on the positions with a fidelity of 0.96. In the embodiments of the present invention, the random ranges of the random frequency noise and the random intensity noise are respectively set to ±5% and ±10% of the reference Rabi frequency. In one or other embodiments of the present invention, the random range can also take other values. Each training will randomly select a value within this range and add it to the evolution dynamics on average. By continuously repeating the training, the model finally has a certain resistance to the noise within this range. The shade of gray in the figure represents the fidelity, and the darker the color, the higher the fidelity. As can be seen from the figure, the highest fidelity of the deep reinforcement learning path is slightly lower than the highest fidelity of the pulse, indicating that it is difficult for the model of deep reinforcement learning to further converge near the limit value. However, the range where the final fidelity of the deep reinforcement learning path is greater than 0.96 is larger, indicating that the deep reinforcement learning path has better robustness to frequency and intensity noise. The highest fidelity of the pulse is roughly located at the center of the figure, that is, when the values of the random frequency noise and the random intensity noise are equal to the reference Rabi frequency of the laser sideband pulse. The peak of the fidelity of deep reinforcement learning deviates from the center point, indicating that the model has not fully and evenly learned the experience under different magnitudes of noise during training. In one or other embodiments of the present invention, by increasing the number of training times of deep reinforcement learning, the model evenly learns the experience under different magnitudes of noise.
[0045] Please refer to the appendix Figure 4 , Figure 4 which is a schematic diagram of the evolution of the phonon Fock state population using the optimal evolution path in the embodiments of the present invention. In the figure, the abscissa represents the evolution time, and the ordinate represents the population of the initial evolution phonon state . In the embodiments of the present invention, the total evolution time is set to 46.15 , the initial evolution phonon state is , and the final evolution phonon state is . At this time, the final fidelity is 0.93, which has a certain gap with the theoretical value of 0.99. This indicates that when cooling the single ion in the ion trap to the phonon Fock ground state in step S1, the single ion is not cooled to the state where the phonon number is zero, and the average phonon number at this time is approximately 0.1. And the average phonon number of the phonon Fock ground state being 0.1 means that it will bring at least a 0.1 loss in the final fidelity, and at the same time will cause a decrease in the fidelity of the carrier pulse. As shown in the figure, the phonon state population corresponding to t = 0 is not equal to 1, indicating that the initial phonon state used in the experiment is not perfect 。
[0046] In one or other embodiments of the present invention, only the phonon final state is used as the judgment basis, the higher the population, the greater the reward. At this time, the reward function is set as: , the significance of using the square in the function is that when the final fidelity approaches 1, the gradient of the reward will continuously increase, which can avoid the problem of difficult further convergence near the limit value to a certain extent.
[0047] Finally, it should be noted that the above specific embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for preparing phonon Fock states based on deep reinforcement learning, characterized in that: The method includes: S1: Cooling a single ion in an ion trap to the phonon Fock ground state; obtaining the optimal evolution path through a deep reinforcement learning neural network model; S2: controlling the laser sideband pulse to evolve according to the optimal evolution path, and driving the single ion with the laser sideband pulse and the carrier, so that the single ion evolves from the phonon Fock ground state to the phonon Fock target state; Wherein, the neural network model adopts the Hamiltonian containing non-resonant excitation terms as the evolutionary dynamics.
2. The method for preparing phonon Fock states according to claim 1, characterized in that: The step of obtaining the optimal evolution path through the neural network model of deep reinforcement learning includes: S11: Construct the neural network model, which uses a linear combination of the action reward function and the final fidelity as the total reward function, and the specific form is: , Where R is the total reward function, For the final fidelity, is the weight coefficient of the final fidelity F, k is the weight coefficient of the action reward function, N is the total number of steps for a single model training, is the step number indicator, is the action reward function of the i-th step; S12: Initialize the network parameters of the neural network model; S13: Perform model training in a multi-step progressive manner and output the optimal evolution path.
3. The method for preparing phonon Fock states according to claim 2, characterized in that: The specific form of the Hamiltonian is: , in, is the annihilation operator, is the Hamiltonian at the ith step, is the Rabi frequency of the ith step, is the detuning amount at step i, is the reduced Planck constant, is the rising transition operator of the second energy level of the single ion internal state, is the vibration frequency of the single ion, is the imaginary unit, t is the time, represents Hermitian conjugation.
4. The method for preparing phonon Fock states according to claim 3, characterized in that: Step S11 also includes: The detuning amount and the Rabi frequency in the Hamiltonian are replaced by parameters, and the specific form is: , in, for The replacement parameters, for The replacement parameters, is the Lamb-Dicke parameter, is the reference Rabi frequency, is the AC Stark frequency, and , The value range of .
5. The method for preparing phonon Fock states according to claim 2, characterized in that: Step S12 specifically includes: Setting the phonon initial state and the phonon final state for model training, wherein the phonon final state differs from the phonon initial state by one phonon; The reference Rabi frequency is set, the evolution time is calculated, and the total number of steps for the single model training is set.
6. The method for preparing phonon Fock states according to claim 2, characterized in that: Step S13 specifically includes: The weight coefficient k of the action reward function is set to decrease with the number of training times, and the weight coefficient m of the final fidelity is set to increase with the number of training times; Adjusting the state parameters according to the total reward function R, taking the evolution path of the action parameters as the output of a single model training, until the model converges, and obtaining the optimal evolution path; The state parameters include a weight coefficient m of the final fidelity and a weight coefficient k of the action reward function, and the action parameters include a replacement parameter of the detuning amount and a replacement parameter of the Rabi frequency.
7. The method for preparing phonon Fock states according to claim 2, characterized in that: The phonon Fock state preparation method further comprises: Set the Rabi frequency with noise to ; Set the detuning amount with noise to ; Using the noisy Rabi frequency to replace the Rabi frequency, using the noisy detuning amount to replace the detuning amount, calculating the Hamiltonian, and obtaining the optimal evolution path in a noisy environment; Among them, random frequency noise The value range is the reference Rabi frequency ±5% of random intensity noise The value range is the reference Rabi frequency ±10% of
Citation Information
Patent Citations
Hole machining path planning method based on multi-agent deep reinforcement learning
CN117273235A
Control method and device for AI intelligent calculation integration, electronic equipment and storage medium
CN119690676A
Simulating large cat qubits using a shifted fock basis
US20220156444A1
Reinforcement learning with quantum oracle
US20220253743A1
Methods and systems for improving an estimation of a property of a quantum state
US20230104058A1