A method for intelligent preparation of bosonic codes based on reinforcement learning
By combining reinforcement learning and Flokai engineering to optimize the Boson code preparation process, the problems of long preparation time and noise sensitivity in existing technologies have been solved, achieving efficient and robust quantum state control and promoting the practical application of quantum computing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHWESTERN POLYTECHNICAL UNIV
- Filing Date
- 2026-03-04
- Publication Date
- 2026-06-23
AI Technical Summary
Existing Bose code preparation techniques face problems such as long evolution times and sensitivity to noise, making it difficult to achieve stable and scalable quantum computing applications.
By combining reinforcement learning and Flokai engineering, the boson code preparation process can be autonomously adjusted by optimizing the amplitude and frequency of the external driving field, thereby achieving efficient and robust quantum state control.
It significantly shortens the Boson code preparation time, maintains high fidelity, adapts to different noise environments, and provides a scalable quantum computing solution.
Smart Images

Figure CN122264154A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of quantum computing technology, specifically relating to a method for preparing Boson codes based on reinforcement learning. Background Technology
[0002] Against the backdrop of rapid development in the field of quantum computing, quantum error correction is a core requirement for solving noise interference and improving computational reliability. Boson codes, as a key technology for quantum error correction, can embed quantum information into appropriate non-Gaussian encoded states. This not only specifically suppresses major error sources such as photon loss and depolarization noise, but also provides a hardware-efficient and scalable solution for fault-tolerant quantum computing in various platforms, including superconducting systems, ion traps, and optical systems, due to their lower physical overhead. However, the preparation and storage of Boson codes still face significant technical challenges. Traditional techniques using general gate operations and adiabatic methods not only require extremely long evolution times but are also highly sensitive to various noise sources such as photon loss and depolarization noise, and have stringent requirements for coherence time and control precision. To achieve stable and scalable Boson code applications and build practical continuous-variable fault-tolerant quantum computing systems, it is still necessary to overcome the core bottlenecks in efficiency, robustness, and scalability of existing preparation methods.
[0003] The rapid development of artificial intelligence (AI) technology has greatly promoted progress in many scientific fields, including quantum technology. Among them, the Floquet Engineering method assisted by reinforcement learning, relying on its autonomous parameter optimization capabilities and high noise resistance, has become an innovative solution to the bottleneck of Bose code preparation in continuous-variable quantum systems. This method organically combines the Floquet Engineering scheme with reinforcement learning. The Floquet Engineering scheme is responsible for controlling the quantum system through an external periodic driving field, providing the quantum physical framework for Bose code preparation; the reinforcement learning agent observes the density matrix of the quantum system in real time, uses the target state fidelity as the core reward signal, autonomously updates the driving parameters (amplitude, frequency), and dynamically adjusts the evolution strategy, optimizing the preparation process without human intervention. This collaborative mode not only significantly shortens the Bose code preparation time but also maintains high fidelity in strong noise environments, providing a practical path for high-quality Bose code preparation and promoting the practical application of continuous-variable quantum technology. Summary of the Invention
[0004] To overcome the problems of long evolution times and sensitivity to noise in existing Bose code preparation techniques, this invention provides a reinforcement learning-based intelligent Bose code preparation method. This method combines reinforcement learning algorithms with Floquet engineering, optimizing key parameters such as the amplitude and frequency of the external driving field to efficiently prepare Bose codes. This approach significantly shortens the preparation time of Bose codes and maintains high fidelity even in noisy environments. It not only demonstrates the powerful capabilities of artificial intelligence in quantum control but also establishes a scalable and experimentally feasible approach for fault-tolerant Bose quantum computing. Beyond the specific application of Bose code preparation, this invention also provides a general paradigm for integrating machine learning and Floquet engineering, potentially solving the decoherence problem in next-generation quantum technologies.
[0005] The technical solution adopted by this invention to solve its technical problem is as follows: Step 1: Set the current time step to... Obtain the density matrix in a quantum system And parse it into a state vector ; Step 2: Input into the policy network of reinforcement learning In the process, the action vector is obtained. ,pass The frequencies of the external driving fields were analyzed separately. With amplitude The evolution of the quantum system is driven by this external driving field; Step 3: Obtain the density matrix of the next state from the quantum system From the target state and Calculate the fidelity ; Step 4: Verify fidelity Parse the reward information using the reward function. ; Step 5: Obtain the information from the above steps. Store in the experience pool In, among them, For a current time step The state vector, For the current time step The action vector, For the current time step Reward information, For the next time step The state vector; repeat steps 1 to 5 until the minimum training requirement is met; Step 6: Continuously draw experience from the pool A batch of data of a random size is selected for training. Step 7: After training, proceed to the prediction phase; obtain the density matrix from the quantum system. From the density matrix Parse the state vector ; Step 8: Convert the state vector Input into the trained policy network Obtain action vectors , by action vector Frequency analysis and amplitude And input it into the quantum system for evolution; Step 9: Repeat steps 7 and 8 until the evolution ends.
[0006] Preferably, step 2 specifically comprises: The state vector Input to policy network , to obtain action The specific formula is shown below: ,
[0007] pass Solve for the amplitude of the external driving force of the quantum system and frequency The specific formula is as follows: ,
[0008] in, , , These are preset fixed parameters.
[0009] Preferably, step 3 specifically comprises: Calculate using the following formula and fidelity :
[0010]
[0011]
[0012]
[0013] in, This indicates finding the trace, which is the sum of the elements on the diagonal of a square matrix; Indicates the target Boson code state; Represents a single quantum bit; Represents the normalization constant; The multiplicity of a rotation-symmetric code; Represents the normalization base; The density matrix representing the target Boson code state; Indicates actual The density matrix of the boson code states at time t.
[0014] Preferably, step 4 specifically comprises: Through the fidelity in step 3 Parse the reward information using the reward function. :
[0015]
[0016]
[0017]
[0018] in, Indicates a reward for success; Indicates the preset time window size; express Fidelity at that time; , As an intermediate variable, record The number of times the number of times the threshold is exceeded; Indicates a stability reward; Preferably, step 6 specifically comprises: Step 6-1: From the experience pool Get a batch size of data (in) ); Step 6-2: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require and Input into the action value network Obtain the action value vector The formula is shown below:
[0019] Step 6-3: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require Input to the target policy network The action vector for the next state is obtained from this. The formula is shown below:
[0020]
[0021] in, To edit noise, It is random noise. c This is the noise clipping threshold constant; It's a clipping function, and the upper and lower bounds are set to - c arrive c ; Step 6-4: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require and Input into the dual-objective value network The time-series difference objective is obtained, as shown in the following formula:
[0022] in, This represents the reward information at time t. Indicates the discount factor; Step 6-5: Update the action value network by minimizing the temporal difference loss. The formula is as follows:
[0023] in, Indicates the network parameters of the action value; This indicates the number of samples in a batch; Step 6-6: Using the deterministic gradient descent algorithm Update policy network To maximize the objective function The formula is as follows:
[0024]
[0025] in, Indicates the experience pool The expected value is obtained from the sampled state vector.
[0026] Therefore, the expression for the gradient descent algorithm is:
[0027] Steps 6-7: Soft update target policy network Soft update target value network and The formula is as follows:
[0028]
[0029] in, For the first Parameters of the value network for each target action; These are the parameters of the target policy network; These are the parameters of the current policy network; This is the soft update coefficient.
[0030] The beneficial effects of this invention are as follows: Existing subsystem control methods largely rely on manual experience to design drive pulse sequences, failing to adapt to inherent noise characteristics such as quantum state decoherence and photon loss, thus limiting the fidelity of quantum state evolution and the robustness of control strategies. This invention, based on the core framework of the reinforcement learning TD3 algorithm, can more accurately capture the dynamic evolution of quantum systems, significantly improving the stability of the training process and enhancing the fidelity and convergence efficiency of quantum state preparation and manipulation tasks.
[0031] Furthermore, the method of this invention possesses strong scalability and is applicable to various quantum technology scenarios. By flexibly adjusting noise parameters, it can adapt to quantum platforms with different noise intensities and diverse control requirements, further optimizing performance in specific scenarios. This invention provides a new technical paradigm for the autonomous optimization control of quantum systems, promotes the deep integration of reinforcement learning and quantum technology, and has significant theoretical innovation value and broad engineering application prospects. Attached Figure Description
[0032] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a schematic diagram illustrating the specific values of amplitude and frequency during the evolution process; Figure 3 This is a schematic diagram illustrating fidelity during the evolution process; Figure 4 This is a schematic diagram illustrating the evolution of the stroboscopic time of the prepared state fidelity under different environmental noise levels when using the method of the present invention. Detailed Implementation
[0033] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0034] The innovation of this invention lies in introducing algorithms from reinforcement learning and designing corresponding reward functions to dynamically optimize the external driving parameters during the quantum Bose code state preparation process. The reinforcement learning agent can observe the density matrix of the quantum system in real time and autonomously update the key parameters (amplitude and frequency) of the external driving field, using the reward function as the optimization objective, achieving breakthroughs in efficiency and robustness.
[0035] The following section will provide a detailed explanation of the methods, procedures, and limitations of the parameters.
[0036] The first step is the main training process: Step 1: Set the current time step to... Obtain the density matrix of the quantum system And parse it into a state vector .
[0037] Step 2: Obtain the above-mentioned... Input into the policy network of reinforcement learning In the process, the action vector is obtained. .pass Analyze the parameter frequencies of the external driving field respectively With amplitude The evolution of the quantum system is driven by this external driving field.
[0038] The above obtained Input to policy network , thus obtaining the action vector The specific formula is as follows: ,
[0039] pass Solve for the frequency of external drive of the quantum system With amplitude The specific formula is as follows: ,
[0040] in, , , These are preset fixed parameters.
[0041] Step 3: Obtain the density matrix of the next state from the quantum system From the target state and Calculate the fidelity .
[0042] Calculate using the following formula and fidelity :
[0043]
[0044]
[0045]
[0046] Step 4: Verify the fidelity in Step 3 Parse the reward information using the reward function. The specific formula is as follows:
[0047]
[0048]
[0049]
[0050] Step 5: Obtain the information from the above steps. Store in the experience pool In, among them, For a current time step The state vector, For the current time step The action vector, For the current time step Reward information, For the next time step The state vector; repeat steps 1 to 5 until the minimum training requirement is met; Step 6: Continuously draw experience from the pool A batch of data of a random size is selected for training.
[0051] Step 6.1: From the experience pool To obtain a batch size of data, use ( ).
[0052] Step 6.2: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require and Input into the action value network Obtain the action value vector The formula is shown below:
[0053] Step 6.3: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require Input to the target policy network The action vector for the next state is obtained from this. The formula is shown below:
[0054]
[0055] Step 6.4: [The sentence is incomplete and requires more context to be translated accurately.] and Input into the dual-objective value network And obtain the time-series difference objective, as shown in the following formula:
[0056] Step 6.5: Update the action value network by minimizing the temporal difference loss. The formula is as follows:
[0057] Step 6.6: Using the deterministic gradient descent algorithm Update policy network To maximize the objective function The formula is as follows:
[0058]
[0059] Therefore, the expression for the gradient descent algorithm is:
[0060] Step 6.7: Soft update target policy network Soft update target value network The formula is as follows:
[0061]
[0062] Secondly, the main steps of reasoning: Step 7: After training, proceed to the prediction phase. Obtain the density matrix from the quantum system. From the density matrix Parse the state vector .
[0063] Step 8: Convert the state vector Input into the trained policy network Obtain action vectors , by action vector Frequency analysis and amplitude And it is input into the quantum system for evolution.
[0064] Obtained from step 7 Input to policy network , to obtain action ,pass Calculate the frequency of the external drive of the quantum system and amplitude The specific formula is the same as in step 2. Check the fidelity at this point. If the fidelity is too low in high noise, retraining can be performed, using the same formula as in step 3.
[0065] Step 9: Repeat steps 7 and 8 until the evolution ends.
[0066] Example: This invention proposes a quantum state control scheme that integrates reinforcement learning and Floquet Engineering. It integrates the reinforcement learning TD3 algorithm mechanism and is supplemented by a phased exploration of noise attenuation and retraining strategy in high-noise environments. The aim is to make full use of the autonomous optimization capability of reinforcement learning to adapt to the dynamic evolution characteristics of quantum systems and solve key problems in existing quantum state preparation methods, such as low efficiency of artificially designed driving field parameters, poor noise adaptability, and large fluctuations in the training process.
[0067] The purpose of this invention is to provide a high-fidelity and robust autonomous preparation and control scheme for quantum states suitable for noisy quantum platforms, so as to realize the precise evolution control of quantum systems in complex noisy environments.
[0068] like Figure 1 The present invention is further illustrated with reference to the accompanying drawings and using the construction of a four-component cat state capable of correcting single-photon loss error as an example. The specific inventive steps are as follows: Step 1: Set the current time step to... Obtain the density matrix of the quantum system And parse it into a state vector .
[0069] In step 2, the above-obtained Input to policy network , to obtain action The obtained driving field parameters satisfy: ,
[0070] In this case, the setting is as follows , .
[0071] In step 3, the density matrix of the next state is obtained from the quantum system. From the target state and Calculate the fidelity Calculate using the following formula and fidelity :
[0072]
[0073]
[0074]
[0075] in, This indicates finding the trace, which is the sum of the elements on the diagonal of a square matrix; Indicates the target Boson code state; Represents a single quantum bit; Represents the normalization constant; The multiplicity of a rotation-symmetric code; Represents the normalization base; The density matrix representing the target Boson code state; Indicates actual The density matrix of the boson code states at time t.
[0076] Step 4: Parse the reward information using fidelity and reward function. The specific formula is as follows:
[0077]
[0078]
[0079]
[0080] Step 5: Place ( Store in the experience pool Repeat steps 1 to 5 above until you reach 20,000 training sessions.
[0081] Step 6: From the experience pool A batch of data of a random size is selected for training.
[0082] Step 6.1: From the experience pool To obtain a batch size of data, use ( ).
[0083] Step 6.2: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require and Input into the action value network Obtain the action value vector The formula is shown below:
[0084] Step 6.3: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require Input to the target policy network The action vector for the next state is obtained from this. The formula is shown below:
[0085]
[0086] Step 6.4: [The sentence is incomplete and requires more context to be translated accurately.] and Input into the dual-objective value network And obtain the time-series difference objective, as shown in the following formula:
[0087] Step 6.5: Update the action value network by minimizing the temporal difference loss. The formula is as follows:
[0088] Step 6.6: Using the deterministic gradient descent algorithm Update policy network To maximize the objective function The formula is as follows:
[0089]
[0090]
[0091] Therefore, the expression for the gradient descent algorithm is:
[0092] Step 6.7: Update the target policy network Update the target value network The formula is as follows:
[0093]
[0094] Step 7: After training, proceed to the prediction phase. Obtain the density matrix from the quantum system. From the density matrix Parse the state vector .
[0095] Step 8: Obtained from Step 7 Input to policy network , to obtain action ,pass Solve for the amplitude of the external driving force of the quantum system and frequency , the specific formula is the same as step 2.
[0096] Step 9: Repeat Step 7. Step 8 directly ends the evolution process, as shown below. Figure 2 The amplitude and frequency are shown in the figure. Figure 3 The fidelity is shown.
[0097] This embodiment uses reinforcement learning to prepare quadrupedal cat states and compares it with the traditional adiabatic evolution scheme. The preparation time of the method of this invention is significantly reduced compared with the traditional scheme. For specific results, see [link to documentation]. Figure 2 Meanwhile, the solution using this invention exhibits high robustness and adaptability to high noise levels after retraining, such as... Figure 3 and Figure 4 As shown. Among them, Figure 3 The horizontal axis represents the evolution period. Figure 4 The solid black line in the middle represents the area where Single-photon loss rate noise, The original strategy evolution fidelity value under decoherence rate; the black dashed line represents the value at which the original strategy evolves. The fidelity value of the original strategy evolution under single-photon loss rate noise; the black dotted line represents... The fidelity of the retrained policy evolution under single-photon loss rate noise. Figure 4 middle , These represent the single-photon loss rate and the decoherence rate, respectively. The larger the environment, the more severe the photon loss. The larger the phase noise, the stronger the interference to the quantum state, as shown in the following formula.
[0098]
[0099] This result not only demonstrates the effectiveness of Floquet Engineering-based control methods in continuous variable systems but also showcases the ability of machine learning techniques to discover nontrivial optimal control strategies. The findings highlight the potential of combining artificial intelligence with quantum control, a combination that could serve as a powerful paradigm for realizing fault-tolerant Bose quantum computing. This framework is expected to be further extended to various quantum platforms and noise models, providing a scalable path for the practical application of quantum error-correcting codes in near-term quantum hardware.
Claims
1. A method for preparing Boson codes based on reinforcement learning, characterized in that, Includes the following steps: Step 1: Set the current time step to... Obtain the density matrix in a quantum system And parse it into a state vector ; Step 2: Input into the policy network of reinforcement learning In the process, the action vector is obtained. ,pass The frequencies of the external driving fields were analyzed separately. With amplitude The evolution of the quantum system is driven by this external driving field; Step 3: Obtain the density matrix of the next state from the quantum system From the target state and Calculate the fidelity ; Step 4: Verify fidelity Parse the reward information using the reward function. ; Step 5: Obtain the information from the above steps. Store in the experience pool In, among them, For a current time step The state vector, For the current time step The action vector, For the current time step Reward information, For the next time step The state vector; repeat steps 1 to 5 until the minimum training requirement is met; Step 6: Continuously draw experience from the pool A batch of data of a random size is selected for training. Step 7: After training, proceed to the prediction phase; obtain the density matrix from the quantum system. From the density matrix Parse the state vector ; Step 8: Convert the state vector Input into the trained policy network Obtain action vectors , by action vector Frequency analysis and amplitude And input it into the quantum system for evolution; Step 9: Repeat steps 7 and 8 until the evolution ends.
2. The method for intelligent preparation of Boson codes based on reinforcement learning according to claim 1, characterized in that, Step 2 specifically involves: The state vector Input to policy network , to obtain action The specific formula is shown below: , pass Solve for the amplitude of the external driving force of the quantum system and frequency The specific formula is as follows: , in, , , These are preset fixed parameters.
3. The method for intelligent preparation of Boson codes based on reinforcement learning according to claim 2, characterized in that, Step 3 specifically involves: Calculate using the following formula and fidelity : in, This indicates finding the trace, which is the sum of the elements on the diagonal of a square matrix; Indicates the target Boson code state; Represents a single quantum bit; Represents the normalization constant; The multiplicity of a rotation-symmetric code; Represents the normalization base; The density matrix representing the target Boson code state; Indicates actual The density matrix of the boson code states at time t.
4. The method for intelligent preparation of Boson codes based on reinforcement learning according to claim 3, characterized in that, Step 4 specifically involves: Through the fidelity in step 3 Parse the reward information using the reward function. : in, Indicates a reward for success; Indicates the preset time window size; express Fidelity at that time; , As an intermediate variable, record The number of times the number of times the threshold is exceeded; This indicates a stability bonus.
5. The method for intelligent preparation of Boson codes based on reinforcement learning according to claim 4, characterized in that, Step 6 specifically involves: Step 6-1: From the experience pool Get a batch size of data (in) ); Step 6-2: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require and Input into the action value network Obtain the action value vector The formula is shown below: Step 6-3: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require Input to the target policy network The action vector for the next state is obtained from this. The formula is shown below: in, To edit noise, It is random noise. c This is the noise clipping threshold constant; It's a clipping function, and the upper and lower bounds are set to - c arrive c ; Step 6-4: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require and Input into the dual-objective value network The time-series difference objective is obtained, as shown in the following formula: in, This represents the reward information at time t. Indicates the discount factor; Step 6-5: Update the action value network by minimizing the temporal difference loss. The formula is as follows: in, Indicates the network parameters of the action value; This indicates the number of samples in a batch; Step 6-6: Using the deterministic gradient descent algorithm Update policy network To maximize the objective function The formula is as follows: in, Indicates the experience pool The expected value of the sampled state vector is taken. Therefore, the expression for the gradient descent algorithm is: Steps 6-7: Soft update target policy network Soft update target value network and The formula is as follows: in, For the first Parameters of the value network for each target action; These are the parameters of the target policy network; These are the parameters of the current policy network; This is the soft update coefficient.
6. An electronic device, characterized in that, include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the method as described in any one of claims 1 to 5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.
8. A chip, characterized in that, include: A processor for retrieving and running a computer program from memory, causing a device on which the chip is mounted to perform the method as described in any one of claims 1 to 5.
9. A computer program product, characterized in that, The computer program product includes a computer storage medium storing a computer program, the computer program including instructions executable by at least one processor, which, when executed by the at least one processor, implement the method as described in any one of claims 1 to 5.