A method for preparing a spin compressed GKP state based on quantum reinforcement learning
Patent Information
- Application Number
- CN202610959378.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-18
AI Technical Summary
[0004]尽管上述现有技术为GKP态制备提供了可行路径,但它们在应对自旋压缩GKP态这一复合体系的高保真度、高效率制备时,仍存在显著缺点
1.针对自旋压缩GKP态这一容错量子计算关键态的高保真度制备难题,本发明设计了适配强化学习的自旋压缩GKP态制备环境。该环境基于量子发射体系统的动力学特性,将相干旋转和自旋压缩操作作为基本控制,并以自旋压缩GKP态的保真度作为强化学习的奖励函数,为后续智能控制算法的训练奠定了物理基础。
Smart Images

Figure CN122779313A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of GKP state preparation technology, and in particular to a spin-squeezed GKP state preparation method based on quantum reinforcement learning. Background Technology
[0002] In the field of quantum computing, spin-squeezed GKP (Gottesman-Kitaev-Preskill) states, as composite quantum states combining spin-squeezing technology and GKP quantum error-correcting codes, have become a key resource for realizing fault-tolerant quantum computing due to their unique advantages in continuous-variable quantum computing and fault-tolerant mechanisms. However, their high-fidelity experimental preparation has always been a recognized technical challenge. Existing technologies mainly focus on the preparation of GKP states and related non-Gaussian states.
[0003] Traditional techniques can be broadly categorized into two types: one is based on deterministic nonlinear interactions, such as the preparation of Schrödinger cat states using instantaneous nonlinearity or direct squeezing in optical parametric oscillators; the other is based on probabilistic methods of post-measurement selection, such as photon number-resolved detection combined with subtractive photon manipulation or catalytic manipulation. To overcome the efficiency bottleneck of traditional methods, researchers have recently begun exploring the application of classical machine learning methods, such as deep reinforcement learning, to quantum control. A typical example is the use of time-multiplexed measurement-based quantum optical processors. By training deep neural networks to learn optimal control pulse sequences, squeezed cat states (GKP state precursors) can be generated with an average success rate of up to 98% and with high deterministic accuracy. This approach demonstrates the enormous potential of machine learning in optimizing complex quantum control problems, enabling adaptive parameter adjustment to approximate the target quantum state.
[0004] While the aforementioned existing technologies provide feasible pathways for GKP state preparation, they still have significant drawbacks when dealing with the high-fidelity and high-efficiency preparation of spin-squeezed GKP states in complex systems. Deterministic schemes in traditional methods are often limited by insufficient strength of nonlinear interactions, resulting in low fidelity and limited compression amplitude of the prepared GKP states, making it difficult to meet the requirements of fault-tolerant computation. Probabilistic schemes, although capable of generating non-Gaussian states with higher fidelity, suffer from inherent randomness leading to extremely low preparation success rates and require complex post-selection circuits and large data buffers. Deep reinforcement learning schemes, while far superior to other traditional methods, are mostly limited to purely optical platforms and have not yet been deeply integrated with spin-squeezing techniques in solid-state spin systems (such as diamond color centers or ion traps). Therefore, a robust control scheme that can combine the enhancement effect of spin-squeezing with the adaptive optimization capabilities of reinforcement learning and dynamically suppress decoherence effects is urgently needed. Summary of the Invention
[0005] The present invention aims to at least partially solve one of the technical problems in the related art.
[0006] The main objective of this invention is to provide a spin-squeezed GKP state preparation method based on quantum reinforcement learning. By constructing a spin-squeezed GKP state preparation environment adapted to reinforcement learning and introducing the exponential expression capabilities of quantum superposition and entanglement, high-fidelity and robust preparation of spin-squeezed GKP states can be achieved.
[0007] The second objective of this invention is to provide an electronic device.
[0008] A third objective of this invention is to provide a non-transitory computer-readable storage medium.
[0009] To achieve the above objectives, a first aspect of the present invention proposes a method for preparing spin-squeezed GKP states based on quantum reinforcement learning, comprising: S1. Construct a spin-squeezed GKP state preparation environment, establish a quantum system composed of N two-level quantum emitters, construct a quantum state evolution model based on the collective spin operator of the quantum system, and set the target spin-squeezed GKP state; S2. Construct a quantum-classical hybrid reinforcement learning model, which includes a policy network and a value network, and the policy network includes a parameterized quantum circuit. S3. Obtain the current quantum state output by the quantum state evolution model, input the current quantum state into the policy network, and generate control actions after processing by the parameterized quantum circuit; S4. Map the control action to the control parameters of the quantum state evolution model, update the quantum state evolution model using the control parameters and drive the quantum system to evolve, and obtain the quantum state at the next moment; S5. Calculate the reward value based on the fidelity between the next quantum state and the target spin-squeezed GKP state, and construct an empirical trajectory by combining the current quantum state, control action, reward value, and next quantum state. S6. Based on the empirical trajectory, the proximal policy optimization algorithm is used to update the parameters of the policy network and the value network, and steps S3 to S6 are repeated until the preset training conditions are met to obtain the control policy for preparing the target spin-compressed GKP state.
[0010] Optionally, step S1 includes: Establish a quantum system consisting of N two-level quantum emitters; Construct a collective spin operator based on the Pauli operators of each two-level quantum emitter in the quantum system; The collective spin operator is used to construct a system Hamiltonian that includes a collective spin term, a phase spin term, and a uniaxial torsional compression term.
[0011] Optionally, the system Hamiltonian, comprising a collective spin term, a phase rotation term, and a uniaxial torsional compression term, is constructed using the collective spin operator, including: Construct a collective rotation term using the Y-axis component of the collective spin operator; A phase rotation term is constructed using the Z-axis component of the collective spin operator; A uniaxial torsional compression term is constructed using the axial quadratic term in the collective spin operator; The collective rotation term, phase rotation term, and uniaxial torsional compression term are combined to form a system Hamiltonian, which is then used to generate quantum state evolution operators.
[0012] Optionally, step S2 includes: The current quantum state output by the quantum state evolution model is converted into environmental state characteristics; The environmental state characteristics are pre-encoded to obtain the quantum circuit input characteristics; The quantum circuit input characteristics are input into the parameterized quantum circuit to obtain quantum measurement results.
[0013] Optionally, the quantum circuit input characteristics are input into a parameterized quantum circuit to obtain quantum measurement results, including: The input features of the quantum circuit are encoded into qubits; Perform parameterized single-qubit rotation operations on the qubit; Entanglement operations are performed between qubits to obtain quantum states that carry environmental state correlation characteristics; The quantum state carrying the environmental state correlation characteristics is measured to obtain the quantum measurement result.
[0014] Optionally, step S4 includes: The control action is decomposed into multiple action components; The multiple motion components are respectively mapped to Y-axis rotation angle, Z-axis rotation angle and single-axis torsional strength; Substitute the Y-axis rotation angle, Z-axis rotation angle, and uniaxial torsional intensity into the quantum state evolution model to generate the quantum state evolution operator for the current control step; The quantum state evolution operator of the current control step is applied to the current quantum state to obtain the quantum state at the next moment.
[0015] Optionally, step S5 includes: Calculate the state overlap between the next quantum state and the target spin-squeezed GKP state; The fidelity is calculated based on the state overlap, and a reward value is generated based on the fidelity. The current quantum state, control action, reward value, next quantum state, and action probability are stored as an experience trajectory, which is then used to update the near-end policy optimization algorithm in step S6.
[0016] Optionally, step S6 includes: Read the state sequence, action sequence, reward sequence, and action probability from the experience trajectory; Calculate the state value corresponding to the state sequence using a value network; Calculate the temporal difference error based on the reward sequence and state value; The dominance function is calculated based on the time-series difference error. The pruning objective function of the near-end policy optimization algorithm is constructed based on the action probability and the advantage function. The policy network parameters are updated using the pruning objective function, and the value network parameters are updated using the value loss function until the preset training conditions are met.
[0017] To achieve the above objectives, a second aspect of this application provides an electronic device, including a processor and a memory; wherein the processor runs a program corresponding to the executable program code by reading executable program code stored in the memory, for implementing the method described in the first aspect embodiment.
[0018] To achieve the above objectives, a third aspect of this application provides a non-transitory computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method described in the first aspect.
[0019] The embodiments of the present invention have the following beneficial effects: 1. To address the challenge of high-fidelity preparation of spin-squeezed GKP states, a key state for fault-tolerant quantum computing, this invention designs a spin-squeezed GKP state preparation environment adapted for reinforcement learning. Based on the dynamic characteristics of the quantum emitter system, this environment uses coherent rotation and spin-squeezing operations as fundamental controls, and employs the fidelity of the spin-squeezed GKP state as the reward function for reinforcement learning, laying the physical foundation for the subsequent training of intelligent control algorithms.
[0020] 2. The classical PPO algorithm requires the design of complex neural networks in high-dimensional quantum systems, leading to a surge in parameters, high training costs, weak generalization ability, and difficulty in capturing the time-series dependence of the quantum environment. This invention proposes a quantum-classical hybrid proximal policy optimization algorithm. Using the classical PPO algorithm as a stable framework, it introduces parameterized quantum circuits as feature extractors for the policy network. Leveraging the quantum superposition and entanglement properties of PQC, it efficiently extracts quantum state features in high-dimensional Hilbert space with fewer parameters, thereby achieving higher fidelity and lower training overhead in spin-squeezed GKP state preparation. Attached Figure Description
[0021] The above-described and additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, in which: Figure 1 A text flowchart illustrating a spin-squeezed GKP state preparation method based on quantum reinforcement learning, provided for an embodiment of the present invention; Figure 2 A flowchart illustrating a spin-squeezed GKP state preparation method based on quantum reinforcement learning, provided for an embodiment of the present invention; Figure 3 A schematic diagram of the structure of the quantum-classical hybrid reinforcement learning model provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a policy network provided in an embodiment of the present invention; Figure 5 The comparative experimental results and parameter setting diagrams provided for embodiments of the present invention are shown in the figure. Detailed Implementation
[0022] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0024] This embodiment provides a spin-squeezed GKP state preparation method based on quantum reinforcement learning, comprising four parts: construction of the spin-squeezed GKP state preparation environment, design of a quantum-classical hybrid algorithm, model training, and experimental results. In the environment construction part, based on the dynamic characteristics of the quantum emitter system, a quantum state evolution model including coherent rotation and spin-squeezing operations is established. In the algorithm design part, the classical proximal policy optimization (PPO) algorithm is used as the core framework, and a parameterized quantum circuit (PQC) is innovatively used as the nonlinear feature extractor of the policy network. The exponential Hilbert space representation capability of the quantum circuit is utilized to enhance the policy network's ability to model complex quantum control policies. In the model training part, empirical trajectories of the spin-squeezed GKP state preparation process are collected through interactive sampling between the agent and the quantum environment, and the pruning of the objective function by the PPO algorithm is used to iteratively optimize the policy network parameters. Experimental results show that this method achieves superior fidelity performance compared to the classical PPO algorithm in the spin-squeezed GKP state preparation task, while significantly reducing the number of training parameters, thus verifying the effectiveness of the quantum-classical hybrid architecture in quantum control problems.
[0025] like Figure 1 and Figure 2 As shown, the method includes the following steps: S1. Construct a spin-squeezed GKP state preparation environment, establish a quantum system composed of N two-level quantum emitters, construct a quantum state evolution model based on the collective spin operator of the quantum system, and set the target spin-squeezed GKP state.
[0026] In practice, a spin-squeezed GKP state preparation environment is first constructed, and a quantum system consisting of N two-level quantum emitters is established, wherein each two-level quantum emitter contains the ground state |0. and excited state|1 Two energy levels, each with its own quantum emitter as the basic quantum unit, together form a many-body quantum system. To describe the overall dynamical properties of this quantum system, a spin operator is first defined for each individual two-level quantum emitter, corresponding to the spin components in the X, Y, and Z directions of the Bloch sphere, respectively, and the reduced Planck constant is used. It is represented using the natural unit system of =1.
[0027] Therefore, the following three monoatomic spin operators can be defined. These correspond to the X, Y, and Z directions of the Bloch sphere, respectively.
[0028] in To reduce Planck's constant, take = 1. For an atomic ensemble consisting of N spin atoms, its collective spin operator can be represented as: ,
[0029] in Let represent the k-th second-level atom. Using... shaft and Spin-squeezed GKP states were prepared by axial rotation and one-axis twisting (OAT).
[0030] Furthermore, a system Hamiltonian is established based on the constructed collective spin operator. This system Hamiltonian consists of a collective spin term, a phase spin term, and a uniaxial torsional compression term, and its expression is:
[0031] in, This is the Y-axis collective rotation control parameter, used to control the overall rotation intensity of the quantum state in the Y-axis direction of the Bloch sphere. By adjusting this parameter, the overall orientation of the quantum state can be rapidly adjusted. These are global phase rotation control parameters for the Z-axis, used to control the system to generate global phase rotation around the Z-axis, thereby achieving phase compensation and orbit correction between different quantum paths. This is a uniaxial torsional (OAT) strength parameter used to control the nonlinear compression process generated by the collective spin secondary interaction. The introduction of nonlinear dynamics allows the initial coherent spin state to gradually evolve into a spin-squeezed state with non-Gaussian properties, providing resources for subsequent GKP state preparation. All three control parameters are time-dependent, allowing dynamic adjustment of their values at different control moments, enabling the system to complete different stages of quantum state evolution according to the control strategy. Specifically, the collective rotation term primarily handles quantum state orientation adjustment, the phase rotation term handles global phase correction, and the uniaxial torsional squeezing term generates the spin-squeezing effect; all three work together to gradually construct the target quantum state.
[0032] Based on the established system Hamiltonian, the corresponding quantum state evolution operator is further generated. Within each control time step Δt, the system completes one composite evolution according to the set control parameters, and its quantum state evolution operator is represented as:
[0033] in: : Single-step unitary transformation operator, performing unitary transformation; : Around the Bloch sphere The collective rotation of the axes, with a rotation angle of . ; : around Global phase rotation of the axis, with a rotation angle of . It is typically used for phase compensation or orbit adjustment; : Uniaxial torsion, strength from Control is the core nonlinear term that generates spin compression and non-Gaussian characteristics.
[0034] S2. Construct a quantum-classical hybrid reinforcement learning model, which includes a policy network and a value network, wherein the policy network includes a parameterized quantum circuit.
[0035] After completing the construction of the quantum system and establishing the quantum state evolution model, a quantum-classical hybrid reinforcement learning model is further constructed. This model adopts an asymmetric Actor-Critic architecture, including a policy network (Actor) and a value network (Critic). The policy network uses a quantum-classical hybrid structure of "classical precoding layer + parameterized quantum circuit (PQC) + classical post-processing layer," utilizing the parameterized quantum circuit to extract features from high-dimensional quantum states and generate policies. The value network uses a pure classical multilayer perceptron (MLP) structure to perform state value estimation, ensuring the stability of the value function calculation and the continuity of gradient updates.
[0036] In one embodiment of the present invention, the network structure diagram of the quantum-classical hybrid reinforcement learning model is as follows: Figure 3 As shown.
[0037] Specifically, the current quantum state output by the quantum state evolution model in step S1 is used as the observation information of the reinforcement learning environment and converted into environmental state features. The environmental state features may include the spin distribution information corresponding to the current quantum state, the collective spin expectation value, the quantum state probability amplitude, or feature parameters that can characterize the evolution state of the system, so that the reinforcement learning model can generate corresponding control strategies based on the current quantum system state.
[0038] Furthermore, the environmental state features are pre-encoded. The pre-coding layer is constructed using a classic fully connected neural network. It extracts features and compresses the dimension of the input state through multiple linear mappings and nonlinear activation functions, mapping the original high-dimensional environmental state into quantum circuit input features suitable for parameterized quantum circuit processing. This pre-coding layer effectively reduces the quantum resource consumption caused by directly inputting high-dimensional quantum states into the quantum circuit, while retaining the main feature information of the environmental state, thus improving the computational efficiency of the subsequent quantum policy network.
[0039] After obtaining the input characteristics of the quantum circuit, they are input into a parameterized quantum circuit for quantum feature extraction. The initial input state of the parameterized quantum circuit. Composed of n qubits, The ground state of each qubit is represented by the initial state, which is initialized to form the initial quantum state. First, a Hadamard gate is applied to each qubit, which transforms all qubits from the ground state to a quantum superposition state, thereby constructing an initial quantum space with parallel expression capabilities, providing a foundation for subsequent high-dimensional feature mapping.
[0040] A parameterized quantum circuit is composed of multiple quantum layers stacked sequentially. Each quantum layer includes a variational block, an entangled block, and a coding block, as shown in the schematic diagram below. Figure 4 As shown.
[0041] The variational block executes the following steps sequentially for each qubit: Revolving doors and In the rotating door operation, the rotation angle corresponding to each rotating door is used as a trainable parameter to participate in model training. By continuously adjusting the rotation parameters, the parameterized quantum circuit can adaptively learn the nonlinear characteristics in the environmental state and improve the policy network's ability to express complex quantum control policies.
[0042] After completing a single-qubit rotation, the system enters an entangled block, where two-qubit CZ gate operations are performed between adjacent qubits to establish quantum entanglement among multiple qubits. Since quantum entanglement can establish correlations between different qubits, the exponential state space of a quantum system can be used to express the correlation characteristics between complex environmental states. This allows parameterized quantum circuits to effectively capture the coupling relationships between different state variables during quantum control, thereby improving the ability of policy networks to represent high-dimensional quantum systems.
[0043] Then the code block is entered, and execution is repeated sequentially on all qubits. Revolving doors and The rotating door operation, combined with the quantum circuit input features output from the classical precoding layer, encodes classical environmental state information into quantum state space, achieving a mapping between classical information and quantum information. Specifically, the encoded phase parameters are constrained by the tanh function, limiting the corresponding rotation angle to a specific range. Within a certain range, to ensure continuous parameter changes and stable training process.
[0044] After continuous processing through multiple quantum layers, the final quantum state is measured, and the expected value of each qubit on the Pauli Z operator is obtained as the quantum measurement result. The quantum measurement result carries the high-dimensional correlation characteristics of the environment state after quantum superposition, quantum entanglement, and parameterization transformation, and serves as the input to the post-processing layer of the policy network.
[0045] The post-processing layer employs a classic fully connected neural network to further perform feature mapping on the quantum measurement results, converting the quantum measurement output into control parameters that match the action space, and generating a corresponding control action probability distribution. This control action probability distribution represents the probability of each candidate control action being selected. The policy network performs action sampling based on this distribution to obtain the control actions required in step S3, thereby achieving dynamic adjustment of the quantum system's control parameters.
[0046] Simultaneously, environmental state characteristics are input into the value network for state value estimation. The value network employs a classic multilayer perceptron structure, with the current environmental state as input and the corresponding state value estimate as output. The formula is:
[0047] The environmental input state, i.e., the observed value; For state The corresponding predicted value, All trainable parameters of the value network; For value network computing; These are the weights and bias parameters of the MLP. The value network calculates the cumulative revenue prediction corresponding to the current environmental state through multi-layer linear mapping and non-linear activation functions. This prediction is used to evaluate the performance of the current policy and serves as a benchmark value in the subsequent policy optimization process for calculating the advantage function.
[0048] Because the policy network employs a quantum-classical hybrid structure while the value network uses a purely classical structure, an asymmetric quantum-classical hybrid reinforcement learning architecture is formed. The parameterized quantum circuit utilizes quantum superposition and quantum entanglement to achieve efficient representation of the high-dimensional state space, reducing the network parameter size while improving the learning ability of complex quantum control policies. The value network maintains the classical MLP structure, ensuring good numerical stability and gradient smoothness in state value estimation. This allows the entire reinforcement learning model to balance quantum expressive power with classical optimization stability, providing a reliable foundation for subsequent iterative updates of the control policy.
[0049] S3. Obtain the current quantum state output by the quantum state evolution model, input the current quantum state into the policy network, and generate control actions after processing by the parameterized quantum circuit.
[0050] In specific implementation, after the quantum-classical hybrid reinforcement learning model is constructed, the current quantum state output by the quantum state evolution model in step S1 is obtained, and the current quantum state is used as the input of the policy network. The policy network generates corresponding quantum control actions to drive the quantum system to complete the evolution at the next moment.
[0051] First, the action space for reinforcement learning control tasks is constructed. The action space is defined as a continuous multidimensional parameter space. It is represented as:
[0052] in, For unitary evolution operator The number of cascaded layers; For continuous interval constraints, the action space is defined as the policy network outputting three action components in each iteration. , , corresponding to a group Each action component It can be linearly mapped to rotation angle and compressive strength:
[0053] in, The action scaling factor is used to map the standardized actions output by the policy network to the physical control range allowed by the quantum control system, ensuring that the control parameters meet the actual driving requirements of the quantum system. Through action scaling, both numerical stability during reinforcement learning training and the control parameters generated satisfy the control constraints corresponding to the system's Hamiltonian are guaranteed.
[0054] Since quantum control processes typically occur in incompletely observable environments, and to avoid disrupting quantum coherence during GKP state preparation by mid-process measurement, it is impossible to obtain complete system state information in real time. Therefore, a reinforcement learning observation space is further constructed. This observation space is defined as a low-dimensional observation space. It is represented as:
[0055] In this embodiment, the environment returns fixed observation values. As input to the policy network, an open-loop control mode is realized. Throughout the entire GKP state preparation process, no intermediate quantum measurements are performed. Instead, policy optimization is based on the performance indicators corresponding to the final state, thereby avoiding quantum state collapse and coherence loss caused by continuous measurements and improving the accuracy of target quantum state preparation.
[0056] Within each control cycle, the current quantum state obtained in step S1 is converted into the corresponding environmental state features and input into the precoding layer of the policy network. The precoding layer first performs feature mapping and dimensionality compression on the environmental state to generate quantum circuit input features suitable for parameterized quantum circuit processing.
[0057] Subsequently, the input features of the quantum circuit are encoded into the qubits corresponding to the parameterized quantum circuit. Parameterized single-qubit rotation operations are then performed on each qubit sequentially, and a nonlinear transformation of the quantum state is completed through a trainable rotation angle. After the single-qubit rotation is completed, entanglement operations are performed between adjacent qubits to establish quantum correlations among multiple qubits, thereby utilizing the properties of quantum superposition and quantum entanglement to achieve high-dimensional feature representation of the environmental state. Encoding operations are then performed to further map the classical environmental features to the quantum state space, completing the fusion representation of classical and quantum information.
[0058] After completing the calculations for all quantum layers of the parameterized quantum circuit, the final quantum state is measured to obtain the expected value of the Pauli Z operator corresponding to each qubit. The quantum measurement results are then input into the post-processing layer of the policy network. The post-processing layer generates a control action probability distribution based on the quantum measurement results, samples the data according to the control action probability distribution, and outputs the corresponding control action.
[0059] S4. Map the control action to the control parameters of the quantum state evolution model, update the quantum state evolution model using the control parameters and drive the quantum system to evolve to obtain the quantum state at the next moment.
[0060] In this embodiment, the initial state of the quantum system is defined as a coherent spin state, and a multi-layer cascaded evolution method is used to complete the state update. The entire control sequence includes M layers of quantum evolution structure, each layer containing three basic operations in sequence: Y-axis rotation gate, Z-axis rotation gate, and single-axis torsion gate. The control parameters of each layer are output and dynamically adjusted in real time by the policy network.
[0061] In a single-step state update process, the current quantum state is processed by the quantum state evolution operator corresponding to the current control step to obtain the quantum state at the next time step. Its evolution process can be represented as follows:
[0062] in, for The quantum state at a given moment; The quantum state at the next time step after the first-order unitary transformation is completed; For the first floor to the second floor The layer-unitary evolution operator acts on the current quantum state in a preset order, realizing multi-level continuous quantum control.
[0063] Since each stage of evolution involves three physical processes—collective rotation, phase correction, and uniaxial torsional compression—the multi-layered cascaded structure can continuously adjust the quantum state orientation, correct the evolution phase, and enhance the spin compression effect, enabling the quantum system to gradually evolve towards the target spin-compressed GKP state.
[0064] After completing all M-level quantum evolutions in the current control cycle, the updated quantum state for the next time step is obtained. This next time step quantum state is then used as the current quantum state input to the policy network for the next reinforcement learning decision cycle, and also serves as the environmental state basis for reward calculation and value assessment, thus realizing a closed-loop iterative process between control action generation, quantum state evolution, and policy optimization.
[0065] S5. Calculate the reward value based on the fidelity between the next quantum state and the target spin-squeezed GKP state, and construct an empirical trajectory by combining the current quantum state, control action, reward value and the next quantum state.
[0066] In practice, after the quantum state evolution of the current control cycle is completed, the reward value is calculated based on the similarity between the quantum state at the next moment and the target spin-squeezed GKP state. The state transition information in the current control process is used to construct the reinforcement learning experience trajectory, providing training data for the policy update of the subsequent proximal policy optimization algorithm.
[0067] First, the next-time quantum state obtained in step S4 is taken as the final output quantum state, and the state overlap between the next-time quantum state and the pre-set target spin-squeezed GKP state is calculated. The target spin-squeezed GKP state is the target quantum state pre-set during the system optimization process, which corresponds to the ideal state that the quantum control task ultimately hopes to achieve; the next-time quantum state is the final output quantum state after all the control sequences of the M layers have been executed.
[0068] Furthermore, the state overlap relationship between the target spin-squeezed GKP state and the final output quantum state is calculated based on their inner product, and the corresponding fidelity is calculated using the state overlap. The fidelity is defined as:
[0069] in, The final output quantum state after all layer control operations have been completed.
[0070] Since the fidelity value ranges from 0 to 1, it is 1 when the final output quantum state is completely identical to the target spin-squeezed GKP state; as the difference between the two increases, the fidelity gradually decreases. Therefore, this embodiment directly uses the fidelity value as the reward value corresponding to the reinforcement learning reward function, so that the reward value can truly reflect the quality of the current control strategy's effect on the preparation of the target quantum state.
[0071] Using fidelity as the reward function has clear physical significance, as it can directly measure the optimization objective of quantum control tasks and avoid the parameter tuning problems caused by artificially designing complex reward functions. Furthermore, since the reward is calculated uniformly only after the entire control sequence is completed, the entire quantum control process employs a terminal reward mechanism, eliminating the need for intermediate measurements during the control process. This effectively avoids quantum state collapse and coherence destruction caused by quantum measurements, ensuring that the target spin-squeezed GKP state maintains a complete coherent evolution process.
[0072] Furthermore, after calculating the reward value, the data corresponding to the current control process is constructed into a reinforcement learning experience trajectory, and the constructed experience trajectory is stored in the experience cache. Complete training batches are then formed sequentially according to the control sequence. After collecting data for a preset number of control rounds, the experience trajectory is uniformly input into the proximal policy optimization algorithm in step S6. State transition information, reward information, and action probability information are used to update the parameters of the policy network and value network, thereby continuously improving the reinforcement learning model's control capability over the target spin-compressed GKP state, enabling the quantum system to gradually learn a quantum control policy with higher fidelity.
[0073] S6. Based on the empirical trajectory, the proximal policy optimization algorithm is used to update the parameters of the policy network and the value network, and steps S3 to S6 are repeated until the preset training conditions are met to obtain the control policy for preparing the target spin-compressed GKP state.
[0074] In this embodiment of the invention, the algorithm flow during model training can be divided into an intelligent agent-environment interaction process and a parameter update process. During the interaction phase, after initializing network parameters and the environment, the policy network receives the state of the environment. Output the probability distribution of actions, and sample actions based on this distribution. According to the environment Evolving to the next state And give a reward for this action. The intelligent agent records the interaction trajectory. , The agent then proceeds to the experience replay area. Once a sufficient number of trajectories have accumulated in the experience replay area, the agent pauses interaction and enters the parameter update phase.
[0075] During parameter updates, the value network processes the states in the trajectories in batches. Predict its value Actions are calculated using the Generalized Advantage Estimation (GAE) method. Advantage function:
[0076] in Indicates timing difference error. This represents the maximum time step in a single round. It is a discount factor. These are the smoothing parameters for GAE. Advantage function. Measuring actions Compared to the strengths and weaknesses of average policies, it guides the optimization of policy networks. Through GAE, PPO balances the variance and bias of advantage estimation, improving training stability.
[0077] To prevent large policy updates from causing training instability, the PPO algorithm introduces a pruning mechanism. Specifically, in each iteration, PPO uses the current policy. Collect a batch of trajectory data and calculate the ratio of the old and new strategies.
[0078] in These are the trainable parameters for the current policy network; It's the old strategy from the previous iteration. (Ratio) The objective function of PPO is defined as: [Measurement of the differences in action selection between the old and new strategies.]
[0079] The clipping function restricts the ratio range to [ ]; To prune hyperparameters, use standard values. To constrain the policy update magnitude and prevent training oscillations, the total loss of PPO combines the objective function loss, value function loss, and policy entropy regularization term.
[0080] Value loss is defined as:
[0081] in For the goal of return, It is a strategy exist Entropy in a state, and These are weighting coefficients used to balance the contributions of each loss term. Indicates all Calculate the expected value. The policy entropy regularization term encourages the agent to remain exploratory and avoids premature convergence to a suboptimal policy. Through mini-batch gradient updates, PPO achieves efficient sample utilization and stable convergence.
[0082] After a parameter update is completed, steps S3 to S6 are repeated, that is, the current quantum state is reacquired, control actions are generated using the updated policy network to drive the evolution of the quantum system, the reward value based on fidelity is calculated and a new empirical trajectory is constructed, and then the proximal policy optimization algorithm is used again to update the policy network and value network parameters.
[0083] Training stops when preset training conditions are met. These preset training conditions may include: the number of training rounds reaching a preset maximum value; the fidelity of the target spin-compressed GKP state reaching a preset threshold; the reward value change being less than a preset convergence threshold over several consecutive training rounds; or the update magnitude of the policy network parameters being lower than a preset threshold. After training, a control strategy for preparing the target spin-compressed GKP state is obtained. This control strategy outputs the corresponding Y-axis rotation angle, Z-axis rotation angle, and uniaxial torsional intensity based on the current state of the quantum system, enabling the quantum system to evolve to the target spin-compressed GKP state under the action of a multi-level control sequence.
[0084] In one embodiment of the present invention, a comparative experiment was conducted on the classical PPO and the quantum-classical hybrid algorithm (PPO-Q) proposed in this paper. The experimental results are as follows: Figure 5 Using the classical PPO algorithm to prepare spin-squeezed GKP states, a fidelity of 97% can be achieved; using a quantum-classical hybrid algorithm, a fidelity of 98% can be achieved. This demonstrates that the robust control method for preparing spin-squeezed GKP states based on quantum reinforcement learning proposed in this application can achieve good performance while significantly reducing the number of parameters.
[0085] To implement the methods of the above embodiments, the present invention also provides an electronic device, which includes a memory and a processor; wherein the processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement the various steps of the methods described above.
[0086] To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the foregoing embodiments.
[0087] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0088] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0089] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
Claims
1. A method for preparing spin-squeezed GKP states based on quantum reinforcement learning, characterized in that, Includes the following steps: S1. Construct a spin-squeezed GKP state preparation environment, establish a quantum system composed of N two-level quantum emitters, construct a quantum state evolution model based on the collective spin operator of the quantum system, and set the target spin-squeezed GKP state; S2. Construct a quantum-classical hybrid reinforcement learning model, which includes a policy network and a value network, and the policy network includes a parameterized quantum circuit. S3. Obtain the current quantum state output by the quantum state evolution model, input the current quantum state into the policy network, and generate control actions after processing by the parameterized quantum circuit; S4. Map the control action to the control parameters of the quantum state evolution model, update the quantum state evolution model using the control parameters and drive the quantum system to evolve, and obtain the quantum state at the next moment; S5. Calculate the reward value based on the fidelity between the next quantum state and the target spin-squeezed GKP state, and construct an empirical trajectory by combining the current quantum state, control action, reward value, and next quantum state. S6. Based on the empirical trajectory, the proximal policy optimization algorithm is used to update the parameters of the policy network and the value network, and steps S3 to S6 are repeated until the preset training conditions are met to obtain the control policy for preparing the target spin-compressed GKP state.
2. The GKP state preparation method based on quantum reinforcement learning according to claim 1, characterized in that, Step S1 includes: Establish a quantum system consisting of N two-level quantum emitters; Construct a collective spin operator based on the Pauli operators of each two-level quantum emitter in the quantum system; The collective spin operator is used to construct a system Hamiltonian that includes a collective spin term, a phase spin term, and a uniaxial torsional compression term.
3. The GKP state preparation method based on quantum reinforcement learning according to claim 2, characterized in that, The system Hamiltonian, comprising a collective spin term, a phase rotation term, and a uniaxial torsional compression term, is constructed using the collective spin operator, including: Construct a collective rotation term using the Y-axis component of the collective spin operator; A phase rotation term is constructed using the Z-axis component of the collective spin operator; A uniaxial torsional compression term is constructed using the axial quadratic term in the collective spin operator; The collective rotation term, phase rotation term, and uniaxial torsional compression term are combined to form a system Hamiltonian, which is then used to generate quantum state evolution operators.
4. The GKP state preparation method based on quantum reinforcement learning according to claim 1, characterized in that, Step S2 includes: The current quantum state output by the quantum state evolution model is converted into environmental state characteristics; The environmental state characteristics are pre-encoded to obtain the quantum circuit input characteristics; The quantum circuit input characteristics are input into the parameterized quantum circuit to obtain quantum measurement results.
5. The GKP state preparation method based on quantum reinforcement learning according to claim 4, characterized in that, The quantum circuit input characteristics are input into a parameterized quantum circuit to obtain quantum measurement results, including: The input features of the quantum circuit are encoded into qubits; Perform parameterized single-qubit rotation operations on the qubit; Entanglement operations are performed between qubits to obtain quantum states that carry environmental state correlation characteristics; The quantum state carrying the environmental state correlation characteristics is measured to obtain the quantum measurement result.
6. The GKP state preparation method based on quantum reinforcement learning according to claim 1, characterized in that, Step S4 includes: The control action is decomposed into multiple action components; The multiple motion components are respectively mapped to Y-axis rotation angle, Z-axis rotation angle and single-axis torsional strength; Substitute the Y-axis rotation angle, Z-axis rotation angle, and uniaxial torsional intensity into the quantum state evolution model to generate the quantum state evolution operator for the current control step; The quantum state evolution operator of the current control step is applied to the current quantum state to obtain the quantum state at the next moment.
7. The GKP state preparation method based on quantum reinforcement learning according to claim 1, characterized in that, Step S5 includes: Calculate the state overlap between the next quantum state and the target spin-squeezed GKP state; The fidelity is calculated based on the state overlap, and a reward value is generated based on the fidelity. The current quantum state, control action, reward value, next quantum state, and action probability are stored as an experience trajectory, which is then used to update the near-end policy optimization algorithm in step S6.
8. The GKP state preparation method based on quantum reinforcement learning according to claim 7, characterized in that, Step S6 includes: Read the state sequence, action sequence, reward sequence, and action probability from the experience trajectory; Calculate the state value corresponding to the state sequence using a value network; Calculate the temporal difference error based on the reward sequence and state value; The dominance function is calculated based on the time-series difference error. The pruning objective function of the near-end policy optimization algorithm is constructed based on the action probability and the advantage function. The policy network parameters are updated using the pruning objective function, and the value network parameters are updated using the value loss function until the preset training conditions are met.
9. An electronic device, characterized in that, Including processor and memory; The processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement the method as described in any one of claims 1-8.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-8.