Quantum calculation simulation acceleration method for multi-quantum-bit Foeli revolving door
By decomposing a multi-qubit Pauli rotation gate into a Pauli matrix tensor product generating matrix consisting of an eigenvector matrix and its conjugate transpose, and directly manipulating the quantum state vector, the problems of high resource consumption, low efficiency, and large error in existing technologies are solved, thus achieving efficient and accurate quantum simulation.
Patent Information
- Application Number
- CN202511690754.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-03-03
AI Technical Summary
Existing technologies for simulating multi-qubit Pauli rotation gates suffer from high resource overhead, low computational efficiency, large memory requirements, difficulty in hardware parallelization, and the introduction of approximation errors, which limit the efficiency and feasibility of large-scale quantum simulations.
The Pauli rotation gate for multi-qubit quantum bits is decomposed into a Pauli matrix tensor product generating matrix consisting of an eigenvector matrix and its conjugate transpose. By directly manipulating the quantum state vector through unitary matrix transformation, the CNOT gate and SWAP operation in traditional schemes are avoided, thus achieving high-precision rotation operations.
Significantly reduces resource consumption, improves simulation efficiency, eliminates performance overhead due to hardware topology limitations, and enables high-precision quantum simulation, suitable for high-fidelity scenarios such as quantum chemistry and materials science.
Smart Images

Figure CN121599150A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of quantum circuit simulation and high-performance computing technology, and specifically relates to a computational method for efficiently accelerating the interaction of multi-qubit Pauli rotation gates with quantum states during quantum simulation on classical computing hardware (such as FPGA and ASIC). Background Technology
[0002] Running quantum algorithms directly on current NISQ hardware still faces significant challenges. Therefore, utilizing classical computing devices for high-performance simulation is a key approach to researching and developing quantum algorithms. Quantum simulation plays an indispensable role not only in the design, verification, and debugging of quantum algorithms, but is also a crucial tool for studying complex phenomena in fields such as quantum many-body physics and quantum chemistry (see Affronte M. Molecular nanomagnets and related phenomena[M]. Berlin: Springer, 2015.).
[0003] However, the dimension of a quantum state increases exponentially with the number of qubits. A quantum state of a system consisting of n qubits requires a length of... The state vector is used to describe it (see Wang Y. Quantum computation and quantum information[J]. 2012.). Simulating a system with only 50 qubits on a classical computer requires storage A complex number. Calculated in double-precision (16-byte) floating-point terms, storing the state vector alone would require approximately 18 petabytes of memory, far exceeding the available resources of any single classical computer. This exponential demand for storage and computing resources poses a tremendous challenge to classical computing hardware.
[0004] Multi-qubit Pauli rotation gates are core components in many important quantum algorithms (such as the quantum approximate optimization algorithm QAOA) (see Farhi E, Goldstone J, Gutmann S. A quantum approximate optimization algorithm[J]. arXiv preprint arXiv:1411.4028, 2014.). Currently, most mainstream simulations of multi-qubit Pauli rotation gates are based on gate decomposition (see Sriluckshmy PV, Pina-Canelles V, Ponce M, et al. Optimal, hardware native decomposition of parameterized multi-qubit Pauli gates[J]. Quantum Science and Technology, 2023, 8(4):045029.), without direct computation, and suffer from problems such as low computational efficiency, large memory overhead, and difficulty in hardware parallelization.
[0005] Currently, the simulation scale of multi-qubit Pauli rotation gates implemented in existing hardware such as FPGAs is mostly no more than 2 qubits (see Choi S, Lee K, Lee JJ, et al. Standalone FPGA-Based QAOAEmulator for Weighted-MaxCut on Embedded Devices[J]. arXiv preprint arXiv:2502.11316, 2025.). In quantum simulation software, single-qubit Pauli rotation gates are easy to implement, while for multi-qubit Pauli rotation gate simulations, the memory overhead of direct tensor product expansion is... Therefore, existing quantum compilers (such as Qiskit) propose decomposing it into a combination of multiple single-bit quantum gates and CNOT gates (see Li G, Wu A, Shi Y, et al. Paulihedral: a generalized block-wise compiler optimization framework for quantum simulation kernels[C] / / Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems. 2022: 554-569.). However, this leads to increased circuit depth, a significant increase in the number of gates, and introduces additional mapping overhead.
[0006] The aforementioned existing gate decomposition-based technical solutions have the following significant drawbacks, which severely restrict the efficiency and feasibility of large-scale quantum simulations, and are reflected in the following aspects:
[0007] (1) Extremely high resource overhead: This scheme requires decomposing high-level gates into a large number of native gates (such as CNOT, single-bit rotating gate), which leads to a significant increase in circuit depth and width, consuming a large amount of valuable resources.
[0008] (2) Huge swapping overhead: In practical near-neighbor connection architecture hardware, a large number of additional SWAP operations must be introduced to transfer quantum states in order to perform the remote two-bit gates required for decomposition, which introduces huge performance overhead.
[0009] (3) Introducing approximation errors makes it difficult to meet the requirements of high-precision simulation: To achieve operability, most gate decomposition schemes cannot accurately realize multi-bit Pauli gates with arbitrary rotation angles. They usually approximate the target gate into a set of discrete, finite basic gate sequences. This approximation introduces inherent, non-zero systematic errors (i.e., approximation errors). To achieve higher accuracy, longer gate sequences must be used for more refined approximations, leading to a contradiction between accuracy and resource consumption.
[0010] (4) Limited compilation and optimization space: By solidifying multi-bit Pauli rotation gates into a fixed sequence of basic gates, the quantum compiler has a very limited space for subsequent circuit optimization. The compiler can only make local, syntactic-level fine-tunings, and it is difficult to understand the overall intent of the sequence from a global perspective.
[0011] Based on the above analysis, existing research is limited by quantum computing hardware and has to approximate the decomposition of multi-qubit Pauli rotation gates into multiple Pauli gates, which introduces a lot of additional computation and memory overhead and needs to be improved. This is the reason for this case. Summary of the Invention
[0012] The purpose of this invention is to provide a quantum computing simulation acceleration method for multi-qubit Pauli rotation gates, which has low hardware resource overhead, high computational efficiency, and strong versatility. It is specifically designed to accelerate multi-qubit Pauli rotation gate calculations in quantum simulations, thereby improving the overall performance of quantum simulators. At the same time, it can provide an efficient quantum circuit simulation scheme for future quantum chips that support multi-qubit Pauli rotation gates.
[0013] To achieve the above objectives, the solution of the present invention is:
[0014] A method for accelerating quantum computing simulations using multi-qubit Pauli rotation gates includes the following steps:
[0015] Step 1, using a multi-qubit Pauli rotation gate Decomposed into ,in, The eigenvector matrix, for The conjugate transpose of the matrix. It is a diagonal matrix;
[0016] The multi-qubit Pauli revolving door Represented as ,in, The eigenvector matrix is the generator matrix of the Pauli tensor product. satisfy ;
[0017] Step 2, Normalization yields its normalized matrix. Then for Simplifying, we get
[0018] ,
[0019] in, N represents a non-negative integer; , for The conjugate transpose of ; The modulus of all diagonal elements is 1;
[0020] Step 3, divide the matrix into... , , Multiply by the quantum state vector, and adjust the elements of the state vector according to the elements in each matrix;
[0021] Step 4: Shift the adjusted state vector to the right as a whole. Bit.
[0022] In step 1 above, ,
[0023] in ;
[0024] ,
[0025] in ;
[0026] .
[0027] In step 1 above, the matrix It is generated by the tensor product of Pauli X, Y, and Z matrices, where the eigenvalues of the Pauli X, Y, and Z matrices are decomposed into...
[0028] ,
[0029] ,
[0030] .
[0031] In step 3 above, the matrix Multiplying by the quantum state vector initializes the resulting vector to zero, according to the matrix. The transformation of the elements of the state vector is determined by the inner element, and the specific rules are as follows.
[0032] If the matrix element is 1, add the vector element directly to the result vector;
[0033] If the matrix element is -1, invert the vector elements first and then add them to the result vector;
[0034] If the matrix element is i, swap the real and imaginary parts of the vector elements, take the negative sign of the new real part, and then add it to the result vector;
[0035] If the matrix element is -i, swap the real and imaginary parts of the vector elements, take the negative sign of the new imaginary part, and then add it to the result vector;
[0036] If the matrix element is 0, the resulting vector remains unchanged.
[0037] In step 3 above, the matrix Multiply by the quantum state vector to perform a rotation operation on the state vector.
[0038] In step 3 above, the matrix Multiplying by the quantum state vector initializes the resulting vector to zero, according to the matrix. The transformation of the elements of the state vector is determined by the inner element, and the specific rules are as follows.
[0039] If the matrix element is 1, add the vector element directly to the result vector;
[0040] If the matrix element is -1, invert the vector elements first and then add them to the result vector;
[0041] If the matrix element is i, swap the real and imaginary parts of the vector elements, take the negative sign of the new real part, and then add it to the result vector;
[0042] If the matrix element is -i, swap the real and imaginary parts of the vector elements, take the negative sign of the new imaginary part, and then add it to the result vector;
[0043] If the matrix element is 0, the resulting vector remains unchanged.
[0044] After adopting the above solution, the beneficial effects of the present invention compared with the prior art are as follows:
[0045] (1) Significantly reduce resource consumption: By decomposing the unitary matrix transformation, this invention completely avoids the serial decomposition and execution of a large number of CNOT gates and single-bit gates in the traditional scheme, resulting in a reduction in circuit depth and running time by orders of magnitude, which greatly improves the simulation efficiency.
[0046] (2) Elimination of swap overhead: The core transformation steps of this invention are all operated directly at the state vector level without the need to introduce any SWAP operation, which fundamentally solves the performance overhead problem caused by hardware topology limitations and demonstrates strong versatility and robustness to various hardware architectures.
[0047] (3) It can achieve high-precision execution and meet the requirements of high-fidelity simulation: The present invention can accurately, or with a precision much higher than that of the gate decomposition scheme, achieve arbitrary rotation angles, fundamentally avoiding the introduction of decomposition approximation errors. This makes the present invention particularly suitable for high-fidelity simulation scenarios such as quantum chemistry and materials science. Attached Figure Description
[0048] Figure 1 This is a flowchart of the present invention. Detailed Implementation
[0049] The technical solution and beneficial effects of the present invention will be described in detail below with reference to the accompanying drawings.
[0050] To achieve the above objectives, this invention provides a method for accelerating quantum computing simulations using a multi-qubit Pauli rotation gate:
[0051] For the Pauli X matrix, Pauli Y matrix, and Pauli Z matrix, their eigenvalue decomposition is as follows: , , . The matrix It is generated by the tensor product of Pauli's X, Y, and Z matrices.
[0052] set up Let A be an eigenvector of matrix A, with corresponding eigenvalues. , Let B be an eigenvector of matrix B, with corresponding eigenvalues. ,but yes The eigenvectors and corresponding eigenvalues Let Pauli matrix ∈{X,Y,Z}, and each It can be decomposed by eigenvalues, denoted as ,in The eigenvector matrix, For the eigenvalue matrix, for The conjugate transpose of .
[0053] Therefore, the Pauli matrix The matrix generated by the tensor product The eigenvalue decomposition can be expressed as:
[0054] (1)
[0055] eigenvector matrix of the Pauli matrix It can be expressed as equation (2), where .
[0056] (2)
[0057] And because of sets Since the multiplication operation is closed, we have equation (3), where , Representation matrix The normalized matrix.
[0058] (3)
[0059] Since the set {1, -1} is also closed under multiplication, we can obtain equation (4). ,in .
[0060] (4)
[0061] Let matrix The feature vector is The corresponding feature value is ,satisfy Due to the matrix It can be broken down into Multiply both sides of the equation by a matrix. eigenvectors Equation (5) is obtained.
[0062] (5)
[0063] because and will Substituting into equation (5), we simplify it to equation (5). (6):
[0064] (6)
[0065] In the formula , Also a matrix A set of feature vectors, and the corresponding feature values are . Therefore, the Pauli rotation gate with multiple qubits... It can be decomposed into equation (7).
[0066] (7)
[0067] The method includes the following steps:
[0068] S1: Pauli Rotation Gate Decomposition of Multi-Qubit Quantum Bits
[0069] S1.1: Decomposition Rule Definition
[0070] The specific rules are as follows:
[0071] Pauli Revolving Door with Multiple Quantum Bits Decomposed into , where the matrix and diagonal It can be derived from formulas (3)(4)(7).
[0072] S2: Calculation of pipeline structure
[0073] From the formula (3), It can be normalized to Therefore It can be simplified to:
[0074] (8)
[0075] in , The modulus of all diagonal elements is 1. According to equation (8), after... The resulting quantum state is determined through the following steps:
[0076] The first step is to calculate the matrix. Multiply by the quantum state vector. Initialize the result vector to zero, according to the matrix. The transformation of the state vector elements is determined by the inner element. The specific rules are as follows:
[0077] If the matrix element is 1, then the vector element is directly added to the result vector;
[0078] If the matrix element is -1, then the vector elements are inverted and added to the result vector.
[0079] If the matrix element is i, then swap the real and imaginary parts of the vector elements, take the negative sign of the new real part, and add it to the result vector.
[0080] If the matrix element is -i, then the real and imaginary parts of the vector elements are swapped, the new imaginary part is negative, and then added to the resulting vector.
[0081] If the matrix element is 0, the resulting vector remains unchanged.
[0082] The second step is to calculate the matrix. Multiply by the quantum state vector. Because The diagonal elements have a modulus of 1, so the state vector is rotated directly.
[0083] The third step is to calculate the matrix. Multiply by the quantum state vector. The steps are the same as in the first step.
[0084] Fourth step, shift the entire state vector to the right. Bit.
[0085] The second and third steps can be designed as pipelined structures for parallel computation, thereby significantly reducing the overall computational latency and improving the throughput of large-scale quantum simulations.
[0086] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0087] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0088] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0089] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0090] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0091] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for accelerating quantum computing simulation using a multi-qubit Pauli rotation gate, characterized in that... Includes the following steps: Step 1, using a multi-qubit Pauli rotation gate Decomposed into ,in, The eigenvector matrix, for The conjugate transpose of the matrix. It is a diagonal matrix; The multi-qubit Pauli revolving door Represented as ,in, The eigenvector matrix is the generator matrix of the Pauli tensor product. satisfy ; Step 2, Normalization yields its normalized matrix. Then for Simplifying, we get , in, N represents a non-negative integer; , for The conjugate transpose of ; The modulus of all diagonal elements is 1; Step 3, divide the matrix into... , , Multiply by the quantum state vector, and adjust the elements of the state vector according to the elements in each matrix; Step 4: Shift the adjusted state vector to the right as a whole. Bit.
2. The method as described in claim 1, characterized in that: In step 1, , in ; , in ; 。 3. The method as described in claim 1, characterized in that: In step 1, the matrix It is generated by the tensor product of Pauli X, Y, and Z matrices, where the eigenvalues of the Pauli X, Y, and Z matrices are decomposed into... , , 。 4. The method as described in claim 1, characterized in that: In step 3, the matrix Multiplying by the quantum state vector initializes the resulting vector to zero, according to the matrix. The transformation of the elements of the state vector is determined by the inner element, and the specific rules are as follows. If the matrix element is 1, add the vector element directly to the result vector; If the matrix element is -1, invert the vector elements first and then add them to the result vector; If the matrix element is i, swap the real and imaginary parts of the vector elements, take the negative sign of the new real part, and then add it to the result vector; If the matrix element is -i, swap the real and imaginary parts of the vector elements, take the negative sign of the new imaginary part, and then add it to the result vector; If the matrix element is 0, the resulting vector remains unchanged.
5. The method as described in claim 1, characterized in that: In step 3, the matrix Multiply by the quantum state vector to perform a rotation operation on the state vector.
6. The method as described in claim 1, characterized in that: In step 3, the matrix Multiplying by the quantum state vector initializes the resulting vector to zero, according to the matrix. The transformation of the elements of the state vector is determined by the inner element, and the specific rules are as follows. If the matrix element is 1, add the vector element directly to the result vector; If the matrix element is -1, invert the vector elements first and then add them to the result vector; If the matrix element is i, swap the real and imaginary parts of the vector elements, take the negative sign of the new real part, and then add it to the result vector; If the matrix element is -i, swap the real and imaginary parts of the vector elements, take the negative sign of the new imaginary part, and then add it to the result vector; If the matrix element is 0, the resulting vector remains unchanged.