GPU-oriented quantum circuit cutting simulation optimization method
Through the GPU-oriented quantum circuit cutting simulation optimization method, customized cutting strategies and GPU acceleration technology are used to solve the problems of high memory footprint and low computing efficiency in large-scale quantum circuit simulation, and efficient quantum circuit simulation and optimization are achieved.
Patent Information
- Application Number
- CN202510377621.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-01
AI Technical Summary
The memory usage and low computing efficiency in large-scale quantum circuit simulations lead to exhaustion of computing resources, making it difficult to effectively verify quantum algorithms and debug quantum hardware.
The quantum circuit cutting simulation optimization method is adopted for GPU, and the simulation process of quantum circuits is optimized through customized cutting strategies and efficient GPU acceleration technology, including circuit preprocessing, heuristic cutting, hybrid integer programming model, probability bucket allocation and state merging.
It significantly reduces the memory footprint and computing complexity of large-scale quantum circuit simulation, improves GPU resource utilization and computing efficiency, and provides powerful tools to support the verification of quantum algorithms and debugging of quantum hardware.
Smart Images

Figure CN120235264A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of quantum computing, and particularly relates to a method for optimizing the simulation of quantum circuit cutting for GPUs, aiming to solve the problems of excessive memory occupation and low computing efficiency in large-scale quantum circuit simulation through customized cutting strategies and efficient GPU acceleration technologies. Background Art
[0002] As a revolutionary computing paradigm, quantum computing utilizes characteristics such as quantum superposition and entanglement, and is expected to achieve computing capabilities far beyond classical computers for specific problems. For example, quantum algorithms such as Shor's algorithm and Grover's algorithm have shown exponential and quadratic acceleration potentials in prime factorization and database search respectively. However, the scarcity of quantum hardware, noise problems, and the stability limitations of qubits make the verification of quantum algorithms and the debugging of quantum hardware highly dependent on quantum simulators on classical computers. Traditional quantum simulation methods simulate quantum circuits through full-amplitude state vectors, but as the number of qubits increases, the required memory grows exponentially, leading to a rapid depletion of computing resources and becoming the main bottleneck restricting quantum simulation performance. For example, simulating a circuit with 50 qubits requires more than 1 PB of memory, which far exceeds the capabilities of current classical computers.
[0003] To solve this problem, the quantum circuit cutting method has emerged. This method decomposes large-scale quantum circuits into multiple small-scale sub-circuits, simulates them separately, and then reconstructs the complete result through classical post-processing, thus significantly reducing memory requirements and computational complexity. The core idea of the quantum cutting method is "divide and conquer", by reducing the number of qubits in a single sub-circuit, transforming the exponential memory requirement into multiple manageable sub-problems. However, simply relying on classical computers to simulate the cut circuits still faces the problem of low computing efficiency, especially when dealing with complex quantum algorithms, the limitations of classical computing resources are still significant.
[0004] In recent years, the combination of classical simulators and quantum cutting methods has become a research hotspot. By leveraging the high-performance computing capabilities of classical computers (such as the parallel computing advantages of GPUs), the simulation process of sub-circuits is further accelerated. The high computing power of GPUs provides new optimization space for the quantum cutting method, and its large-scale parallel computing ability can significantly improve the simulation efficiency of sub-circuits.
[0005] In previously existing patented inventions, Yu Lei proposed a full-amplitude quantum computing simulation method (CN111832144B), which accelerates the simulation efficiency in a parallel manner; Dou Menghan proposed a new quantum state encoding route, reducing the number of logic gates required for computing tasks (CN117787428A). However, these patents do not involve the splitting and reuse of circuits and still treat them as a whole for computing, resulting in a relatively rapid increase in complexity when the number of qubits increases. Nevertheless, how to efficiently utilize the video memory resources of GPUs and reduce data transmission overhead remains a key challenge in achieving efficient quantum simulation. For example, the video memory capacity of commercial GPUs is limited, and data transmission between GPUs and CPUs may become a performance bottleneck.
[0006] Glossary:
[0007] 1. Quantum Computing: A new type of computing method based on the principles of quantum mechanics that uses qubits for computing. Different from classical computing, quantum computing utilizes the properties of quantum superposition and quantum entanglement to be able to process a large number of computing tasks in parallel, significantly enhancing the computing efficiency for certain problems.
[0008] 2. Qubit: The basic unit of quantum computing, which is different from bits in traditional computing and can simultaneously be in a superposition state of 0 and 1. A qubit can be represented as a linear combination of the two basis states |0> and |1> to describe quantum information.
[0009] 3. Quantum Circuit: A circuit composed of quantum gate operations used to implement the process of quantum computing. The quantum gates in a quantum circuit operate on qubits to achieve the evolution of quantum states and the execution of computing tasks.
[0010] 4. Circuit Cutting: A technique for decomposing a complex quantum circuit into several sub-circuits. Through cutting, the original circuit can be divided into smaller and more manageable sub-circuits, thereby reducing computing complexity and resource consumption.
[0011] 5. State Vector Reuse: A method for reducing the amount of computation by repeatedly using the intermediate state vectors of similar sub-circuits. It can effectively reduce the need for repeated computation and improve the efficiency of quantum circuit simulation.
[0012] 6. Mixed-Integer Programming: An optimization technique used to solve optimization problems involving integer and continuous variables. In quantum circuit simulation, the mixed-integer programming model is used to optimize the circuit cutting process to determine the optimal cutting points and schemes. Summary of the Invention
[0013] The object of the present invention is to propose a simulation optimization method for quantum circuit cutting for GPUs. Through customized cutting strategies and efficient GPU acceleration technologies, the present invention significantly optimizes the simulation efficiency of large-scale quantum circuits.
[0014] A simulation optimization method for quantum circuit cutting for GPUs proposed by the present invention is specifically as follows:
[0015] (1) Circuit preprocessing and establishing dependencies between qubits
[0016] (1.1) Preprocess the target quantum circuit, clean single-qubit gates and other redundant parts in the quantum circuit to simplify the target quantum circuit;
[0017] (1.2) Analyze the dependencies between qubits in the quantum circuit preprocessed in step (1.1), construct the quantum circuit as a directed acyclic graph (DAG), where nodes represent quantum gates and edges represent the data flow between qubits; subsequently, use the mixed integer programming (MIP) method to model the directed acyclic graph, transform the optimization problem of the quantum circuit into an objective function and constraint conditions, so as to achieve efficient allocation and optimization of quantum circuit resources;
[0018] (2) Circuit cutting
[0019] Adopt a heuristic cutting method, determine the best cutting point according to the complexity of the quantum circuit and the dependencies between qubits in step (1.2); the heuristic cutting method selects a scheme that can minimize the overall simulation burden by evaluating the computational overhead and resource requirements of different cutting schemes;
[0020] (3) Evaluation preprocessing
[0021] (3.1) Determine the evaluation order: Dynamically determine the evaluation order of sub-circuits according to the dimensions after contraction of each sub-circuit and its adjacent sub-circuits; by analyzing the dependencies between sub-circuits and the scale of the state vector, adopt the minimum dimension first method to give priority to evaluating sub-circuits with smaller dimensions to reduce computational complexity and memory occupancy;
[0022] (3.2) Allocate probability buckets: Determine the corresponding number of probability buckets for each sub-circuit, and dynamically allocate the number of probability buckets according to the output state distribution and accuracy requirements of the sub-circuit to ensure efficient storage and merging of intermediate results in subsequent evaluation and post-processing, while balancing computational overhead and simulation accuracy;
[0023] (4) Hybrid calculation of sub-circuit evaluation and post-processing based on probability buckets
[0024] According to the evaluation order determined in step (3.1), the pair of sub-circuits cut in step (2) are read into the GPU at one time and processed as the front-end circuit and the back-end circuit respectively; the evaluation of the front-end circuit uses state vector multiplexing optimization, and at the same time dynamically loads the initialization tasks of the back-end circuit to reduce the video memory occupancy; the state vectors of the two sub-circuits are subjected to state merging and contraction processing using the probability bucket allocation result in step (3.2), and the useless resources are released, and the remaining calculations of the back-end circuit are completed and the results are returned.
[0025] In the present invention, for the heuristic cutting method in step (2), the mixed integer programming model is solved by the Gurobi solver. Gurobi can quickly find an approximate optimal solution within the given time limit through its efficient solution algorithm and parallel computing ability, while ensuring that the quality of the solution meets the actual requirements. This method can not only effectively balance the computational load and communication overhead of the sub-circuits, but also provide a feasible cutting method within a limited time, providing reliable support for the simulation optimization of large-scale quantum circuits.
[0026] In the present invention, the specific steps for determining the evaluation order in step (3.1) are as follows:
[0027] (3.1.1) Small dimension first method: In each iteration, select a pair of adjacent sub-circuits with the smallest current dimension for merging, and add the result to the list to be processed; this process continues until all the derivative evaluation contraction tasks of all sub-circuits in the first contraction stage are completed, or all intermediate results in the second contraction stage are merged.
[0028] (3.1.2) Abstraction and calculation: Abstract the cut sub-circuits into dimension representations, and determine the dimension of each sub-circuit through the output qubits and the number of cutting knives; for each pair of adjacent sub-circuits, calculate the dimension of the intermediate result after their contraction, and select a pair with the smallest generated intermediate result dimension for priority merging.
[0029] (3.1.3) Constraints and optimization: The sub-circuits participating in the contraction in each round cannot participate in subsequent contraction operations. For sub-circuits or intermediate states for which no suitable pairing can be found, they are directly placed at the end of the merging order list.
[0030] (3.1.4) Time complexity: The time complexity is O(k 2 ), where k is the number of sub-circuits or intermediate states participating in the sorting.
[0031] In the present invention, in the evaluation and post-processing hybrid calculation in step (4):
[0032] (1) Front - end circuit evaluation and reuse optimization: Store the evaluation results of the state vectors under the I - basis in the shared memory of the thread block to provide reuse optimization for all measurement bases of the front - end circuit, significantly reducing repeated calculations;
[0033] (2) Back - end circuit evaluation: Only load the initialization tasks required for the first part of the back - end circuit into the GPU for evaluation, avoiding loading all derivative tasks at once, thereby reducing the video memory occupancy;
[0034] (3) State merging and contraction processing: Use the probability bucket allocation results output in step (3.2) to perform state merging on the state vectors of the front - end and back - end sub - circuits, greatly reducing the video memory occupancy. Subsequently, use the contraction method to perform contraction processing on some results;
[0035] (4) Resource release and final calculation: Release the storage space of the state vectors under the Z - basis of the front - end circuit, and reuse the state vectors under the basis in the shared memory to calculate the state vectors under the X - basis and Y - basis, and release them after completion; Finally, calculate the remaining initialization derivative tasks of the back - end circuit, perform bucket processing and contraction operations, and return the final result.
[0036] The beneficial effects of the present invention are as follows: The present invention proposes a method for quantum circuit cutting and simulation optimization combined with GPU acceleration, aiming to maximize the computing potential of the GPU and provide an efficient solution for the simulation of large - scale quantum circuits. By optimizing video memory utilization, reducing data transfer overhead, and designing efficient sub - circuit evaluation and post - processing strategies, the present invention provides a powerful tool for the verification of quantum algorithms and the debugging of quantum hardware. Aiming at the problems of limited GPU video memory capacity and large data transfer overhead, the present invention designs an optimization strategy for the cutting algorithm based on a heuristic method, combines a mixed - integer optimization model to generate a cutting scheme adapted to the GPU hardware characteristics; at the same time, proposes an evaluation - post - processing hybrid simulation algorithm based on probability buckets, which significantly improves the GPU video memory utilization rate and computing efficiency through dynamic state merging and post - processing operations. The present invention is applicable to a single - GPU environment, can effectively reduce the complexity of large - scale quantum circuit simulation, improve the simulation efficiency, and provide efficient support for the verification and debugging of complex quantum algorithms. Description of the Drawings
[0037] Figure 1 The solution process of the heuristic cutting method;
[0038] Figure 2 It is a schematic timeline of sub - circuit evaluation and post - processing based on probability buckets. Detailed Embodiments
[0039] The present invention will be further described below through embodiments in combination with the drawings.
[0040] Example 1: The present invention proposes a quantum circuit cutting simulation optimization method for GPUs, which significantly optimizes the simulation efficiency of large-scale quantum circuits through customized cutting strategies and efficient GPU acceleration technologies.
[0041] Step 1: Circuit preprocessing and establishing qubit dependencies
[0042] By analyzing the gate operation sequence of the quantum circuit, a complete dependency graph between qubits is established. In the circuit preprocessing stage, the quantum circuit is first streamlined by removing all single-qubit gate operations (such as X, H, S gates, etc.), and only keeping two-qubit and multi-qubit gates as key operation nodes, thus abstracting the original quantum circuit into a streamlined directed acyclic graph (DAG). In this DAG representation, each qubit corresponds to a node, and multi-qubit gate operations are transformed into directed edges between nodes. The direction of the edge is determined by the timing of the gate operation, and the weight of the edge reflects the complexity of the gate operation. This representation method not only clearly shows the essential dependency relationship between qubits but also significantly simplifies the subsequent circuit analysis and cutting decision-making process by eliminating the interference brought by single-qubit gates. The preprocessing stage also includes topological feature analysis of the DAG, such as calculating the in-degree / out-degree of nodes, identifying critical paths, etc., providing important basis for subsequent heuristic cutting. The entire preprocessing process has a linear time complexity and can efficiently handle large-scale quantum circuits.
[0043] Step 2: Circuit cutting optimization
[0044] A heuristic cutting method combined with a mixed-integer programming model is used to determine the optimal cutting point. This method evaluates the computational overhead, resource requirements, and variance of qubits in subcircuits of different cutting schemes to ensure that the scale of each subcircuit is within the GPU memory capacity. Although the circuit cutting method can reduce the number of qubits to be processed by decomposing large circuits, in practical applications, key limiting factors of different platforms need to be comprehensively considered. On quantum computers, circuit depth and complexity are the main bottlenecks; while on classical simulators, memory capacity and computing power (such as GPU memory) become the key constraints. In the classical quantum simulator environment, the circuit cutting method must fully consider the dual limitations of the GPU hardware: on the one hand, the storage requirement of quantum states grows exponentially with the number of qubits, and the limited memory capacity strictly restricts the scale of subcircuits that can be processed; on the other hand, there is an inherent contradiction between the efficient parallel computing ability of the GPU and the memory occupation - increasing parallelism can accelerate the calculation but will exacerbate the memory pressure, and reducing parallelism can save memory but cannot fully utilize computing resources. This balance problem between memory capacity and parallel efficiency makes the optimal cutting scheme need to be carefully weighed among the scale of subcircuits, the granularity of computing tasks, and data transfer overhead, becoming the key factor restricting the GPU-accelerated quantum simulation effect.
[0045] Such as Figure 1As shown, the Gurobi solver is used to find an approximate optimal solution within a limited time to minimize the overall simulation burden.
[0046] Step 3: Evaluate preprocessing
[0047] The "minimum dimension first method" is adopted to optimize the sub-circuit merging order. The algorithm preferentially merges adjacent sub-circuits with the smallest dimension, adds the results to the list to be processed until all derivative evaluation contraction tasks are completed. For sub-circuits that cannot be paired, they are placed at the end of the merging order, effectively avoiding GPU memory overflow. At the same time, the state vector reuse technology is applied to store the I-basis evaluation results in the shared memory, significantly reducing repeated calculations; the "dynamic bucket division method" is used to obtain the state vector length limit of each sub-circuit before evaluation and utilize it in the subsequent evaluation process, significantly reducing the data transmission overhead while ensuring the calculation accuracy.
[0048] Step 4: Hybrid calculation based on probability buckets
[0049] The information obtained in the preprocessing is used to optimize the sub-circuit merging process. The sub-circuit contraction process is optimized and designed, and the operation is divided into two key stages: First, in the state vector evaluation of the front-end circuit in the I-basis and Z-basis, efficient utilization is achieved by storing the reusable I-basis state vectors in the shared memory; the back-end circuit only needs to load the |0> and |1> initialization tasks. After the evaluation, the state merging is performed using the bucket results of the "dynamic bucket division method", significantly reducing the video memory occupancy. Subsequently, an intelligent memory management strategy is adopted: the Z-basis state vector space is released in a timely manner, the I-basis state vectors in the shared memory are reused to generate the X / Y-basis state vectors and then released immediately, and finally the |+> and |i> derivative tasks of the back-end circuit are calculated and bucket processed. This phased memory optimization scheme improves the video memory usage efficiency while ensuring the calculation accuracy by dynamically managing the GPU video memory resources. Figure 2 Describes the timeline of the steps executed for each pair of sub-circuits during the hybrid calculation process.
[0050] In specific experiments, six circuit sets, namely Supremacy, AQFT, Adder, HWEA, BV, and Regular, were tested. At the medium-scale qubit level, average accelerations of 40.43%, 64.05%, 55.05%, 19.87%, 21.59%, and 40.36% were achieved respectively. In addition, for the simulation of high-scale qubits from 50 to 60, the present invention for the first time achieved efficient simulation within an acceptable time on a single GPU, providing a new solution for large-scale quantum circuit simulation. For example, in the experiment of the Supremacy circuit, the combination of the heuristic cutting algorithm and the state vector reuse technology significantly improved the simulation efficiency, with an acceleration effect of more than 40%; in the AQFT circuit, the simulation time was reduced by 64.05%. These experimental results demonstrate the wide applicability and high efficiency of the present invention in different types of quantum circuits.
[0051] Through the above method, the present invention has achieved efficient quantum circuit cutting and simulation optimization, significantly improving the GPU resource utilization rate and calculation efficiency, and providing reliable support for the simulation of large-scale quantum circuits.
Claims
1. A quantum circuit cutting simulation optimization method for GPU, characterized in that: The specific steps are as follows: (1) Circuit preprocessing and establishing dependencies between qubits (1.1) Preprocess the target quantum circuit and clean up the single quantum gates and other redundant parts in the quantum circuit to simplify the target quantum circuit; (1.2) analyzing the dependencies between qubits in the quantum circuit preprocessed in step (1.1), and constructing the quantum circuit as a directed acyclic graph (DAG), where nodes represent quantum gates and edges represent data flows between qubits; then, using a mixed integer programming (MIP) method to model the directed acyclic graph, and transforming the optimization problem of the quantum circuit into an objective function and constraints, thereby achieving efficient allocation and optimization of quantum circuit resources; (2) Circuit cutting A heuristic cutting method is used to determine the best cutting point based on the complexity of the quantum circuit and the dependencies between the qubits in step (1.2). The heuristic cutting method selects a scheme that minimizes the overall simulation burden by evaluating the computational overhead and resource requirements of different cutting schemes. (3) Evaluation preprocessing (3.1) Determine the evaluation order: Dynamically determine the evaluation order of subcircuits based on the dimensions of each subcircuit after it is condensed and merged with its adjacent subcircuits; By analyzing the dependencies between subcircuits and the size of the state vector, the minimum dimension priority method is used to prioritize the evaluation of subcircuits with smaller dimensions, thereby reducing computational complexity and memory usage; (3.2) Allocate probability buckets: Determine the number of corresponding probability buckets for each subcircuit. Dynamically allocate the number of probability buckets based on the output state distribution and accuracy requirements of the subcircuit to ensure efficient storage and merging of intermediate results during subsequent evaluation and post-processing, while balancing computational overhead and simulation accuracy. (4) Hybrid calculation of subcircuit evaluation and post-processing based on probability bucket According to the evaluation order determined in step (3.1), the pair of sub-circuits cut in step (2) are read into the GPU at one time and processed as the front-end circuit and the back-end circuit respectively; the evaluation of the front-end circuit uses state vector reuse optimization, and the initialization task of the back-end circuit is dynamically loaded to reduce the memory usage; the state vectors of the two sub-circuits are merged and contracted using the probability bucket allocation result in step (3.2), and useless resources are released, and the remaining calculations of the back-end circuit are completed and the results are returned.
2. The GPU-oriented quantum circuit cutting simulation optimization method according to claim 1, characterized in that: The heuristic cutting method described in step (2) solves the mixed integer programming model through the Gurobi solver. Gurobi quickly finds the approximate optimal solution within a given time limit through its efficient solution algorithm and parallel computing capabilities.
3. The GPU-oriented quantum circuit cutting simulation optimization method according to claim 1, characterized in that: Determine the evaluation order as described in step (3.1), the specific steps are as follows: (3.1.1) Small dimension first method: In each iteration, a pair of adjacent subcircuits with the smallest current dimension are selected for merging, and the result is added to the list to be processed; this process continues until the derivative evaluation and merging tasks of all subcircuits in the first merging stage are completed, or all intermediate results in the second merging stage are merged; (3.1.2) Abstraction and calculation: The cut subcircuits are abstracted into dimensional representations, and the dimension of each subcircuit is determined by the output qubits and the number of cuts. For each pair of adjacent subcircuits, the dimension of their intermediate results after contraction is calculated, and the pair with the smallest intermediate result dimension is selected and merged first. (3.1.3) Constraints and optimization: Subcircuits participating in each round of merging cannot participate in subsequent merging operations. Subcircuits or intermediate states that cannot find a suitable pairing are placed directly at the end of the merging order list. (3.1.4) Time complexity: The time complexity is O(k 2 ), where k is the number of subcircuits or intermediate states involved in the sorting.
4. According to a GPU-oriented quantum circuit cutting simulation optimization method according to claim 1, it is characterized in that: In the mixed calculation of evaluation and post-processing described in step (4): (1) Front-end circuit evaluation and reuse optimization: The state vector evaluation results under the I basis are stored in the shared memory of the thread block to provide reuse optimization for all measurement bases of the front-end circuit, significantly reducing repeated calculations; (2) Back-end circuit evaluation: Only the initialization tasks required for the first part of the back-end circuit are loaded into the GPU for evaluation, avoiding loading all derivative tasks at once, thereby reducing video memory usage; (3) State merging and contraction processing: Use the probability bucket allocation results output from step (3.2) to merge the state vectors of the front-end and back-end subcircuits, greatly reducing the memory usage. Then, the contraction method is used to contract some of the results; (4) Resource release and final calculation: Release the storage space of the state vector under the Z basis of the front-end circuit, and reuse the state vector under the basis in the shared memory to calculate the state vector under the X basis and Y basis, and release it after completion; finally, calculate the remaining initialization derivative tasks of the back-end circuit, perform bucket processing and contraction operations, and return the final result.
Citation Information
Patent Citations
A full-amplitude quantum computing simulation method
CN111832144B
Quantum state coding circuit, quantum calculation method and related device
CN117787428A