A quantum circuit simulation method and device suitable for a multi-GPU system
By employing dynamic partitioning of qubits and constructing communication channels using the NCCL library in a multi-GPU system, the problems of insufficient computing resources in a single GPU and insufficient data exchange in multiple GPUs are solved, thus achieving efficient quantum circuit simulation.
Patent Information
- Application Number
- CN202411164533.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-23
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2044-08-23
AI Technical Summary
Existing quantum circuit simulation methods have limited computing resources in a single-GPU environment, making them unsuitable for the computational needs of large-scale quantum circuits. Furthermore, they lack effective data transmission and exchange mechanisms in a multi-GPU environment, resulting in low simulation efficiency.
By employing dynamic partitioning of qubits, the initial state vector of the quantum system is divided among various GPUs. An efficient communication channel is constructed using the NCCL library to achieve data exchange and quantum gate operations between multiple GPUs. The quantum circuit simulation operations are then executed in parallel by multiple GPUs to obtain the final state vector.
The efficiency of quantum circuit simulation is improved in a multi-GPU environment, overcoming the storage limitations of a single GPU, reducing data exchange requirements, and enhancing simulation performance.
Smart Images

Figure CN119005352B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of quantum computing technology, specifically to the field of quantum circuit simulation, and more specifically, to a quantum circuit simulation method and apparatus suitable for multi-GPU systems. Background Technology
[0002] Quantum computing, a revolutionary technology based on the principles of quantum mechanics, is rapidly gaining prominence in the global scientific research field. It encodes and processes information using qubits, leveraging unique phenomena such as quantum superposition and quantum entanglement to achieve highly efficient parallel operations on information. Quantum algorithms, such as Shor's factorization algorithm and Grover's search algorithm, have demonstrated significant speed improvements over traditional computing methods when dealing with specific types of problems. With continuous technological advancements, quantum computing has shown immense potential application value in various fields, including materials science, cryptography, and combinatorial optimization.
[0003] Currently, quantum computing is still in the so-called "Noisy Medium-Scale Quantum (NISQ) era." In this stage, quantum circuit simulation has become an important tool for researchers to verify the correctness and effectiveness of quantum algorithms and quantum systems. Quantum circuit simulation involves representing the state of a quantum system as a complex vector, the length of which grows exponentially with the number of qubits. During the simulation, by continuously multiplying the state vector with the unitary matrices corresponding to each quantum gate in the quantum circuit, the evolution of the quantum system's state can be gradually derived until the final quantum state is obtained. Despite numerous challenges, the prospects for quantum computing are promising, and it is expected to open a new chapter in the field of computing.
[0004] As the number of qubits increases, the computational complexity of quantum circuit simulation grows exponentially. In a single-machine environment, existing quantum circuit simulation technologies are struggling to meet the time and resource requirements for simulating large-scale quantum circuits. Current quantum circuit simulation techniques typically use state vectors to simulate quantum circuits, focusing only on the functional implementation of quantum system state vectors and quantum gate matrix operations. As the scale of quantum circuits increases, a single CPU or GPU struggles to meet the computational demands of large-scale quantum circuit simulation. Furthermore, existing quantum circuit simulation techniques generally do not consider how to perform quantum circuit simulations across multiple GPUs. Firstly, in a multi-GPU environment, the state vectors of the quantum system are difficult to store in the same memory space, meaning each GPU has incomplete access to quantum states, making simulation impossible with existing quantum computing methods. Secondly, how to transmit and exchange quantum state vectors between GPUs is another issue that existing quantum computing simulation techniques have not considered. With the development of quantum computing, the number of qubits in quantum computers is constantly increasing; existing quantum computers have reached scales of over one hundred qubits, and quantum circuit simulation technology faces increasingly greater challenges.
[0005] In summary, existing quantum circuit simulation methods suffer from the following problems: 1) Limited computational resources in a single-GPU environment: The dimension of the state vector of a quantum system increases exponentially with the number of qubits, making the memory and computing power of a single GPU insufficient for simulating large-scale quantum circuits. 2) Inability to simulate quantum circuits in a multi-GPU environment: In a multi-GPU environment, the incomplete storage of the state vectors encountered on each GPU in the same memory space is a limitation, making existing simulation methods unsuitable for multi-GPU environments. 3) Lack of inter-GPU data transfer and exchange mechanisms: Existing quantum circuit simulation methods do not consider data exchange mechanisms in multi-GPU environments. Therefore, a method suitable for simulating large-scale quantum circuits in multi-GPU systems is urgently needed.
[0006] It should be noted that the background information presented here is only for illustrating relevant information about the present invention to aid in understanding the technical solutions of the present invention, and does not imply that the relevant information is necessarily prior art. In the absence of evidence indicating that the relevant information was disclosed before the filing date of this invention, the relevant information should not be considered prior art. Summary of the Invention
[0007] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide a quantum circuit simulation method and apparatus suitable for multi-GPU systems.
[0008] The objective of this invention is achieved through the following technical solution:
[0009] According to a first aspect of the present invention, a quantum circuit simulation method suitable for multi-GPU systems is proposed for simulating the operation of a quantum circuit on a specified quantum system to obtain the final state vector of the quantum system. The method includes: step S1, obtaining the number of GPUs and constructing a communication channel between GPUs, obtaining all qubits of the quantum circuit to be simulated, the order of the qubits, the quantum gate operation sequence, and the initial state vector of the specified quantum system, wherein the quantum gate operation sequence includes multiple ordered quantum gate operation layers, each quantum gate operation layer includes multiple parallel-executable quantum gate operations, and each quantum gate operation acts on one or more qubits; step S2, using a preset method to select multiple partition qubits from all qubits of the quantum circuit for each quantum gate operation layer to partition the initial state vector of the quantum system, so as to construct a dynamic partition qubit plan; step S3, based on the dynamic partition qubit plan, for each partition qubit acting on the partition qubit... The process involves: Step S4: Adding GPU data exchange operations between each pair of consecutive quantum gate operation layers with different partitioned qubits, and transforming all quantum gate operations into multi-GPU supported quantum gate operations. A quantum circuit simulation operation sequence is constructed using all GPU data exchange operations and all multi-GPU supported quantum gate operations. Step S5: Based on the partitioned qubits of the first quantum gate operation layer in the quantum circuit, the initial state vector of the specified quantum system is evenly divided among each GPU to initialize the state vector in each GPU. All operations in the quantum circuit simulation operation sequence are executed in parallel on each GPU to update its state vector. When executing quantum gate operations on partitioned qubits, each GPU uses a preset data exchange method to perform data exchange operations between the current GPU and other GPUs. Step S6: Collecting the updated state vectors from all GPUs and rearranging them according to the order of the qubits in the quantum circuit to obtain the final state vector of the quantum system.
[0010] Preferably, the communication channel between GPUs is built based on the NCCL library.
[0011] Preferably, in step S2, the preset method involves selecting partitioned qubits for each quantum gate operation layer from all qubits of the quantum circuit through the following steps: Step S21: Configure the partitioning cost of each quantum gate operation according to the data exchange and computational load required for different types of quantum gate operations to act on the partitioned qubits, wherein the larger the data exchange and computational load, the larger the partitioning cost; Step S22: Obtain the partitioning cost of each qubit in the quantum circuit to be simulated when it is selected as a partitioned qubit in each quantum gate operation layer based on the partitioning cost configured for different quantum gate operations in step S21, wherein the partitioning cost of a qubit in a quantum gate operation layer is the partitioning cost of all quantum gate operations acting on that qubit in that quantum gate operation layer. S23. Divide all quantum gate operation layers into multiple quantum gate operation layer intervals, where each quantum gate operation layer interval contains multiple consecutive quantum gate operation layers, and the partitioned qubits of each quantum gate operation layer are the same; S24. For every two consecutive quantum gate operation layer intervals, calculate the sum of the partitioning cost of the second quantum gate operation layer and the first preset threshold, and the partitioning cost when the partitioned qubits of the second quantum gate operation layer are the same as those of the first quantum gate operation layer, and compare the two values. If the former is less than or equal to the latter, then update the partitioned qubits of all quantum gate operation layers in the second quantum gate operation layer interval to the same partitioned qubits as those of the quantum gate operation layers in the first quantum gate operation layer interval.
[0012] Preferably, step S23 includes: step S231, dividing the quantum circuit into multiple qubit groups containing different qubits based on all the qubits in the quantum circuit, wherein each qubit group contains A different quantum bit, The number of GPUs is given. Step S232: For each qubit group, the qubits in the qubit group are used as the partition qubits of all quantum gate operation layers. The partition cost of each quantum gate operation layer is calculated. The partition cost of each quantum gate operation layer is accumulated layer by layer according to the order of the quantum gate operation layers. If the accumulated value is less than a preset threshold, the interval containing the most quantum gate operation layers is found, and the number of quantum gate operation layers in the interval and the quantum gate operation layers it contains are obtained. The partition cost of each quantum gate operation layer is the sum of the partition costs of all partition qubits of the quantum gate operation layer. Step S233: The interval corresponding to the qubit group with the most quantum gate layers is divided into a quantum gate operation interval, and the qubits in the qubit group are used as the partition qubits of all quantum gate operation layers in the quantum gate operation interval. Step S234: Based on the remaining quantum gate operation layers in the quantum gate operation sequence, steps S231 to S233 are repeated until all quantum gate operation layers are divided.
[0013] Preferably, the GPU data exchange operation includes: a full data exchange operation, which is an operation of exchanging all state vectors in the current GPU with all state vectors in other GPUs; a half data exchange operation, which is an operation of exchanging the first half of the state vector of the current GPU with the second half of the state vector of other GPUs, or exchanging the second half of the state vector of the current GPU with the first half of the state vector of other GPUs; a quarter data exchange operation, which is an operation of dividing the state vector of the current GPU into four equal parts, selecting three parts and exchanging them with the state vectors of three other GPUs respectively; and a quantum bit data exchange operation, which is an operation of selecting the state vectors to be exchanged and performing data exchange based on the quantum bit information of the current GPU and other GPUs as needed.
[0014] Preferably, the state vector of each GPU is initialized as follows: Based on the segmented qubits of the first quantum gate operation layer, the state of the segmented qubits mapped on each GPU is initialized as follows:
[0015]
[0016]
[0017] in, Indicates the GPU serial number. Indicates the sequence number is The segmented qubit states mapped on the GPU. The value is the number of split qubits minus 1. This represents the binary string formed by concatenating the states of each segmented qubit. The GPU index is equal to the decimal value corresponding to the binary string formed by concatenating the states of the segmented qubits mapped on that GPU. The state of each segmented qubit is either 1 or 0. For each GPU, the portion of the initial state vector of the specified quantum system that is identical to the state of the segmented qubits mapped on that GPU is obtained, and the state vectors of the other qubits besides the segmented qubits are extracted from it as the state vector of that GPU to initialize the state vector of each GPU.
[0018] Preferably, the preset data exchange method is configured as follows: obtaining the portion of the current GPU's state vector that needs to be exchanged and the target GPU for data exchange, and obtaining the portion of the state vector and transmitting it to the target GPU; transmitting the portion of the target GPU's state vector that needs to be exchanged to the current GPU, and storing the state vector transmitted from the current GPU; performing a quantum gate operation on the current GPU to update the quantum state vector, extracting the portion of the state vector transmitted from the target GPU from the updated quantum state vector and returning it to the target GPU, and returning the stored state vector of the current GPU in the target GPU to the current GPU.
[0019] According to a second aspect of the present invention, a quantum circuit simulation apparatus is provided for implementing the method described in any of the first aspects of the present invention, for simulating the operation of a quantum circuit on a specified quantum system to obtain the final state vector of the quantum system. The apparatus includes: a communication module for using a communication channel built based on the NCCL library to perform data exchange operations between GPUs; and a quantum system state vector segmentation module for selecting quantum gate operation layers from all qubits of the quantum circuit to be simulated. The system employs a partitioning qubit module to construct a dynamic partitioning qubit plan by using the partitioning cost to partition the initial state vector of the quantum system; a quantum circuit simulation module to construct a sequence of quantum circuit simulation operations including all GPU data exchange operations and all multi-GPU supported quantum gate operations, and to initialize the state vector on the GPUs, and execute the sequence of quantum circuit simulation operations in parallel on all GPUs to update the state vector on each GPU; and a state vector integration module to collect the updated state vectors from all GPUs and sort them according to the order of the qubits in the quantum circuit to be simulated to obtain the final state vector of the quantum device.
[0020] According to a third aspect of the present invention, a computer-readable storage medium is provided thereon storing a computer program that can be executed by a processor to implement the steps of the method described in any one of the first or second aspects of the present invention.
[0021] According to a fourth aspect of the present invention, an electronic device is provided, comprising: one or more processors; and a memory for storing executable instructions; wherein the one or more processors are configured to implement the steps of the method described in any one of the first or second aspects of the present invention by executing the executable instructions.
[0022] Compared with the prior art, the advantages of the present invention are as follows:
[0023] This invention enables quantum circuit simulation tasks to be performed in a multi-GPU environment, overcoming the limitations of single-machine GPUs in storing large-scale quantum circuit state vectors and significantly improving simulation efficiency. In this method, a specific algorithm is used to plan the dynamic partitioning of qubits with the lowest possible cost, thereby rationally distributing the state vectors across GPUs and reducing the need for cross-GPU data exchange. Furthermore, an efficient communication channel built using the NCCL library further accelerates information transmission between GPUs, thus enhancing the simulation performance of quantum circuits. Attached Figure Description
[0024] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:
[0025] Figure 1 This is a flowchart of a quantum circuit simulation method for multi-GPU systems according to an embodiment of the present invention;
[0026] Figure 2 This is a flowchart of a quantum circuit simulation method for multi-GPU systems according to an embodiment of the present invention;
[0027] Figure 3 This is a structural diagram of a quantum circuit simulation device suitable for multi-GPU systems according to an embodiment of the present invention;
[0028] Figure 4 This is a structural diagram of another quantum circuit simulation device suitable for multi-GPU systems according to an embodiment of the present invention. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the invention is further described in detail below through specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0030] As mentioned in the background section, existing quantum circuit simulation methods suffer from the following problems: 1) Limited computing resources in a single GPU environment: The dimension of the state vector of a quantum system increases exponentially with the number of qubits, making the memory and computing power of a single GPU environment insufficient to simulate large-scale quantum circuits. 2) Inability to use multi-GPU environments for quantum circuit simulation: In a multi-GPU environment, the incomplete state vectors encountered on each GPU are difficult to store in the same memory space, making existing simulation methods unsuitable for multi-GPU environments. 3) Lack of inter-GPU data transfer and exchange mechanisms: Existing quantum circuit simulation methods do not consider data exchange mechanisms in multi-GPU environments.
[0031] To address these issues, this invention proposes a quantum circuit simulation scheme suitable for multi-GPU systems. In this scheme, the initial state vector of the quantum system is partitioned across each GPU, and the operation sequence of the quantum circuit is executed in parallel on each GPU to update the state vector on each GPU. Finally, the state vectors of each GPU are collected and reordered to obtain the final state vector of the quantum system, thereby enabling the simulation of large-scale quantum circuits on a multi-GPU system. Furthermore, to address the problem of the large amount of data exchange between GPUs affecting the efficiency of quantum circuit simulation in multi-GPU system simulations, this invention constructs a dynamic quantum bit partitioning plan with low data exchange and computational cost, and uses this plan to construct a quantum gate operation sequence for quantum circuit simulation, thereby reducing data exchange operations between GPUs. Specifically, when partitioning the initial state vector of the quantum system, a dynamic quantum bit partitioning plan with the minimum partitioning cost is constructed, and the state vector of the quantum system is partitioned to each GPU accordingly. Further, addressing the lack of efficient data transmission mechanisms between GPUs in existing quantum circuit simulation methods, this invention constructs a communication channel between GPUs based on NCCL and proposes data exchange operations based on this communication channel to achieve various efficient communication methods between GPUs.
[0032] Before describing the embodiments of the present invention in detail, some of the terms used therein are explained as follows:
[0033] NCCL (NVIDIA Collective Communications Library): NVIDIA Collective Communications Library. NCCL is a high-performance communication library specifically designed for parallel computing applications using NVIDIA GPUs to achieve efficient data transfer and communication. It supports various communication methods, such as point-to-point (send / receive), broadcast, reduce, and reduce scatter.
[0034] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0035] In summary, this invention proposes a quantum circuit simulation method suitable for multi-GPU systems, used to simulate the operation of quantum circuits on a specified quantum system to obtain the final state vector of the quantum system. The flowchart of the method is shown below. Figure 1As shown, it includes: Step S1, obtaining the number of GPUs and constructing communication channels between GPUs, obtaining all qubits of the quantum circuit to be simulated, the order of the qubits, the sequence of quantum gate operations, and the initial state vector of the specified quantum system, wherein the sequence of quantum gate operations includes multiple ordered quantum gate operation layers, each quantum gate operation layer includes multiple parallel-executable quantum gate operations, and each quantum gate operation acts on one or more qubits; Step S2, using a preset method to select multiple partition qubits from all qubits of the quantum circuit for each quantum gate operation layer to partition the initial state vector of the quantum system, so as to construct a dynamic partition qubit plan; Step S3, based on the dynamic partition qubit plan, adding GPU data intersection between each quantum gate operation acting on the partition qubit and between every two consecutive quantum gate operation layers with different partition qubits. The process involves several steps: Step S4, where the initial state vector of the specified quantum system is evenly divided among the partitioned qubits of the first quantum gate operation layer in the quantum circuit to initialize the state vector in each GPU. All operations in the quantum circuit simulation operation sequence are then executed in parallel on each GPU to update its state vector. During the execution of quantum gate operations on the partitioned qubits, each GPU performs data exchange operations between the current GPU and other GPUs using a preset data exchange method. Step S5, the updated state vectors in all GPUs are collected and rearranged according to the order of the qubits in the quantum circuit to obtain the final state vector of the quantum system.
[0036] To better understand the present invention, each step will be described in detail below with reference to specific embodiments.
[0037] According to one embodiment of the present invention, in step S1, a communication channel between GPUs is constructed as follows: a data transmission channel between N GPUs in a multi-GPU system is constructed based on the NCCL library. This inter-GPU communication channel allows for efficient data exchange between multiple GPUs, providing methodological support for quantum state data exchange between multiple GPUs during quantum circuit simulation. The present invention provides eight data communication methods: point-to-point (send / receive), broadcast, reduce, reduce scatter, all reduce, scatter, gather, and all gather. It should be understood that the data communication methods listed herein are exemplary and not exhaustive; the data communication methods are not limited to the above eight types, and other data communication methods capable of realizing quantum state data exchange operations are also acceptable.
[0038] According to an embodiment of the present invention, step S2 includes: step S21, configuring the partitioning cost of each quantum gate operation according to the data exchange and computational cost required for different types of quantum gate operations to partition qubits, wherein the larger the data exchange and computational cost, the larger the partitioning cost; step S22, obtaining the partitioning cost of each qubit in the quantum circuit to be simulated when it is selected as a partitioned qubit in each quantum gate operation layer based on the partitioning cost configured for different quantum gate operations in step S21, wherein the partitioning cost of a qubit in a quantum gate operation layer is the sum of the partitioning costs of all quantum gate operations acting on that qubit in that quantum gate operation layer; step S23, all quantum gates... The operation layer is divided into multiple quantum gate operation layer intervals, where each quantum gate operation layer interval contains multiple consecutive quantum gate operation layers, and the partitioned qubits of each quantum gate operation layer are the same. In step S24, for every two consecutive quantum gate operation layer intervals, the sum of the partitioning cost of the second quantum gate operation layer and a first preset threshold, and the partitioning cost when the partitioned qubits of the second quantum gate operation layer are the same as those of the first quantum gate operation layer, are calculated. The two values are compared; if the former is less than or equal to the latter, the partitioned qubits of all quantum gate operation layers in the second quantum gate operation layer interval are updated to the same partitioned qubits as those in the first quantum gate operation layer interval. By executing the above steps S21-S24, a dynamic qubit planning can be constructed.
[0039] According to one embodiment of the present invention, in step S21, when configuring the partitioning cost of each quantum gate operation based on the amount of data exchange and computation required for the quantum gate operation to partition the qubits, it can be configured as follows:
[0040] 1) For a general quantum gate matrix, when it is used to partition a qubit, it is necessary for the GPUs that have the state vectors corresponding to the state 0 and state 1 of the qubit to perform half of the quantum state vector data exchange, and the partitioning cost is configured to be 2.
[0041] 2) For diagonal quantum gate matrices, when they are used to split qubits, no additional data exchange operation is required, and the splitting cost is configured to be 0;
[0042] 3) When the inverse diagonal quantum gate matrix is used to split a qubit, the GPUs that have the state vectors of the qubit state 0 and state 1 need to exchange all the state vectors. This quantum gate operation avoids data exchange between GPUs by adjusting the GPU index, and its splitting cost is configured to be 0.5.
[0043] 4) For measurement and reset quantum gate operations, when applied to splitting qubits, a measurement data exchange and merging operation is required to calculate the final measurement result of the quantum state. The splitting generation configuration is 1.
[0044] 5) For other quantum gate matrices, when they are used to split qubits, it is necessary to add a state vector exchange operation between GPUs based on the qubits, and the splitting cost is configured as 3.
[0045] It should be understood that the partitioning cost of each type of quantum gate operation configuration described above is only one of the specific embodiments, and those skilled in the art can modify the partitioning cost of each type of quantum gate to obtain other embodiments.
[0046] According to one embodiment of the present invention, in step S22, the partitioning cost when each qubit in the quantum circuit to be simulated is selected as a partitioned qubit is obtained as follows: The partitioning cost of all qubits in the quantum circuit to be simulated is calculated based on the partitioning costs of different types of quantum gates, wherein the partitioning cost corresponding to each qubit at each quantum gate operation layer is the sum of the partitioning costs of all quantum gate operations acting on that qubit at that quantum gate operation layer. The partitioning cost when each qubit is selected as a partitioned qubit at each quantum gate operation layer can be obtained in the above manner. Based on all the obtained partitioning costs, a partitioning cost table D can be constructed, with the number of columns being the number of qubits in the quantum circuit and the number of rows being the number of quantum gate operation layers in the quantum circuit, wherein each element D in the table... ij (i represents the row, j represents the column) represents the partitioning cost when the j-th qubit is selected as a partition bit on the i-th quantum gate operation layer.
[0047] According to an embodiment of the present invention, step S23 includes: step S231, dividing the quantum circuit into multiple qubit groups containing different qubits based on all qubits in the quantum circuit, wherein each qubit group contains A different quantum bit, The number of GPUs is given. Step S232: For each qubit group, the qubits in the qubit group are used as the partition qubits of all quantum gate operation layers. The partition cost of each quantum gate operation layer is calculated. The partition cost of each quantum gate operation layer is accumulated layer by layer according to the order of the quantum gate operation layers. If the accumulated value is less than a preset threshold, the interval containing the most quantum gate operation layers is found, and the number of quantum gate operation layers in the interval and the quantum gate operation layers it contains are obtained. The partition cost of each quantum gate operation layer is the sum of the partition costs of all partition qubits in the quantum gate operation layer. The partition cost of each qubit in each quantum gate operation layer is obtained by querying the constructed partition cost table. Step S233: The interval corresponding to the qubit group with the most quantum gate layers is divided into a quantum gate operation interval, and the qubits in the qubit group are used as the partition qubits of all quantum gate operation layers in the quantum gate operation interval. Step S234: Based on the remaining quantum gate operation layers in the quantum gate operation sequence, steps S231 to S233 are repeated until all quantum gate operation layers are divided. Perform steps S231 to S234 above to divide all quantum gate operation layers into multiple quantum gate operation layer intervals.
[0048] According to one embodiment of the present invention, in step S231, the number of qubit groups divided is: That is, selecting qubits without replacement from all the qubits contained in the quantum circuit. The types of ways to select a qubit, among which... This indicates the number of qubits contained in a quantum circuit. The calculation formula is as follows:
[0049]
[0050] According to one embodiment of the present invention, in step S3, constructing the quantum circuit simulation operation sequence includes: 1) converting all quantum gate operations into multi-GPU supported quantum gate operations; 2) adding GPU data exchange operations between each quantum gate operation acting on a segmented qubit and between every two consecutive quantum gate operation layers with different segmented qubits. The GPU data exchange operations in the quantum circuit simulation operation sequence are used to handle data exchange operations between GPUs when quantum gate operations act on segmented qubits or when the segmented qubits change, thereby realizing the simulated quantum circuit on a multi-GPU system. It should be understood that converting quantum gate operations into multi-GPU supported quantum gate operations is a technique well-known to those skilled in the art and will not be described in detail here.
[0051] According to an embodiment of the present invention, in step S3, the data exchange operation between GPUs includes a full data exchange operation, a half data exchange operation, a quarter data exchange operation, a bit-based data exchange operation, and a measurement data exchange merging operation. The full data exchange operation involves exchanging all state vectors in the current GPU with all state vectors in other GPUs. The half data exchange operation involves exchanging the first half of the current GPU's state vector with the second half of the state vectors in other GPUs, or exchanging the second half of the current GPU's state vector with the first half of the state vectors in other GPUs. The quarter data exchange operation involves dividing the current GPU's state vector into four equal parts and selecting three parts to exchange with the state vectors of three other GPUs. The bit-based data exchange operation involves selecting the state vectors to be exchanged based on the qubit information of the current GPU and other GPUs and performing the data exchange. The measurement data exchange merging operation involves using a global reduction data merging operation to merge the quantum state measurement results from all GPUs into the data exchange in the current GPU. This invention provides multiple data exchange methods for dynamically selecting based on the required amount of data when simulating quantum circuits, effectively controlling the amount of data transmitted each time, reducing data transmission time, and improving the efficiency of quantum circuit simulation.
[0052] According to an embodiment of the present invention, in step S4, the state vector of each GPU is initialized as follows: the state of the segmented qubits mapped on each GPU is initialized based on the segmented qubits of the first quantum gate operation layer, wherein the state of the segmented qubits mapped on each GPU is represented as:
[0053]
[0054]
[0055] in, Indicates the GPU serial number. Indicates the sequence number is The segmented qubit states mapped on the GPU. The value is the number of split qubits minus 1. This represents the binary string formed by concatenating the states of each segmented qubit. The GPU index is equal to the decimal value corresponding to the binary string formed by concatenating the states of the segmented qubits mapped on that GPU. The state of each segmented qubit is either 1 or 0. For each GPU, the portion of the initial state vector of the specified quantum system that is the same as the state of the segmented qubits mapped on that GPU is obtained. From this, the state vectors of the other qubits besides the segmented qubits are extracted as the state vector of that GPU to initialize the state vector of each GPU.
[0056] To better understand the process of initializing the state vector on each GPU, the following description will be provided with examples.
[0057] Assume the quantum circuit has 10 qubits (Q0...Q9), and the multi-GPU system has 4 GPUs (GPU0, GPU1, GPU2, GPU3). The first quantum gate operation layer of the quantum circuit is divided into qubits Q0 and Q1. The initial state vectors on each GPU are GPU0 =
[00] [00000000,…,111111111], GPU1 =
[01] [00000000,…,111111111], GPU2 =
[10] [00000000,…,111111111], GPU3 =
[11] [00000000,…,111111111], where the index of each GPU is equal to the decimal value corresponding to the binary state of the segmented qubit, that is, the state of the segmented qubit corresponding to GPU0 is
[00] . The initial quantum state vector on each GPU is the state vector of the remaining qubits whose state of the segmented qubit is equal to its GPU index in the initial state vector of the quantum system. For example, the state vector corresponding to GPU0 is the decimal value corresponding to the binary state of the segmented qubits Q0 and Q1, which is equal to 0. The state vectors corresponding to all remaining qubits (Q2~Q9) in the initial state vector of the quantum system are, that is, the initial state vector on GPU0 is, when the segmented qubits Q0Q1=00, all the state vectors corresponding to (Q2……Q9)∈(00000000~11111111).
[0058] According to one embodiment of the present invention, in step S4, after initializing the state vector on each GPU, all operations in the quantum circuit simulation operation sequence are executed in parallel on each GPU to update its state vector. Specifically, during quantum gate operations in the quantum circuit simulation operation sequence, a quantum gate matrix multiplication algorithm is selected to execute each quantum gate operation based on its properties. The quantum gate matrix multiplication algorithms include: general quantum gate matrix multiplication, diagonal quantum gate matrix multiplication, inverse diagonal quantum gate matrix multiplication, single-bit controlled quantum gate matrix multiplication, quantum gate matrix multiplication with control bits, and special two-bit quantum gate matrix multiplication. It should be understood that the quantum gate matrix multiplication algorithms listed herein are exemplary and not exhaustive; other quantum gate matrix multiplication algorithms that can be used to perform quantum gate operations are also possible.
[0059] According to an embodiment of the present invention, in step S4, when all operations in the quantum circuit simulation operation sequence are executed in parallel in each GPU, the data exchange operation between GPUs in the quantum circuit simulation operation sequence is performed as follows: The portion of the current GPU's state vector that needs to be exchanged for data and the target GPU for data exchange are obtained, and the portion of the state vector is obtained and transmitted to the target GPU; the portion of the target GPU's state vector that needs to be exchanged for data is transmitted to the current GPU, and the state vector transmitted from the current GPU is stored; a quantum gate operation is performed on the current GPU to update the quantum state vector, the portion of the state vector transmitted from the target GPU is extracted from the updated quantum state vector and returned to the target GPU, and the state vector of the current GPU stored in the target GPU is returned to the current GPU.
[0060] According to one embodiment of the present invention, in step S4, after all operations in the quantum circuit simulation operation sequence are executed in parallel in each GPU, the updated state vectors in all GPUs are collected and rearranged according to the order of the qubits in the quantum circuit to obtain the final state vector of the quantum system.
[0061] According to an embodiment of the present invention, the program flow corresponding to the quantum circuit simulation method for multi-GPU systems proposed in this invention is as follows: Figure 2 As shown in the figure, the program flow includes:
[0062] 1) Initialize the communication channels between multiple GPUs;
[0063] 2) Input quantum circuit and the required number of GPUs;
[0064] 3) Based on the input quantum circuits and the number of GPUs, select the optimal number of qubits to divide and generate a series of quantum circuit simulation operations, including quantum gate matrix multiplication operations and quantum state data exchange operations;
[0065] 4) Parallel multi-process startup of multi-GPU quantum circuit simulator, with each process corresponding to one graphics processing unit (GPU);
[0066] 5) Based on the segmented qubits, construct the corresponding partial quantum state vector in each GPU;
[0067] 6) Traverse the series of quantum circuit simulation operations, obtain and execute quantum circuit simulation operations. For data exchange operations, obtain the corresponding quantum state vector and target GPU information, send the quantum state data to be exchanged to the target GPU, and update the state vector on the GPU after waiting for the data to return. For quantum gate operations, perform matrix multiplication between the quantum gate matrix representing the quantum gate operation and the state vector on the GPU to update the state vector.
[0068] 7) Wait for all processes to finish their computations, and then summarize the state vectors from all GPUs;
[0069] 8) Reorder the aggregated quantum state vectors according to the order of the qubits in the quantum circuit;
[0070] 9) Output the final quantum system state vector. According to one embodiment of the present invention, a quantum circuit simulation device for implementing the quantum circuit simulation method suitable for multi-GPU systems is proposed, the structure of which is as follows: Figure 3 As shown, it includes: a communication module for building a communication channel based on the NCCL library and performing data exchange operations between GPUs; and a quantum system state vector segmentation module for selecting quantum gate operation layers from all qubits of the quantum circuit to be simulated. The system employs a partitioning qubit module to construct a dynamic partitioning qubit plan by using the partitioning cost to partition the initial state vector of the quantum system; a quantum circuit simulation module to construct a sequence of quantum circuit simulation operations including all GPU data exchange operations and all multi-GPU supported quantum gate operations, and to initialize the state vector on the GPUs, and execute the sequence of quantum circuit simulation operations in parallel on all GPUs to update the state vector on each GPU; and a state vector integration module to collect the updated state vectors from all GPUs and sort them according to the order of the qubits in the quantum circuit to be simulated to obtain the final state vector of the quantum device.
[0071] According to another embodiment of the present invention, the present invention proposes a quantum circuit simulation device for implementing the quantum circuit simulation method suitable for multi-GPU systems, the device having the following structure: Figure 4As shown, it includes: a communication module for building a communication channel based on the NCCL library and performing data exchange operations between GPUs; a quantum circuit construction module for generating quantum circuits, which can be customized to generate quantum circuits containing any number of qubits and any type of quantum gates; and a quantum system state vector segmentation module for selecting quantum gate operation layers from all qubits of the quantum circuit to be simulated. The system employs a partitioning qubit module to construct a dynamic partitioning qubit plan by using the partitioning cost to partition the initial state vector of the quantum system; a quantum circuit simulation module to construct a sequence of quantum circuit simulation operations including all GPU data exchange operations and all multi-GPU supported quantum gate operations, and to initialize the state vector on the GPUs, and execute the sequence of quantum circuit simulation operations in parallel on all GPUs to update the state vector on each GPU; and a state vector integration module to collect the updated state vectors from all GPUs and sort them according to the order of the qubits in the quantum circuit to be simulated to obtain the final state vector of the quantum device.
[0072] To better illustrate the beneficial effects of this invention, the inventors designed a comparative experiment to evaluate the performance of the proposed quantum circuit simulation method. The comparative experiment involved simulating the same quantum circuit using both a single-GPU quantum circuit simulation method in a single-GPU environment and a dual-GPU environment using the proposed quantum circuit simulation method suitable for multi-GPU systems. The time required for each quantum circuit simulation method to simulate the same quantum circuit was then recorded. The experimental results are shown in Table 1. The data in the table show that as the number of qubits increases, the proposed quantum circuit simulation method suitable for multi-GPU systems significantly reduces the time required compared to the traditional single-GPU quantum circuit simulation method, confirming the significant advantage of this invention in terms of efficiency for large-scale quantum circuit simulation.
[0073] Table 1 Experimental Results
[0074] The number of qubits contained in a quantum circuit Simulation time for a single GPU environment (s) Environment simulation duration (s) for two GPUs 25 0.163 0.198 28 1.42 1.12 30 6.12 4.43
[0075] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.
[0076] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.
[0077] Computer-readable storage media can be tangible devices that hold and store instructions for use by an instruction execution device. Computer-readable storage media can include, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof.
[0078] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A quantum circuit simulation method suitable for a multi-GPU system, for simulating a running process of a quantum circuit on a specified quantum system to obtain a final state vector of the quantum system, characterized in that, The method comprises: Step S1, obtaining the number of GPUs and constructing the communication channel between the GPUs, obtaining all the quantum bits of the quantum circuit to be simulated, the order of the quantum bits, the sequence of quantum gate operations, and the initial state vector of the specified quantum system, wherein the sequence of quantum gate operations comprises a plurality of ordered quantum gate operation layers, each quantum gate operation layer comprises a plurality of quantum gate operations that can be executed in parallel, and each quantum gate operation acts on one or more quantum bits; Step S2, selecting a plurality of partition quantum bits for partitioning the initial state vector of the quantum system for each quantum gate operation layer from all the quantum bits of the quantum circuit using a preset method to construct a dynamic partition quantum bit plan; Step S3, based on the dynamic partition quantum bit plan, adding a GPU data exchange operation between each quantum gate operation acting on the partition quantum bit and each quantum gate operation layer that is continuous and different from the partition quantum bit, and converting all the quantum gate operations into multi-GPU supported quantum gate operations to construct a quantum circuit simulation operation sequence with all the GPU data exchange operations and all the multi-GPU supported quantum gate operations; Step S4, based on the partition quantum bits of the first quantum gate operation layer in the quantum circuit, equally partitioning the initial state vector of the specified quantum system to each GPU to initialize the state vector in each GPU, and executing all the operations in the quantum circuit simulation operation sequence in parallel in each GPU to update the state vector thereon, wherein each GPU performs a data exchange operation between the current GPU and other GPUs using a preset data exchange method when executing the quantum gate operation acting on the partition quantum bit; Step S5, collecting the updated state vectors in all the GPUs and rearranging them according to the order of the quantum bits in the quantum circuit to obtain the final state vector of the quantum system.
2. The method of claim 1, wherein, The communication channel between the GPUs is constructed based on the NCCL library.
3. The method of claim 1, wherein, In the step S2, the preset method is to select partition quantum bits for each quantum gate operation layer from all the quantum bits of the quantum circuit by the following steps: Step S21, configuring the partition cost of each quantum gate operation according to the data exchange amount and the computation amount required for the quantum gate operation to act on the partition quantum bit, wherein the larger the data exchange amount and the computation amount, the larger the partition cost; Step S22, obtaining the partition cost of each quantum bit in the quantum circuit to be simulated when it is selected as a partition quantum bit in each quantum gate operation layer based on the partition cost of different quantum gate operations configured in step S21, wherein the partition cost of a quantum bit in a quantum gate operation layer is the sum of the partition costs of all the quantum gate operations acting on the quantum bit in the quantum gate operation layer; Step S23, dividing all the quantum gate operation layers into a plurality of quantum gate operation layer intervals, wherein each quantum gate operation layer interval contains a plurality of continuous quantum gate operation layers, and the partition quantum bits of each quantum gate operation layer are the same. In step S24, for each two continuous quantum gate operation layer intervals, a sum of a split cost of a second quantum gate operation interval and a first preset threshold, and a split cost of the second quantum gate operation interval in a case that split qubits of the second quantum gate operation interval are the same as split qubits of a first quantum gate operation interval are calculated, and the two values are compared, if the former is less than or equal to the latter, split qubits of all quantum gate operation layers in the second quantum gate operation layer interval are updated to be the same as split qubits of the first quantum gate operation layer interval.
4. The method of claim 3, wherein, The step S23 comprises: Step S231, based on all the quantum bits in the quantum circuit, a plurality of quantum bit groups containing different quantum bits are divided, wherein each quantum bit group contains different quantum bits, is the number of GPUs; In step S232, for each qubit group, taking qubits in the qubit group as split qubits of all quantum gate operation layers, a split cost of each quantum gate operation layer is calculated, and split costs of quantum gate operation layers are accumulated layer by layer in order, and an interval containing the most quantum gate operation layers is found under the condition that the accumulated value is less than a preset threshold, and the number of quantum gate operation layers in the interval and the quantum gate operation layers contained in the interval are obtained, wherein the split cost of each quantum gate operation layer is the sum of split costs of all split qubits of the quantum gate operation layer. In step S233, the interval corresponding to the qubit group with the most quantum gate layers is divided into a quantum gate operation interval, and qubits in the qubit group are taken as split qubits of all quantum gate operation layers in the quantum gate operation interval. In step S234, based on the remaining quantum gate operation layers in the quantum gate operation sequence, steps S231 to S233 are repeated until all quantum gate operation layers are divided.
5. The method of claim 1, wherein, The GPU data exchange operation comprises: A full data exchange operation, which is an operation of exchanging all state vectors in a current GPU with all state vectors in other GPUs; A half data exchange operation, which is an operation of exchanging a front half of a state vector in the current GPU with a back half of a state vector in other GPUs, or exchanging a back half of a state vector in the current GPU with a front half of a state vector in other GPUs; A quarter data exchange operation, which is an operation of equally dividing a state vector in the current GPU into four parts, and selecting three parts to perform data exchange with state vectors in three other GPUs respectively; According to the qubit data exchange operation, which is an operation of selecting a state vector to be exchanged according to qubit information of the current GPU and other GPUs to be exchanged, and performing data exchange.
6. The method of claim 1, in the step S3, the state vector of each GPU is initialized by the following way: Based on the split qubits of the first quantum gate operation layer, the split qubit state mapped on each GPU is initialized in the following way: wherein, a sequence number of the GPU, a sequence number of the GPU, a partitioned qubit state mapped on the GPU with the sequence number the value of the sequence number is the number of partitioned qubits minus 1, a binary number string formed by splicing the state of each partitioned qubit, wherein the sequence number of the GPU is equal to the decimal number value corresponding to the binary number string formed by splicing the state of the partitioned qubits mapped on the GPU, and the state of each partitioned qubit is 1 or 0. For each GPU, a part state vector in the initial state vector of the specified quantum system which is the same as the split qubit state mapped on the GPU is obtained, and a state vector of other qubits except the split qubits is extracted therefrom as the state vector of the GPU, so as to initialize the state vector of each GPU.
7. The method of claim 1, wherein, The preset data exchange method is configured to: Obtain the part of the state vector of the current GPU that needs to perform data exchange and the target GPU that performs data exchange, and obtain the part of the state vector and transmit it to the target GPU; Transmit the part of the state vector of the target GPU that needs to perform data exchange to the current GPU, and store the state vector transmitted by the current GPU; Update the quantum state vector by performing quantum gate operation on the current GPU, extract the part of the state vector transmitted by the target GPU from the updated quantum state vector and return it to the target GPU, and return the state vector of the current GPU stored in the target GPU to the current GPU.
8. A quantum circuit simulation apparatus for implementing the method of any one of claims 1 to 7, for simulating a running process of a quantum circuit on a specified quantum system to obtain a final state vector of the quantum system, characterized in that, The device comprises: A communication module for performing data exchange operation between GPUs based on the communication channel constructed by the NCCL library; a quantum system state vector partitioning module configured to select, for a layer of quantum gate operations, from all qubits of a quantum circuit to be simulated a partitioned qubit for partitioning an initial state vector of the quantum system having a minimum partitioning cost, to construct a dynamic partitioned qubit schedule; A quantum circuit simulation module for constructing a quantum circuit simulation operation sequence including all GPU data exchange operations and all multi-GPU supported quantum gate operations, initializing the state vector on the GPU, and executing the quantum circuit simulation operation sequence in parallel on all GPUs to update the state vector on each GPU; A state vector integration module for collecting the updated state vectors in all GPUs and sorting them based on the order of the quantum bits in the quantum circuit to be simulated to obtain the final state vector of the quantum device.
9. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of the method of any one of claims 1 to 7.
10. An electronic device, comprising: Comprise: One or more processors; And A memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method of any one of claims 1 to 7 by executing the executable instructions.
Citation Information
Patent Citations
Super-computing-oriented quantum search simulation method and system
CN116227615A
Systems and methods for optimizing quantum circuit simulation using graphics processing units
WO2023177846A1