A large-scale quantum circuit simulation method based on the cooperation of CPU and GPU
By using CPU and GPU collaborative methods in quantum circuit simulation, the memory access efficiency is optimized, and the memory access delay problem of quantum circuit simulation software in the prior art is solved, achieving significant computing speed improvement and acceleration effect.
Patent Information
- Application Number
- CN202210232271.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-03-09
AI Technical Summary
Existing quantum circuit simulation software has memory access delay problem, especially in the simulation calculation of large-scale quantum algorithms, resulting in limited computing speed.
A large-scale quantum circuit simulation method based on CPU and GPU collaboration is adopted to reduce memory access latency by optimizing time locality and spatial locality. Specific measures include storing qubit states with full amplitude, storing state arrays with paging, remapping qubit indexes, and dividing quantum circuits into independent sub-circuits to improve computing efficiency.
The speed of quantum circuit simulation has been significantly improved, 455 times acceleration of 34 qubits and 120 times acceleration of 32 qubits has been achieved, solving the problems of memory access delay and memory capacity limitation in the prior art.
Smart Images

Figure CN114595820B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of quantum computing, high-performance computing, etc., and particularly relates to a method for large-scale quantum circuit simulation based on the cooperation of a CPU and a GPU. Background Art
[0002] With the development of quantum computers and quantum algorithms, quantum computing has become the most likely means to surpass classical computing in the future. However, due to reasons such as poor fidelity of quantum computers, high R & D and usage costs, it is very inconvenient to verify quantum algorithms on them. Therefore, using a classical computer to simulate the operation of a quantum circuit has become an effective means. Quantum simulation methods are mainly divided into two categories, including the state vector (full amplitude) method and the tensor network method: the former is limited by the number of qubits, and its memory occupancy grows exponentially with the increase of qubits, doubling the memory occupancy for each additional qubit; the latter can bypass the exponential wall limit by using tensor contraction and decomposition, but has poor performance for deeper circuits and circuits with higher entanglement.
[0003] QuEST is implemented based on the state vector method. It uses an array of state vectors to store the complete state vector and supports single qubit gates with multiple control qubits. Among them, the application of qubit gates is carried out by sequentially applying the corresponding small operators to the entire state vector. Although QuEST supports GPU acceleration, its implementation has an obvious memory access bottleneck and there is a large room for optimization; in addition, the limitation of only supporting a single GPU also restricts the scalability of the problem size of GPU acceleration of QuEST: QuEST stores the state vector in the video memory to avoid the copy overhead between the main memory and the video memory, so when the number of qubits exceeds 31, the required memory is at least 64GB, exceeding the video memory capacity of the vast majority of GPUs, and QuEST will not be able to run. Such large-scale quantum algorithms still cannot be accelerated by using high-performance GPU hardware for simulation calculation by QuEST. Summary of the Invention
[0004] The technical problem solved by the present invention: Overcoming the deficiencies of the prior art, providing a method for large-scale quantum circuit simulation based on the cooperation of a CPU and a GPU, solving the memory access latency problem of existing quantum circuit simulation software, effectively optimizing memory access from the perspectives of temporal locality and spatial locality, and significantly improving the speed of quantum circuit simulation.
[0005] The technical solution of the present invention, a method for large-scale quantum circuit simulation based on the cooperation of a CPU and a GPU, includes the following steps:
[0006] Step 1: Start the quantum circuit simulation process and read the input of the quantum circuit to be calculated. The input includes the number of qubits \(n\), the initial state \(\Psi\) of the qubits, the set of quantum gates \(G\), and the quantum gate layout diagram \(E\).
[0007] Step 2: Call the memory manager to store the initial state \(\Psi\) of the qubits in the state array \(s\) in the way of full amplitude. The specific method is as follows:
[0008] (1) Represent the state of the qubit in the form of probability. If \(n = 1\), that is, the input circuit contains only one qubit \(q\), and its value set is \(\{0, 1\}\), and both values exist in the form of probability, satisfying:
[0009] p q = 0 + p q = 1 = 1
[0010] where \(p\) q=0 is the probability that the state of this qubit is 0, and \(p\) q=1 is the probability that the state of this qubit is 1.
[0011] (2) If \(n>1\), for the \(n\)-qubit \(q\) n , its state \(\Psi\) n has a value set containing \(2\) n elements, which are the \(n\)-bit binary representations from 0 to \(2\) n-1 , that is, \(\{00…000, 00…001, 00…010, ……, 11…111\}\). For each value \(q\) 0 q 1 …q n-2 q n-1 , the value of the \(i\)-th bit \(q\) i corresponds to the value of the \(i\)-th qubit; in the given qubit state, each permutation combination of values has a probability of occurrence, and the sum of the probabilities of all possible states also satisfies the equation:
[0012] p 00..000 + p 00..001 + …… + p 11…111 = 1
[0013] (3) The full amplitude method is to accurately store the probability of each value combination occurrence and store it in the array \(s\). This state array contains \(2\) n elements, which are the probabilities of \(2\) n states occurring.
[0014] Step 3: Call the memory manager to store the state array \(s\) in pages, using pages as the basic unit for subsequent data migration. For the index \(q\) of the \(n\)-qubit 0 q 1 …q n-2 qn-1 , set q n-2k-1 q n-2k …q n-1 The corresponding 2k bits are the page number, and the remaining (n - 2k) bits are the offset within the page. The page number is always even and evenly divided into two parts. The page number q n-k-1 q n-k …q n-1 The corresponding k bits are the task numbers in the task division. The lower k bits and the (n - 2k)-bit offset within the page together form an (n - k)-bit GPU video memory index, and its value range is the size of the memory block calculated by a GPU task. The index schematic diagram is as Figure 3 shown.
[0015] Step 4: Use the quantum gates in the set G of quantum gates as nodes and the directed dependencies between the quantum gates recorded in the quantum gate layout diagram E as edges to establish a directed acyclic graph DAG; let the set of control qubits of the quantum gate corresponding to node P be C P , and the set of target qubits be T P , then for the quantum gates A and B that appear successively in the original circuit, there is a directed edge from A to B if and only if for any qubit x that satisfies the following conditions:
[0016] x ∈ (C A ∪T A ) ∩ (C B ∪T B )
[0017] For any quantum gate C between the quantum gates A and B in the original circuit, it satisfies:
[0018]
[0019] Step 5: Based on the directed acyclic graph established in Step 4, divide the directed acyclic graph into multiple mutually independent sub-circuits S i , forming a set S. The set of quantum gates in S i is denoted as G i ; assume that the set of quantum gates without predecessors in the DAG is H, then the steps for dividing the sub-circuits are:
[0020] (1) If the set is non-empty, take out a quantum gate H from the set H i , and go to step (2), otherwise go to step 6;
[0021] (2) Using H i as the initial node, establish a sub-circuit S i , and go to step (3);
[0022] (3) Denote the number of qubits in the sub-circuit S i as k, Among them, C Gi represents the set of target qubits G i , and T Gi represents the set of control qubits G i . The video memory of the GPU is t bits; traverse all the sets of quantum gates G. If:
[0023] ① There exists a quantum gate G j such that for the set of leading nodes G' of G j in the DAG, for the following is satisfied:
[0024] G′ i ∈S i
[0025] ② For the quantum gate G in ① j , let k' = |∪ Gi∈S {(C Gi ∪T Gi )∪(C Gj ∪T Gj )}|, and it satisfies 2 k '<t;
[0026] If ① and ② are satisfied, then let k = k' and go to step (4); otherwise, go to step (5);
[0027] (4) Add the quantum gate G j to the sub-circuit S i , regard the quantum gate G j as G i , and go to step (3);
[0028] (5) For all the quantum gates g in the sub-circuit S i , that is, delete the quantum gate g and the edges with g as the endpoints from the DAG graph, re-traverse the DAG, set the set H as the set of quantum gates without leading nodes currently, add the sub-circuit S i to the set of sub-circuits S, and return to step (1).
[0029] Step 6: Call the executor. For the set of sub-circuits S divided, take out one sub-circuit S i each time and go to step 7;
[0030] Step 7: Traverse the sub-circuit S i Denote the set of qubits without data dependence in the sub-circuit S i as Q. A qubit q ∈ Q if and only if for any gate g in the set of sub-circuit quantum gates G i , if it acts on the qubit q, it only acts on the qubit q, that is:
[0031]
[0032] Step 8: If the set Q of qubits is an empty set, go to Step 11; otherwise, take out each qubit q from the set Q of qubits one by one, denote the index of q as index(q), if index(q) < n - 2k, that is, the position of q in the index corresponds to the in-page offset, go to Step 9, otherwise go to Step 10;
[0033] Step 9: Traverse the [n - 2k, n - k - 1] bits of the qubit index. If there exists an index x such that the qubit corresponding to x then remap q x and the index of q. The rule is: in each state established in Step 2:
[0034] (1) If there is no qubit q with data dependency x = the taken-out qubit q, remain unchanged;
[0035] (2) If q x ≠ q, then swap s[q 0 q 1 …q x …q…q n-2 q n-1 and s[q 0 q 1 …(1 - q x )…(1 - q)…q n- 2 q n-1 , that is, swap the page number positions of q and q x in the paged storage established in Step 3. As Figure 3 shown, in the figure, the number of qubits n = 15, and the number of page numbers 2k = 8. The light-colored squares in [0, 7] in the first row correspond to q x , and after the data exchange, it is stored in the [8, 15] bits, as shown in the second row.
[0036] If there is no x or the exchange is completed, put the taken-out qubit q back into the set Q of qubits, and return to Step 8.
[0037] Step 10: Traverse the [n - k, n - 1] bits of the qubit index. If there exists an index x such that the qubit corresponding to x then remap q x and the index of q. The rule is: in the page table established in Step 2:
[0038] (1) If q x = q, remain unchanged;
[0039] (2) If qx If it is not equal to q, then swap the entry q in the page table 0 q 1 …q x …q…q 2k-2 q 2k-1 and q 0 q 1 …(1 - q x )…(1 - q)…q 2k-2 q 2k-1 the address recorded in. As Figure 3 shown in the 2nd and 3rd lines, the light-colored squares in [8, 15] correspond to q x , and after data exchange, it is stored in the high memory of [12, 15], as shown in the 4th line.
[0040] If x does not exist or the exchange is completed, then put q back into the set Q and go to step 8;
[0041] Step 11: On a cluster containing multiple computing nodes, simulate the execution of a quantum circuit. For a sub-circuit containing s qubits, each quantum gate is represented as a square matrix A of 2 s *2 s , the state of s qubits at time n is represented as a vector Ψ s of length 2 s (n), and the state of s qubits at time n + 1 is represented as a vector Ψ s of length 2 s (n + 1). Then the action of each quantum gate is represented as:
[0042] Ψ s (n + 1) = A · Ψ s (n)
[0043] During the calculation process, task partitioning is required. The calculation tasks are numbered with the high k bits as the task numbers and evenly sent to the available GPUs, and then go to step 12;
[0044] Step 12: When data needs to be sent between computing nodes, the sender executes the sending task, and the receiver executes the receiving, loading, and running tasks. For each computing node, start a thread to handle the sending task separately; after the sending task is completed, close the thread, and the receiving, loading, and computing tasks are executed using the original thread. After completion, check whether all sub-circuits have been calculated. If so, go to step 13; otherwise, return to step 7;
[0045] Step 13: After calculating all the sub-circuits in the sub-circuit set S in step 6, output the final qubit state Ψn: Ψn = [p 00..000 , p 00..001 ,......, p 11..111 ;
[0046] where p 00…000 is the probability that the state of the qubit is 00…000 after all quantum gates act on it. is the probability that the state of the qubit is a 0 a 1 …a n-1 after all quantum gates act on it.
[0047] The advantages of the present invention compared with the prior art are as follows: Based on the underlying mechanism of quantum computing, the present invention re-plans the sorting of quantum gates, improving the temporal locality; remaps the storage of qubits, and the qubit information to be calculated each time is continuously stored in memory, improving the spatial locality and significantly enhancing the memory access efficiency.
[0048] (1) The present invention designs a structured quantum circuit simulation method. Each link of the quantum circuit simulation is encapsulated, reflecting the idea of high cohesion and low coupling, which is convenient for subsequent sub-module optimization. Compared with the CPU implementation of QuEST, the speedup ratio of this method on 34 qubits is 455 times, and compared with Qibo on one node, the speedup ratio of this method on 32 qubits is 120 times.
[0049] (2) The present invention uses the full amplitude method to store the state information of qubits in an array. This method realizes the continuous storage of qubit information, which is convenient for searching.
[0050] (3) The present invention stores the state array of qubits in pages, implementing a memory management mechanism based on page tables. Since the page number mapping of the page table is a logical mapping, when the mapping method of the qubit changes and involves the index bit where the page number is located, only the mapping of the page table needs to be modified to complete the change, without exchanging the stored data.
[0051] (4) The present invention converts the arrangement information of quantum gates into a storage method of a directed acyclic graph. Compared with the traditional linear storage, the directed acyclic graph can analyze the dependence relationship between quantum gates in the circuit, providing a basis for the division of the sub-circuit in claim 5.
[0052] (5) The present invention divides the original quantum circuit into multiple independent sub-circuits. Compared with the original circuit, the divided sub-circuits have better temporal locality.
[0053] (6) The present invention exchanges the qubits on the same page in the page table, improving the spatial locality.
[0054] (7) The present invention performs remapping between pages. Because only the page table needs to be modified instead of migrating data during page remapping, the mapping can be efficiently performed, improving the spatial locality. Description of the Drawings
[0055] Figure 1 is the overall flowchart of the method of the present invention;
[0056] Figure 2 is the overall structure diagram of the method of the present invention;
[0057] Figure 3 is the state index and remapping schematic diagram of the memory manager in the present invention. Detailed Embodiments
[0058] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0059] The implementation process steps of the present invention are as shown in Figure 1 and the visual composition of each module is as shown in Figure 2 . Among them, steps 1, 2, and 3 correspond to the circuit input, steps 4 and 5 correspond to the gate aggregator and task allocator on the left in Figure 2 , steps 6-11 correspond to the executor on the right in Figure 2 , and step 12 corresponds to the communicator. The specific execution process is as follows:
[0060] Step 1: Start the quantum circuit simulation process and read the input of the quantum circuit to be calculated. The input includes the number of qubits n, the initial state Ψ of the qubits, the set of quantum gates G, and the quantum gate layout diagram E;
[0061] Step 2: Call the memory manager and store the initial state Ψ of the qubits in the state array s in the form of full amplitudes; the specific method is as follows:
[0062] (1) Represent the state of the qubits in the form of probability. If n = 1, that is, the input circuit only contains one qubit q, and its value set is {0, 1}, and both values exist in the form of probability, satisfying:
[0063] p q = 0 + pq = 1 = 1
[0064] where p q=0 is the probability that the state of this qubit is 0, and p q=1 is the probability that the state of this qubit is 1
[0065] (2) If n>1, for n qubits q n , its state Ψn The value set of n contains 2 n-1 elements, which are the n-bit binary representations of numbers from 0 to 2 0 , i.e., {00…000, 00…001, 00…010, ……, 11…111}. For each value q 1 …q n-2 q n-1 , the value q of the i-th bit i corresponds to the value of the i-th qubit; in a given qubit state, each permutation combination of values has a probability of occurrence, and the sum of the probabilities of all possible states also satisfies the equation:
[0066] p 00..000 + p 00..001 + …… + p 11…111 = 1
[0067] (3) The full amplitude method is to accurately store the probability of occurrence of each value combination and store it in the array s. This state array contains 2 n elements, which are the probabilities of 2 n states occurring;
[0068] Step 3: Call the memory manager to store the state array s in pages, using pages as the basic unit for subsequent data migration. For the index q of the n-bit qubit 0 q 1 …q n-2 q n-1 , set the 2k bits corresponding to q n-2k-1 q n-2k …q n-1 as the page number, and the remaining (n - 2k) bits as the offset within the page. Among them, the page number is always even and is evenly divided into two parts. The k bits corresponding to the page number q n-k-1 q n-k …q n-1 are the task numbers in the task division, and the lower k bits and the (n - 2k)-bit offset within the page together form the GPU video memory index with a value range that is the size of the memory block calculated by a GPU task. The index schematic diagram is as Figure 3 shown.
[0069] Step 4: Use the quantum gates in the quantum gate set G as nodes and the directed dependencies between the quantum gates recorded in the quantum gate layout diagram E as edges to establish a directed acyclic graph DAG; let the set of control qubits corresponding to the node P be C P , and the set of target qubits be T P . Then, for the quantum gates A and B that appear successively in the original circuit, there is a directed edge from A to B if and only if for any qubit x that satisfies the following conditions:
[0070] x ∈ (C A ∪ T A ) ∩ (G B ∪ T B )
[0071] For any quantum gate C between quantum gates A and B in the original circuit, it satisfies that:
[0072]
[0073] Step 5: Based on the established directed acyclic graph in Step 4, divide the directed acyclic graph into multiple mutually independent sub - circuits S i , forming the set S. The set of quantum gates in S i is denoted as G i ; Assume that the set of quantum gates without predecessors in the DAG is H, then the steps to divide the sub - circuit are as follows:
[0074] (1) If the set is non - empty, take out a quantum gate H i from the set H and go to step (2), otherwise go to step 6;
[0075] (2) Taking H i as the initial node, establish the sub - circuit S i and go to step (3);
[0076] (3) Denote the number of qubits in the sub - circuit S i as k, where C Gi represents the set of target qubits of G i , T Gi represents the set of control qubits of G i , and the video memory of the GPU is t bit; Traverse all sets of quantum gates G. If:
[0077] ① There exists a quantum gate G j , for the set of predecessor nodes G' of G j in the DAG, such that for it satisfies that:
[0078] G′ i ∈ S i
[0079] ② For the quantum gate G j in ①, let k' = |∪ Gi∈S {(C Gi ∪ T Gi ) ∪ (C Gj ∪ T Gj )}|, and it satisfies 2 k '< t;
[0080] If ①② are satisfied, let k = k', and go to step (4); otherwise, go to step (5).
[0081] (4) Add the quantum gate G j to the sub-circuit S i , regard the quantum gate G j as G i , and go to step (3).
[0082] (5) For all the quantum gates g in the sub-circuit S i , that is delete the quantum gate g and the edge with g as the end point from the DAG graph, re-traverse the DAG, set the set H as the set of quantum gates without predecessors currently, add the sub-circuit S i to the sub-circuit set S, and return to step (1).
[0083] Step 6: Call the executor. For the sub-circuit set S divided, take out one sub-circuit S i each time, and go to step 7;
[0084] Step 7: Traverse the sub-circuit S i Denote the set of qubits without data dependence in the sub-circuit S i as Q. A qubit q ∈ Q if and only if for any gate g in the sub-circuit quantum gate set G i , if it acts on the qubit q, it only acts on the qubit q, that is:
[0085]
[0086] Step 8: If the qubit set Q is an empty set, go to step 11; otherwise, take out the qubits q from the qubit set Q one by one, denote the index of q as index(q). If index(q) < n - 2k, that is, the position of q in the index corresponds to the in-page offset, go to step 9; otherwise, go to step 10;
[0087] Step 9: Traverse the [n - 2k, n - k - 1] bits of the qubit index. If there exists an index x such that the qubit corresponding to x then remap q x and the index of q. The rule is that in each state established in step 2:
[0088] (1) If the qubit q without data dependence x = the taken-out qubit q, remain unchanged;
[0089] (2) If q x ≠ q, then swap s[q 0 q 1 …qx …q…q n-2 q n-1 q with s 0 q 1 …(1 - q x )…(1 - q)…q n- 2 q n-1 the value of, that is, swap q and q x the page number position in the paged storage established in step 3. As Figure 3 shown, in the figure, the number of qubits n = 15, and the number of page numbers 2k = 8. The light - colored squares in [0, 7] in the first row correspond to q x , and after data exchange, it is stored in the [8, 15] position, as shown in the second row.
[0090] If x does not exist or the exchange is completed, then put the retrieved qubit q back into the qubit set Q, and return to step 8.
[0091] Step 10: Traverse the [n - k, n - 1] bits of the qubit index. If there exists an index x such that the qubit corresponding to x then remap q x and the index of q. The rule is: in the page table established in step 2:
[0092] (1) If q x = q, remain unchanged;
[0093] (2) If q x != q, then swap the entries q 0 q 1 …q x …q…q 2k-2 q 2k-1 and q 0 q 1 …(1 - q x )…(1 - q)…q 2k-2 q 2k-1 recorded addresses in. As Figure 3 shown in the second and third rows, the light - colored squares in [8, 15] correspond to q x , and after data exchange, it is stored in the high - memory [12, 15], as shown in the fourth row.
[0094] If x does not exist or the exchange is completed, then put q back into the set Q, and enter step 8;
[0095] Step 11: On a cluster containing multiple computing nodes, simulate the execution of the quantum circuit. For a sub - circuit containing s qubits, where each quantum gate is represented as 2 s *2 sFor the square matrix A, the state of s qubits at time n is represented by a vector Ψ of length 2 s and the state of s qubits at time n+1 is represented by a vector Ψ of length 2 s (n), then the action of each quantum gate is represented as: s Ψ s (n+1) = A·Ψ
[0096] Ψ s (n+1) = A·Ψ s (n)
[0097] During the calculation process, it is necessary to divide the tasks. The calculation tasks are numbered with the high k bits as the task numbers and evenly sent to the available GPUs, and then enter step 12;
[0098] Step 12: When data needs to be sent between computing nodes, the sender executes the sending task, and the receiver executes the receiving, loading, and running tasks. For each computing node, a thread is started to handle the sending task separately; after the sending task is completed, the thread is closed, and the receiving, loading, and computing tasks are executed using the original thread. After completion, check whether all sub-circuits have been calculated. If so, enter step 13; otherwise, return to step 7;
[0099] Step 13: After calculating all the sub-circuits in the sub-circuit set S in step 6, output the final qubit state Ψn: Ψn = [p 00..000 , p 00..001 ,......, p 11..111 ;
[0100] where p 00…000 is the probability that the qubit state is 00…000 after the action of all quantum gates. is the probability that the qubit state is a 0 a 1 …a n-1 after the action of all quantum gates.
Claims
1. A large-scale quantum circuit simulation method based on the cooperation of CPU and GPU, characterized in that, it includes the following steps: Step 1: Start the quantum circuit simulation process, read the input of the quantum circuit to be calculated, and the input includes the number of qubits n, the initial state Ψ of the qubits, the set of quantum gates G, and the quantum gate layout diagram E; Step 2: Call the memory manager, store the initial state Ψ of the qubits in the form of full amplitudes, and store it in the state array s; Step 3: Call the memory manager, store the state array s in pages, and use the page as the basic unit for subsequent data migration; Step 4: Use the quantum gates in the set of quantum gates G as nodes and the directed dependencies between the quantum gates recorded in the quantum gate layout diagram E as edges to establish a directed acyclic graph DAG; Step 5: Based on the completed directed acyclic graph (DAG) established in Step 4, divide the DAG into multiple mutually independent sub-circuits S i , forming a set of sub-circuits S, where S i the set of quantum gates of the sub-circuits is denoted as G i ; Step 6: Invoke the actuator. For the set S of sub-circuits obtained by partitioning, take out one sub-circuit S each time and proceed to Step 7; i Step 7: Traverse the sub-circuit S i The set of qubits without data dependencies in is Q. For a qubit i q, if and only if for any gate g in the set of sub-circuit quantum gates G Step 8: If the set of qubits Q is an empty set, go to Step 11; otherwise, take out the qubits q one by one from the set of qubits Q, record the index of q as index(q), if index(q) < n - 2k, that is, the position of q in the index corresponds to the in-page offset, go to Step 9, otherwise go to Step 10; Step 9: Traverse the [n - 2k, n - k - 1] bits of the qubit index. If there exists an index x such that the qubit corresponding to this index x , then remap the indices of q x and q, and enter Step 8; Step 10: Traverse the [n-k,n-1] bits of the qubit index. If there exists an index x such that the qubit corresponding to index x , then remap the indices of q x and q, and enter Step 8; Step 11: Simulate the quantum circuit on a cluster containing multiple computing nodes. For a sub-circuit containing s qubits, each quantum gate is represented as a square matrix A of size 2 s 2 s The state of s qubits at time n is represented as a vector Ψ s of length 2 s (n), and the state of s qubits at time n+1 is represented as a vector Ψ s of length 2 s (n+1). Then the action of each quantum gate is represented as: During the calculation process, task division is required. The calculation tasks are numbered with the high k bits as the task number and evenly sent to the available GPUs, and then go to Step 12; Step 12: When data needs to be sent between computing nodes, the sender executes the sending task, and the receiver executes the receiving, loading, and running tasks. For each computing node, start a thread to handle the sending task separately; After the sending task is completed, close the thread, and the receiving, loading, and computing tasks are executed using the original thread. After completion, check whether all sub-circuits have been calculated. If so, go to Step 13, otherwise return to Step 7; Step 13: After calculating all the sub-circuits in the sub-circuit set S of Step 6, output the final qubit state Ψ n : where is the probability that the state of the qubit is 00…000 after the action of all quantum gates, is the probability that the state of the qubit is after the action of all quantum gates; In Step 9, the remapping method is as follows: (1) If the qubit q with no data dependence x = the retrieved qubit q, remains unchanged; (2) If q x ≠ q, then swap s[q 0 q 1 …q x …q…q n-2 q n-1 and s[q 0 q 1 …(1 - q x )…(1 - q)…q n-2 q n-1 , that is, swap the page numbers of q and q x in the memory manager; If there is no x or the swap is completed, put the taken-out qubit q back into the set of qubits Q and return to Step 8.
2. The large-scale quantum circuit simulation method based on the cooperation of CPU and GPU according to Claim 1, characterized in that: In Step 2, the method for storing the initial state Ψ of the qubits in the form of full amplitudes is as follows: (1) Represent the state of the qubits in the form of probability. If n = 1, that is, the input circuit only contains one qubit q, and its value set is {0, 1}, and both values exist in the form of probability, satisfying: where is the probability that the qubit state is 0, is the probability that the qubit state is 1; (2) If n > 1, for an n-bit qubit , its state has a value set containing 2 n elements, which are the n-bit binary representations from 0 to 2 n-1 , i.e., {00…000, 00…001, 00…010, ……, 11…111}. For each such value q 0 q 1 …q n-2 q n-1 , the value q i of the i-th bit corresponds to the value of the i-th qubit; in a given qubit state, each permutation combination of values has a probability of occurrence, and the sum of the probabilities of all possible states also satisfies the equation: (3) The full amplitude method is to accurately store the probability of each combination of values and store it in the state array s. This state array s contains 2 n elements, which are the probabilities of 2 n states occurring.
3. The large-scale quantum circuit simulation method based on the cooperation of CPU and GPU according to Claim 1, characterized in that: In the said step 3, the method for storing the state array s in pages is as follows: for the index q of n qubits 0 q 1 …q n-2 q n-1 , set the corresponding 2k bits as the page number, and the remaining (n - 2k) bits as the offset within the page. Among them, the page number is always an even number and is evenly divided into two parts. The corresponding k bits are the task numbers in the task division. The lower k bits and the (n - 2k)-bit offset within the page together form an (n - k)-bit GPU video memory index, and its value range is the memory block size calculated by a GPU task.
4. The large-scale quantum circuit simulation method based on the cooperation of CPU and GPU according to Claim 1, characterized in that: In Step 4, the method for establishing the directed acyclic graph DAG is as follows: Let the set of control qubits of the quantum gate corresponding to node P be , and the set of target qubits be . Then, for quantum gates A and B that appear successively in the original circuit, there is a directed edge from quantum gate A to quantum gate B if and only if for any qubit x that satisfies the following conditions: For any quantum gate C between quantum gates A and B in the original circuit, it satisfies: 。 5. The large-scale quantum circuit simulation method based on the cooperation of CPU and GPU according to Claim 1, characterized in that: In step 5, it is divided into multiple independent sub-circuits S i The method is as follows: (1) If the set is non-empty, take a quantum gate H from the set H i , and proceed to step (2); otherwise, proceed to step 6. (2) Using H i as the initial node, establish the sub-circuit S i , and proceed to step (3); (3) Denote the number of qubits in the sub - circuit S i as k, , where represents G i the set of target qubits, represents G i the set of control qubits, and the video memory of the GPU is t bit; Traverse all the sets of quantum gates G. If: ① There exists a quantum gate G j , for G j the set of leading nodes G' in the DAG such that for , it satisfies: ② For the quantum gate G in ① j, Let satisfy ; If ①② are satisfied, then let , go to step (4); otherwise, go to step (5). (4) Add the quantum gate G j into the sub-circuit S i Consider the quantum gate G j as G i, and go to step (3); (5) For all quantum gates g in the sub - circuit S i , that is , delete the quantum gate g and the edges with g as endpoints from the DAG. Re - traverse the DAG, set the set H as the set of quantum gates that currently have no predecessors, add the sub - circuit S i to the sub - circuit set S, and return to step (1).
6. A large-scale quantum circuit simulation method based on the cooperation of CPU and GPU according to Claim 1, characterized in that: In the said step 10, the method for remapping is as follows: in the memory manager in step 3: (1) If q x = q, remain unchanged; (2) If q x ≠ q, then swap the item q 0 q 1 …q x …q…q 2k-2 q 2k-1 with q 0 q 1 …(1 - q x )…(1 - q)…q 2k- 2 q 2k-1 in the recorded address; If x does not exist or the swap is completed, then put q back into the set Q and return to step 8.