An Efficient Quantum Circuit Simulation Method Based on Distributed Systems

By classifying state vector sharding and qubits in a distributed system, parallel processing and last step merge operations, the problem of frequent node communication in distributed quantum line simulation is solved, and efficient and scalable quantum line simulation is achieved.

CN117291271BActive Publication Date: 2025-06-10EAST CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311230276.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-22
Publication Date
2025-06-10
Estimated Expiration
2043-09-22

AI Technical Summary

Technical Problem

The existing distributed quantum line simulation technology requires frequent inter-node communication during each layer of line simulation, resulting in large communication overhead, low efficiency, and difficulty in scaling.

Method used

By slicing the state vectors evenly to each distributed node and dividing the qubit into high-bit qubits and low-bit qubits according to the number of nodes, parallel processing and last-step merge operation are adopted to avoid inter-node communication simulated by each layer of line.

Benefits of technology

It significantly reduces the communication overhead in distributed quantum line simulation, improves simulation efficiency, and achieves good scalability. The simulation time accelerates linearly with the increase of the number of nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117291271B_ABST
    Figure CN117291271B_ABST
Patent Text Reader

Abstract

The present invention discloses an efficient quantum circuit simulation method based on a distributed system, which includes four steps in total. Step 1, construct an initial state vector according to the quantum circuit, and evenly slice the state vector and place it on each distributed node respectively; Step 2, divide the qubits into two categories: high-order qubits and low-order qubits according to the number of nodes; Step 3, on each node, perform the processing on the high-order qubits and the operations on the low-order qubits in parallel. Perform multiplication operations between the unitary matrices composed of the quantum gates on the high-order part of each layer of the quantum circuit, and directly update each slice of the state vector for the quantum gates on the low-order part; Step 4, divide the slices of the state vector on each node into multiple blocks, and realize the final merging operation through block reorganization, calculation, and block reallocation among the distributed nodes, and truly integrate the quantum gate operations on the high-order part into the state vector update. The present invention can greatly reduce the communication overhead and significantly accelerate the quantum circuit simulation in the distributed system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of quantum computing, and particularly relates to an efficient quantum circuit simulation method based on a distributed system. Background Art

[0002] Quantum computing, as an emerging field, has demonstrated its quantum supremacy in many applications. Due to the properties of quantum superposition and entanglement, quantum computers exhibit powerful computing potential. However, accessing real physical quantum computers is extremely expensive, and the number of accessible qubits they provide is also limited. Therefore, the true machine verification of quantum algorithms is currently in an immature stage of development. To better understand quantum behavior and verify quantum algorithms, quantum simulators provide an important way to promote the development of quantum computers and quantum algorithms. So far, various types of quantum simulators have emerged to trade off between runtime and space. Among these simulators, full-state quantum circuit simulation has become very important because it is more suitable for deeper quantum circuits and well supports software debugging. It generates all the probability amplitudes of the state vector during the simulation process.

[0003] Currently, quantum simulators generally perform simulations according to the quantum circuit model. Given a quantum circuit with n qubits, all possible states in the quantum circuit form a state vector, which contains 2 n probability amplitudes. For existing quantum algorithms, they apply quantum gates to the quantum circuit to achieve the transformation of the state vector, so as to obtain the desired probability amplitude distribution, and then process the final state vector to obtain the quantum algorithm result. For a quantum circuit, it contains multiple layers of circuits, and there is a series of quantum gates on each layer of the circuit. Full-state quantum circuit simulation performs quantum gate operations layer by layer from front to back. The quantum gates on each layer obtain the operation matrix of this layer through the tensor product operation, and the size of the operation matrix is 2 n ×2 n . Then, the operation matrix and the state vector of the previous layer are used to perform matrix-vector multiplication to obtain the state vector result of this layer. Assuming the initial state vector is The operation matrix obtained through the tensor product operation of the first layer is denoted as M 1 , then the state vector is converted to After that, after passing through T layers of quantum circuits, the state vector is converted to

[0004] In practical applications, 2 n ×2 nThe operation matrix of the size does not need to be fully calculated. A method called state vector simulation can transform the state vector in place using quantum gates on each qubit. For a single-qubit quantum gate, if it is applied to the k-th qubit, only the two probability amplitudes with a step distance of 2 k in the probability amplitudes are paired into probability amplitude pairs. These probability amplitude pairs are respectively multiplied by the 2×2 single-qubit quantum gate matrix-vector to update the corresponding probability amplitudes. That is, where G represents an arbitrary single-qubit quantum gate matrix, "*" ∈ {0, 1}, "*…*" means that the corresponding binary representations in the indices are the same, only the binary representation in the k-th position is different, and all 2 probability amplitude pairs need to be updated. For a controlled two-qubit quantum gate, which probability amplitude pairs need to be updated is determined by the position of the control qubit, and the step of the probability amplitude pairs to be updated is determined by the position of the target qubit. The specific calculation is, where c represents the position of the control qubit, t represents the position of the target qubit; only when the state of the control qubit is |1>, the corresponding probability amplitude will be updated; the step of the probability amplitude pair is 2 ( .

[0005] For the above state vector simulation method, it needs to store a state vector of size 2. As the depth of the simulated circuit increases, the required computing and storage resources will also increase significantly. Since a distributed system has multiple computing nodes and thus has a large amount of computing resources, it can well meet the resource requirements in quantum computing simulation. When the state vector is sharded and placed on each node, the quantum gate operations on some qubits can be locally computed on each node because the step of their probability amplitude pairs is less than the size of the state vector shard on each node. However, there are still some gate operations on qubits with higher weights in the probability amplitude index that cannot complete the state vector update only relying on the local probability amplitudes of each node because the step of their probability amplitude pairs has exceeded the size of the state vector shard. The quantum gates on these higher qubits must rely on inter-node communication to complete.

[0006] However, the communication overhead in a distributed system accounts for a large proportion. The communication overhead is about 80%-90% of the total overhead. For a multi-layer quantum circuit, multiple inter-node communications are required for each layer. This mode is very unfriendly to quantum circuits with a deeper number of layers. How to design a distributed quantum circuit simulation technology that can completely avoid the inter-node communication and synchronization overhead required for each layer of circuit simulation until the last merging operation is a huge challenge. At the same time, how to design the communication volume of the last merging operation to be similar to the data communication volume in the previous operation is also a key problem to be solved. Summary of the Invention

[0007] To overcome the problems existing in the above technologies, the object of the present invention is to provide an efficient quantum circuit simulation method based on a distributed system. First, the present invention evenly slices the state vector according to the number of nodes in the distributed system and places it on each node. At the same time, the qubits are divided into two categories: high-order qubits and low-order qubits according to the number of nodes. Then, the processing on the high-order qubits and the operations on the low-order qubits are executed in parallel on each distributed node. In the last step, the operation fusion is completed by dividing the blocks, realizing the efficient quantum circuit simulation on the distributed system. The present invention can significantly reduce the communication overhead in the distributed quantum circuit simulation, thereby greatly improving the quantum circuit simulation efficiency, and the distributed simulation shows good scalability.

[0008] The specific technical solution for achieving the object of the present invention is as follows:

[0009] An efficient quantum circuit simulation method based on a distributed system, the method comprising:

[0010] Step 1, data partitioning: Initialize the state vector of the circuit according to the initial state of the qubits, evenly divide the state vector into multiple slices, and place each slice on each distributed node;

[0011] Step 2, qubit classification: Divide the qubits in the circuit into two categories: high-order qubits and low-order qubits according to the number of distributed nodes;

[0012] Step 3, parallel processing of high and low order operations: On each node, execute the processing on the high-order qubits and the operations on the low-order qubits in parallel. Perform multiplication operations between the unitary matrices composed of the tensor product operations of the quantum gates on the high-order part of each layer of the quantum circuit, and directly update each state vector slice in place for the quantum gates on the low-order part;

[0013] Step 4, final merging operation: Divide the state vector slices on each node into multiple blocks, and through three steps of block recombination, calculation, and block reallocation between distributed nodes, truly integrate the quantum gate operations on the high-order part into the state vector update.

[0014] The step 1 of evenly dividing the state vector into multiple slices and placing each slice on each distributed node specifically includes:

[0015] Given a quantum circuit with n qubits, the size of its corresponding state vector is 2 n , denoted as N; assuming there are H nodes in the distributed system, then the state vector of size N will be evenly divided into H parts, each state vector slice has N / H probability amplitudes, and each probability amplitude is denoted as α s; The state vector is continuously divided, that is, the probability amplitude index s on each node is continuous; the state vector slices on H nodes are respectively denoted as V 0 , V 1 ,..., V H-1 .

[0016] Each of the state vector slices has N / H probability amplitudes, and each probability amplitude is denoted as α s , specifically including:

[0017] For the probability amplitude α s on each node, its index s consists of two parts, which are respectively denoted as s 1 、s 2 ; s 1 represents the number of the distributed node, ranging from 0 to H - 1; s 2 represents the probability amplitude number within the node, ranging from 0 to N / H - 1; therefore, s 1 on different nodes are different from each other, but s 2 all represent from 0 to N / H - 1.

[0018] As described in step 2, according to the number of distributed nodes, the qubits in the circuit are divided into two categories: high - order qubits and low - order qubits, specifically including:

[0019] Given a quantum circuit with n qubits, the qubits are respectively denoted as q n-1 , q n-2 ,..., q 1 , q 0 ; where q 0 represents the qubit with the lowest weight in the probability amplitude index, that is, s = q n-1 , q n-2 ,..., q 1 , q 0 ; After step 1, each node has N / H probability amplitudes, denoted as L = N / H; according to the state vector simulation method, the quantum gates on the log 2 L qubits with lower weights can directly update the state vector within each distributed node, denoted as l = log 2 L; q l-1 ,..., q 1 , q 0 These l qubits are called low - order qubits; for the remaining n - l qubits, the quantum gate operations on them must rely on the probability amplitudes on multiple nodes to complete the state vector update, so communication between distributed nodes is required, and these n - l qubits are called high - order qubits, denoted as h = n - l; that is, the quantum circuit has a total of n qubits q n-1 , q n-2 ,..., q1 , q 0 , is divided into h high - order qubits q n-1 ,..., q l+1 , q l and l low - order qubits q l-1 ,..., q 1 , q 0 ; where n = log 2 N, h = log 2 H, l = log 2 L.

[0020] Perform multiplication operations between the unitary matrices composed of the tensor product operations of the quantum gates on the high - order part of each layer of the quantum circuit described in step 3, specifically including:

[0021] Given a quantum circuit, which has T layers of circuits, and there are multiple quantum gate operations acting on the corresponding qubits on each layer; for the quantum gate operations on the high - order qubits of each layer of the circuit, each quantum gate can be represented as a small unitary matrix; perform tensor product operations between the unitary matrices of the high - order quantum gates to obtain a unitary matrix of size H×H; then, obtain the final unitary matrix through matrix multiplication operations on the H×H - sized unitary matrices on the T layers of the circuit; denote the H×H - sized unitary matrices on each layer as M high,1 , M high,2 ,..., M high,T , and obtain the final unitary matrix M high,T representing the operations on the high - order qubits through matrix multiplication M high,2 ·...·M high,1 ·M u .

[0022] The quantum gates on the low - order part described in step 3 directly update each state - vector slice in - place, specifically including:

[0023] According to the state - vector simulation method, when updating the state vector, the step size of the probability amplitude used for the gate operation on the k - th qubit is 2 k ; therefore, the maximum step size of the probability amplitude required for the gate operations on the l low - order qubits is 2 l-1 , and the continuous L probability amplitudes on one node can fully satisfy the gate operations on the l low - order qubits; for the T - layer quantum circuit, continuously perform the gate operations on the low - order qubits of each layer on each distributed node according to the state - vector simulation method, and the state vector is updated in - place.

[0024] On each node, perform the processing on the high - order qubits and the operations on the low - order qubits in parallel, specifically including:

[0025] For the processing on the above high - level qubits and the operations on the low - level qubits, these two operations are executed in parallel in a multi - threaded manner on each distributed node.

[0026] The step 4 of dividing the state vector shards on each node into multiple blocks specifically includes:

[0027] After data partitioning, each distributed node has L probability amplitudes; the L probability amplitudes are continuously divided into H blocks within the node, and each block has L / H probability amplitudes; the H blocks on the i - th node are respectively denoted as V i,0 , V i,1 ,..., V i,H-1 .

[0028] The step 4 of truly integrating the quantum gate operations on the high - level into the state vector update through three steps of block recombination, calculation, and block re - distribution among distributed nodes specifically includes:

[0029] The block recombination process recombines the j - th block on each distributed node to the j - th node for subsequent calculation, that is, the j - th block on the i - th node is sent to the i - th block on the j - th node; the calculation process performs a scalar multiplication operation of the vector shards of the obtained unitary matrix M u and the state vector shards on each distributed node after block recombination, denoted as the "⊙" operation; the block re - distribution process is to return the results calculated on each distributed node back to each node according to the block size to obtain the final correct result.

[0030] The scalar multiplication operation of the vector shards, denoted as the "⊙" operation, specifically includes:

[0031] This operation acts on an arbitrary H×H matrix M and a vector w, and the size of the vector w is a multiple of H; the operation is expressed as M⊙w = w′, where the vector w includes H shards, respectively denoted as w 0 , w 1 ,..., w H-1 , and H shards of the vector w′ are calculated, and each shard is respectively denoted as w′ 0 , w′ 1 ,..., w′ H-1 ; the specific operation process is where i is an integer within the range of [0, H - 1], m ij represents the element in the i - th row and j - th column of the matrix M; the ⊙ operation is to perform a scalar multiplication operation on the vector shards with each row of the matrix M, and add up the results of one - row operations to obtain one shard of the final result w′.

[0032] The method proposed by the present invention can significantly reduce the communication overhead between distributed nodes. The communication in each layer of quantum circuit calculation is completely avoided, and only one communication in the last merging operation is required. Moreover, the amount of communication data in the last merging operation is almost the same as that in the previous steps. This greatly improves the efficiency of distributed quantum circuit simulation, and with the increase in the number of distributed nodes, this simulation technology almost achieves linear simulation acceleration, showing good scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is an example of a 3-layer quantum circuit with 4 qubits, and the division of the state vector on 4 distributed nodes;

[0034] Figure 2 It is a schematic diagram of parallel processing of high and low bit operations;

[0035] Figure 3 It is a schematic diagram of the last merging operation on 4 distributed nodes;

[0036] Figure 4 It is a flow chart of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0037] The present invention will be further described below in conjunction with the drawings and embodiments:

[0038] Refer to Figure 4 , the present invention includes the following steps:

[0039] Step 1, data division: First, initialize the state vector according to the quantum circuit, and evenly divide the state vector into multiple vector shards, and each vector shard is placed on each node of the distributed system respectively. This division is a continuous division of the state vector, and the indexes of the probability amplitudes in each shard after division are continuous. Given a quantum circuit with n qubits, the size of the state vector is N = 2 n . There are H nodes in the distributed system, then the number of probability amplitudes in each node is N / H. The probability amplitudes on each node are denoted as α s , where s consists of s 1 , s 2 two parts. s 1 represents the node number in the distributed system, and s 2 represents the number within the node.

[0040] The reason why data partitioning works well is that distributed quantum circuit simulation is not suitable for task partitioning by dividing sub - circuits. When there are many control gates in the circuit and the circuit has a deep number of layers, task partitioning may increase the storage of state vectors and weaken the advantages of the distributed system. Through data partitioning, the number of stored state vectors does not increase, and it is not restricted by a specific quantum circuit. Continuous data storage also facilitates the effective implementation of the state vector simulation method.

[0041] Step 2, qubit classification: According to the number of distributed nodes, qubits are divided into two categories: high - order qubits and low - order qubits. Given a quantum circuit with n qubits, these n qubits will be divided into h high - order qubits and l low - order qubits. Among them, h = log 2 H, where H is the number of nodes in the distributed system; l = log 2 L, where L is the number of probability amplitudes within each node, and L = N / H. According to the state vector simulation method, the quantum gate operations on the h high - order qubits need to complete the state vector update through communication between nodes, while the quantum gate operations on the l low - order qubits can complete the state vector update in - place using the probability amplitudes within each node.

[0042] The reason why qubit classification works well is that by dividing qubits into two categories, the operations on high - order qubits and the operations on low - order qubits can be considered and discussed separately. For the gate operations on high - order qubits that need to be completed through communication, the communication operations required at each step can be considered and fused into one step; while the operations on low - order qubits can continue locally without hindrance and do not need to wait intermittently for the completion of the communication operations on high - order qubits.

[0043] Step 3, parallel processing of high - and low - order operations: By dividing high - order qubits and low - order qubits, the operations on their corresponding qubits are also considered and executed separately, and parallel acceleration is achieved through a multi - thread approach. By dividing the two types of qubits, the multi - layer gate operations on low - order qubits directly update the local state vector shards on each node. At the same time, the operation matrices on high - order qubits are multiplied together to obtain a unitary matrix M of size H×H u , which is used to incorporate the operations on high - order qubits into the state vector update in the last step.

[0044] The reason why the high - low bit operation parallel processing works well is as follows: First, the high - low bit operations are carried out separately, and by using the multi - thread method, the two can be executed simultaneously, improving the overall simulation efficiency of the circuit. Second, converting the operations on the high - order qubits into the consecutive multiplication operations of the operation matrices completely avoids the communication between nodes required in each layer of the previous quantum circuit simulation. The T times of communication are reduced to only 1 time of communication, and the communication overhead with a large proportion is greatly reduced, significantly accelerating the simulation speed of the quantum circuit with a relatively large number of layers.

[0045] Step 4, the last - step merging operation: The present invention proposes an efficient last - step merging technique. On each distributed node, the state vector is sliced into multiple blocks, and then the operations on the high - order qubits are integrated into the state vector update through three steps: block recombination, ⊙ operation, and block redistribution. First, for the L probability amplitudes on each node, they are divided into H blocks, and each block has L / H probability amplitudes. Then, in the block recombination process, the j - th block on each node is sent to the j - th node for organization. The j - th block from the i - th node is placed in the position of the i - th block on the j - th node. Then, the recombined state vector slice on each node is subjected to a ⊙ operation with the unitary matrix M u obtained by performing the operations on the high - order qubits to get the results of the corresponding blocks. The calculated block results are sent back to their original positions through the block redistribution process to obtain the state vector in the correct order.

[0046] The reason why the last - step merging operation works well is as follows: First, through block division and block recombination in the last step, the operations on each node no longer require all the probability amplitudes on other nodes, and only partial blocks on each node are needed to complete the calculation operations of the corresponding blocks, greatly reducing the storage requirements on each node. Second, through the three steps of block recombination, calculation, and block redistribution, the communication volume required for the entire quantum circuit simulation is only 2N - 2L, where N is the number of all probability amplitudes and L is the number of probability amplitudes on each node; while the traditional method requires node - to - node communication for each layer, with a communication volume of hNT. The last - step merging operation proposed by the present invention greatly reduces the total communication volume and significantly improves the simulation efficiency. Suppose n = 20, h = 5, T = 200, then N = 2 20 ,H = 2 5 ,L = 2 15 。For the traditional method, its communication volume is 1000×2 20 ;while the method proposed by the present invention only requires a communication volume of 31×2 16 ,achieving a reduction in communication volume of more than 2 9 times.

[0047] Embodiment

[0048] Data partitioning: As Figure 1(As shown in (a), in this embodiment, a 3-layer quantum circuit with 4 qubits is given. Therefore, there are a total of N = 2 4 = 16 probability amplitudes in the state vector. Among them, G in the quantum circuit refers to an arbitrarily given single-qubit quantum gate. There are 4 distributed nodes in the given distributed system. The state vector is continuously sliced and placed on 4 nodes. Then the corresponding probability amplitude distribution is as shown in Figure 1 (b). The probability amplitudes α 0000 ~α 0011 are placed on node 0, the probability amplitudes α 0100 ~α 0111 are placed on node 1, the probability amplitudes α 1000 ~α 1011 are placed on node 2, and the probability amplitudes α 1100 ~α 1111 are placed on node 3. Correspondingly, the probability amplitude indices on each node can also be represented in the form of the node number and the number within the node. That is, the probability amplitudes on node 0 can also be represented as α 0,0 ~α 0,3 , the probability amplitudes on node 1 can be represented as α 1,0 ~α 1,3 , the probability amplitudes on node 2 can be represented as α 2,0 ~α 2,3 , and the probability amplitudes on node 3 can be represented as α 3,0 ~α 3,3 .

[0049] Qubit classification: In this embodiment, there are 4 nodes in the given distributed system. Then the number of high-order qubits is h = log 2 4 = 2, and the number of low-order qubits is l = n - h = 4 - 2 = 2. Therefore, Figure 1 in (a), q 3 , q 2 are classified as high-order qubits, and q 1 , q 0 are classified as low-order qubits.

[0050] Parallel processing of high and low order operations: For a 3-layer quantum circuit, the operation matrix M u on its high-order qubits is obtained depending on M high,3 ×M high,2 ×M high,1 , such as the processing on the 2 high-order qubits of the 3-layer quantum circuit shown in Figure 2 . Among them, G represents an arbitrary 2×2 single-qubit quantum gate, and I represents that no gate is added to the second high-order qubit, which is a 2×2 identity matrix. It means that no quantum gates are applied to the two high-order qubits. In the present invention, if the high-order qubit serves as the control bit of the control gate, when calculating the unitary matrix on the high-order qubit, this bit is regarded as having no gates applied. This control qubit only functions when the low-order qubits are operated. Since no gates are applied to the two high-order qubits in the first layer of the quantum circuit, then Finally, through M u = M high,3 × M high,2 × M high,1 operations, the 4×4 operation matrix for the high-order part is obtained.

[0051] For the operations on the two low-order qubits, the corresponding probability amplitudes are respectively operated on within each node from the first layer to the third layer, and the operations on the nodes are carried out in parallel. For the first layer of the quantum circuit, a G gate and an I gate (i.e., no gate) are respectively added to the two low-order qubits. Using the state vector simulation method, for the operation on qubit q 1 at each node, (α i , 0, α i,2 ) forms a pair of probability amplitudes, and (α i,1 , α i,3 ) forms a pair of probability amplitudes, where i represents the i-th node. The pair of probability amplitudes and the 2×2 G gate matrix are respectively subjected to matrix-vector multiplication operations to obtain the results, that is For the qubit q 0 with no gate operation, no update of the state vector needs to be done. For the second layer of the quantum circuit, the qubit q 1 serves as the target bit of the control gate, and its gate operation is still on the qubit q 1 , but it needs to be controlled by the qubit q 2 . Therefore, when performing the gate operation on the qubit q 1 , still (α i,0 , α i,2 ) forms a pair of probability amplitudes, and (α i,1 , α i,3 ) forms a pair of probability amplitudes. However, the matrix-vector multiplication operation with the G gate only occurs at the two nodes where i = 1 and 3, because only at this time, the state of the qubit q 2 is |1>. For the G gate on the qubit q 0 , its quantum gates respectively act on the (α i,0 , α i,1 ) and (α i,2 , α i,3 ) pairs of probability amplitudes at each node. For the third layer of the quantum circuit, no gate operations are added to the two low-order qubits, so no update of the state vector is done.

[0052] The last step of the merging operation: further divide the state vector shards into blocks on each distributed node, as Figure 3 shown. In this embodiment, there are a total of 4 distributed nodes. Therefore, the state vector shard V i on the i-th node is divided into 4 blocks, denoted as V i,0 , V i,1 , V i,2 , V i,3 respectively. In this embodiment, since there are 4 probability amplitudes on each node, there is only 1 probability amplitude in each block of each node, that is, there is only the probability amplitude α i,0 in V i,0 . After the block division, a block recombination operation is performed, and the V i,j block is sent to the i-th block of node j. In this embodiment, the blocks V 0,0 , V 1,0 , V 2,0 , V 3,0 are recombined onto node 0, and the blocks V 0,1 , V 1,1 , V 2,1 , V 3,1 are recombined onto node 1, the blocks V 0,2 , V 1,2 , V 2,2 , V 3,2 are recombined onto node 2, and the blocks V 0,3 , V 1,3 , V 2,3 , V 3,3 are recombined onto node 3. Then, the ⊙ operation is performed on each node. The specific calculation process is as Figure 3 shown, and V i,j is updated to After the state vector is updated by the ⊙ operation, the calculation results are put back to the correct state vector positions through block redistribution. That is, is sent back to node 0, is sent back to node 1, is sent back to node 2, and is sent back to node 3.

[0053] According to the above introduction, the present invention realizes an efficient quantum circuit simulation method based on a distributed system through four steps: data partitioning, qubit classification, parallel processing of high and low bit operations, and the final merging operation. In the present invention, the state vector is continuously allocated to each distributed node, and all qubits are divided into two categories: high-order qubits and low-order qubits according to the number of nodes. The processing on high-order qubits and the operations on low-order qubits are executed separately and in parallel. The number of communications required for quantum circuit simulation is reduced from T times to only one time in the final merging operation, significantly reducing the communication overhead in distributed quantum circuit simulation. In the final merging operation, by dividing the state vector shards into blocks and implementing efficient fusion of operations on high-order qubits through block recombination, ⊙ operation, and block redistribution operations. The communication volume in this step is almost the same as that in each previous step. The reduction of communication overhead significantly improves the efficiency of distributed quantum circuit simulation, and the simulation time linearly accelerates with the increase in the number of nodes, showing good scalability.

Claims

1. An efficient quantum circuit simulation method based on a distributed system, characterized in that, the method includes: Step 1, data partitioning: Initialize the state vector of the circuit according to the initial state of the qubits, evenly divide the state vector into multiple slices, and place each slice on each distributed node; Step 2, qubit classification: Divide the qubits in the circuit into two categories, high-order qubits and low-order qubits, according to the number of distributed nodes; Step 3, parallel processing of high and low order operations: On each node, parallelly execute the processing on high-order qubits and the operations on low-order qubits. Perform multiplication operations between the unitary matrices composed of the tensor product operations of the quantum gates on the high-order part of each layer of the quantum circuit, and directly update each slice of the state vector in place for the quantum gates on the low-order part; Step 4, the final merging operation: Divide the slices of the state vector on each node into multiple blocks, and truly integrate the quantum gate operations on the high-order part into the state vector update through three steps of block reorganization, calculation, and block reallocation between distributed nodes; where: The step 1 of evenly dividing the state vector into multiple slices and placing each slice on each distributed node specifically includes: Given a quantum circuit of n qubits, the size of its corresponding state vector is 2 n , denoted as N; if there are H nodes in the distributed system, then the state vector of size N will be evenly divided into N parts, and each state vector shard has H / H probability amplitudes, and each probability amplitude is denoted as α s ; the state vector is continuously divided, that is, the probability amplitude index s on each node is continuous; the state vector shards on H nodes are respectively denoted as V 0 , V 1 ,..., V H-1 ; Each of the state vector shards has N / H probability amplitudes, and each probability amplitude is denoted as α s , specifically including: For the probability amplitude α on each node s , its index s consists of two parts, denoted as s 1 and s 2 respectively; s 1 represents the number of the distributed node, ranging from 0 to H - 1; s 2 represents the number of the probability amplitude within the node, ranging from 0 to N / H - 1; therefore, s 1 on different nodes are different from each other, but s 2 represents values ranging from 0 to N / H - 1 for all of them; The step 2 of dividing the qubits in the circuit into high-order qubits and low-order qubits according to the number of distributed nodes specifically includes: Given a quantum circuit with n qubits, which are denoted as q n-1 , q n-2 ,..., q 1 , q 0 ; where q 0 represents the qubit with the lowest weight in the probability amplitude index, i.e., s = q n-1 , q n-2 ,..., q 1 , q 0 ; After step 1, each node has N / H probability amplitudes, denoted as L = N / H; According to the state vector simulation method, the quantum gates on the log 2 L qubits with lower weights can directly update the state vector within each distributed node, denoted as l = log 2 L; q l-1 ,..., q 1 , q 0 These l qubits are called low-order qubits; For the remaining n - l qubits, the quantum gate operations on them must rely on the probability amplitudes on multiple nodes to complete the state vector update. Therefore, communication between distributed nodes is required. These n - l qubits are called high-order qubits, denoted as h = n - l; That is, the quantum circuit has a total of n qubits q n-1 , q n-2 ,..., q 1 , q 0 , which are divided into h high-order qubits q n-1 ,..., q l+1 , q l and l low-order qubits q l-1 ,..., q 1 , q 0 ; where n = log 2 N, h = log 2 H, l = log 2 L; The step 4 of dividing the slices of the state vector on each node into multiple blocks specifically includes: After data partitioning, each distributed node has L probability amplitudes; the L probability amplitudes are continuously partitioned into H blocks within the node, and each block has L / H probability amplitudes; the H blocks on the i-th node are respectively denoted as V i,0 , V i,1 ,..., V i,H-1 ; The step 4 of truly integrating the quantum gate operations on the high-order part into the state vector update through three steps of block reorganization, calculation, and block reallocation between distributed nodes specifically includes: The block reorganization process reorganizes the j-th block on each distributed node to the j-th node for subsequent calculations, that is, the j-th block on the i-th node is sent to the i-th block on the j-th node; the calculation process performs a scalar multiplication operation of vector sharding on the state vector shards on each distributed node after block reorganization, denoted as the "⊙" operation; the block redistribution process is to return the results calculated on each distributed node back to each node according to the block size to obtain the final correct result. u Perform a scalar multiplication operation of vector sharding with the unitary matrix M obtained above and the state vector shards on each distributed node after block reorganization, denoted as the "⊙" operation; the block redistribution process is to return the results calculated on each distributed node back to each node according to the block size to obtain the final correct result.

2. The efficient quantum circuit simulation method based on a distributed system according to claim 1, characterized in that, The step 3 of performing multiplication operations between the unitary matrices composed of the tensor product operations of the quantum gates on the high-order part of each layer of the quantum circuit specifically includes: Given a quantum circuit with T layers of circuits, and multiple quantum gate operations acting on corresponding qubits on each layer; for the quantum gate operations on the high-order qubits of each layer of circuits, the quantum gates can all be represented as small unitary matrices; perform a tensor product operation between the unitary matrices of the high-order quantum gates to obtain a unitary matrix of size H×H; then, obtain the final unitary matrix through matrix multiplication of the H×H-sized unitary matrices on the T layers of circuits; denote the H×H-sized unitary matrices on each layer as M high,1 , M high,2 ,..., M high,T , and through matrix multiplication M high,T ·...·M high,2 ·M high,1 obtain the final unitary matrix M u that represents the operation on the high-order qubits.

3. The efficient quantum circuit simulation method based on a distributed system according to claim 1, characterized in that, The step 3 of directly updating each slice of the state vector in place for the quantum gates on the low-order part specifically includes: According to the state vector simulation method, when updating the state vector for the gate operation on the k-th qubit, the step size of the probability amplitude used is 2 k ; therefore, the maximum step size of the probability amplitude required for the gate operations on the lower l qubits is 2 l-1 , and the consecutive L probability amplitudes at a node are fully capable of satisfying the gate operations on the lower l qubits; for a quantum circuit with T layers, on each distributed node, according to the state vector simulation method, the gate operations on the lower qubits of each layer are continuously executed, and the state vector is updated in place.

4. The efficient quantum circuit simulation method based on a distributed system according to claim 1, characterized in that, The step 3 of parallelly executing the processing on high-order qubits and the operations on low-order qubits on each node specifically includes: For the parallel execution of the processing on high-order qubits and the operations on low-order qubits, these two are parallelly executed in a multi-threaded manner on each distributed node.

5. The efficient quantum circuit simulation method based on a distributed system according to claim 1, characterized in that, The scalar multiplication operation of the vector slice, denoted as the "⊙" operation, specifically includes: The operation acts on an arbitrary H×H matrix M and a vector w, where the size of the vector w is a multiple of H; the operation is denoted as M⊙w = w′, where the vector w consists of H shards, denoted as w 0 , w 1 ,..., w H-1 , and a vector w′ with H shards is calculated, and each shard is denoted as w′ 0 , w′ 1 ,..., w′ H-1 ; the specific operation process is where i is an integer in the range of [0, H - 1], and m ij represents the element in the i-th row and j-th column of the matrix M; the ⊙ operation is to perform a scalar multiplication operation on each row of the matrix M with the vector shards, and add up the results of one row of operations by shards, which is one shard of the final result w′.

Citation Information

Patent Citations

  • Novel universal quantum gate and quantum circuit optimization method

    CN108334952A

  • Efficient simulation method based on quantum circuit

    CN115759270A