Devices and methods for quantum circuit simulators

By using greedy algorithms with backtracking in quantum circuit simulators to generate the optimal quantum gate sequence and qubit set, the problems of huge memory requirements and low data exchange efficiency in the prior art are solved, and significant performance improvement is achieved.

CN113508404BActive Publication Date: 2025-06-27HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980093261.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-03-29
Publication Date
2025-06-27
Estimated Expiration
2039-03-29

AI Technical Summary

Technical Problem

When existing quantum circuit simulators simulate quantum computing on classical computers, they face problems such as huge memory requirements, low data exchange efficiency and failure to effectively determine the optimal qubit and gate permutation order.

Method used

Greedy algorithms, especially those with backtracking, generate the optimal quantum gate sequence and qubit set, and form a cluster by merging quantum gates to optimize data exchange and arithmetic operations in quantum circuit simulators.

Benefits of technology

It significantly reduces the amount of data transmission and arithmetic operations between nodes, improves the performance of quantum circuit simulators, and even improves several times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113508404B_ABST
    Figure CN113508404B_ABST
Patent Text Reader

Abstract

The present invention provides a device for a quantum circuit simulator and a quantum circuit simulator including at least one such device. The device is configured to: obtain a first quantum gate sequence; generate a second quantum gate sequence as a subsequence of the first quantum gate sequence; calculate a local qubit set and a global qubit set according to the second quantum gate sequence; generate a quantum gate cluster set, where each cluster includes a subset of the quantum gates in the second quantum gate sequence merged together using a greedy algorithm; generate a third quantum gate sequence according to the order of the clusters, the third quantum gate sequence including all the quantum gates in the second quantum gate sequence; provide the local qubit set and the global qubit set to the quantum circuit simulator; and output the third quantum gate sequence to the quantum circuit simulator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of quantum computing, and more particularly, to the simulation of quantum circuits on classical computers. Specifically, the present invention relates to a device for a quantum circuit simulator, and a quantum circuit simulator including at least one such device. In addition, the present invention relates to a method for quantum gate and qubit scheduling for a quantum circuit simulator, wherein the method can be executed by the device. Background Art

[0002] A general quantum circuit simulator stores a mathematical representation of the entire state of a quantum computer in memory. The size of this state scales as 2 n , where n is the number of simulated qubits of the quantum computer. For 40 qubits, the size of this state is 16 TiB. This requires the use of a multi-node computing system in order to distribute the large state across the memories of the nodes. During the simulation of a quantum circuit, a portion of the state needs to be accessed from a remote node.

[0003] To simulate quantum computing on a classical computer, one can use the linear algebra representation of quantum computing (quantum circuits). In this representation, the state of an n-qubit quantum circuit is a vector in a Hilbert space with an orthonormal basis of vectors The dimension of the space is equal to 2 n . According to quantum computing theory, the following relationship holds:

[0004]

[0005] From the above relationship, a direct way to represent the state of a quantum computer in memory is to store 2 n complex numbers {α i}, which are called the amplitudes of the corresponding basis states. The value |α i | 2 determines the probability of observing the basis state i as the output of the quantum circuit / computer.

[0006] Quantum computing can be represented as a linear unitary operator U acting on a vector to produce a resulting state :

[0007]

[0008] Since the basis in the Hilbert space is defined, the operator U is represented by a 2 n × 2 n dimensional matrix.

[0009] In quantum computing, a quantum gate is defined as a fundamental unitary operator acting on one or several qubits. The sizes of practical quantum gates are 1-qubit, 2-qubit, and 3-qubit. Using these quantum gates, any quantum algorithm can be represented. According to the above equation (2), any quantum algorithm can be represented by the relationship between a unitary matrix and a sequence of quantum gates and the operator U:

[0010]

[0011] In other words, a quantum algorithm can be represented as a tensor product of quantum gates, with each quantum gate acting on a subset of qubits.

[0012] Figure 8 A set of typical quantum gates used in most common quantum algorithms is shown. Among them, CNOT and CZ are examples of a special type of quantum gate, which are called controlled gates. Such quantum gates act on two or more qubits, where one or more qubits are used to control certain operations. The qubit on which the operation is performed is called the target, and the other qubits are called controls.

[0013] Using a graphical representation, a quantum circuit can be drawn for a quantum algorithm, as Figure 9 schematically shown. The numbered horizontal lines represent qubits, and the quantum gates acting on the qubits are located on the corresponding lines. The quantum gates are applied in the order from left to right. According to the above relationship (3), it can be concluded from the properties of the tensor product that quantum gates acting on disjoint sets of qubits commute. A set of quantum gates sharing the same horizontal position is called a layer of the quantum circuit.

[0014] As mentioned above, a general quantum circuit simulator stores an array of 2 n complex numbers (the coefficients α i ) in the computer memory. Using, for example, IEEE754 double-precision floating-point representation, this requires 16·2 n bytes of memory. One can easily see that when the number of qubits grows (for example, 40 qubits require 16 TiB of memory), the memory requirements quickly become intractable for a single computer. In such a case, the simulator program has to split the state vector into multiple parts and store them in the memory of several computers (nodes, as described above).

[0015] Let the quantum simulator run on n = L + R qubits. Then, if a single computer can only store 2 L elements of the state vector, the number of computer nodes required is 2 R .

[0016] The natural way to choose a basis in the above relationship (1) is to assign a basis state Assign a state in which, according to the binary representation of the index i, the qubit is |0> or |1>. For example: for three qubits, there are 8 basis states In the basis state all qubits are in the state |0>. In qubit 1 is in the state |1>, and the other two are in the state 0. In qubits 1 and 2 are in the state |1>, and qubit 0 is in the state |0>.

[0017] According to the state vector distribution scheme, it is obvious that when the states of the other R qubits are fixed to be equal to the binary representation of the node rank, each node stores all the amplitudes that determine the probabilities of |0> and |1> of the first L qubits. In this document, the first L qubits are called local qubits, and the last R qubits are called global qubits.

[0018] When a quantum gate is applied to one or more local qubits, the matrix-vector multiplication is performed locally on each node without accessing the amplitudes stored in remote nodes because the other qubits are not affected by the gate. When a quantum gate is applied to one or more global qubits, the matrix-vector multiplication cannot be performed because the computing nodes cannot directly access the memory of remote computers. In this case, a data exchange mechanism is needed.

[0019] A traditional method provides a method for qubit reordering, which is performed in the following cases: the qubits are renumbered, the corresponding amplitudes are transmitted between nodes, and stored in the memory of the corresponding nodes according to the new qubit numbering and the rank of the nodes. This process is called qubit swapping because the qubits and the amplitudes exchange their positions. Figure 11 illustrates this process. This method can be used to simulate a quantum gate that is initially applied to global qubits. In this case, it is only necessary to exchange the numbers between the global qubits involved in the operation and some unused local qubits, and then transmit the corresponding amplitudes between nodes. After that, a quantum gate acting on local qubits can be simulated.

[0020] Distributed computing usually uses the MPI library to perform data exchange between nodes, so the MPI operations are used to represent the data exchange pattern in the program. The qubit swapping operation can be completed using a single MPI_Alltoall operation. Any number of qubits less than or equal to R can be swapped at once. It is easy to prove that the amount of amplitudes transmitted is equal to

[0021]

[0022] Among them, k is the number of globally swapped qubits. It can be clearly seen from the above relationship (4) that the data transfer required to swap several qubits at once is less than that required to swap them sequentially one by one.

[0023] However, a typical quantum circuit can contain hundreds of thousands of gates. Without any optimization techniques, each gate implies matrix-vector multiplication, and in a distributed scenario, the amplitudes have to be transferred many times between nodes. Therefore, in the method described above, without carefully defining the set of qubits to be swapped, if some qubits in the collection are not included in a sufficient number of gate applications, there will be additional overhead in data swapping. This method does not provide any advice on how to determine the optimal set of qubits to be reordered.

[0024] Another method describes an open-source implementation of a distributed quantum circuit simulator (QuEST). In QuEST, the method of qubit reordering described above is used, but the implementation is limited to single-qubit swaps.

[0025] The most sophisticated method of quantum circuit simulation uses a scheduling component (scheduler) that determines the order of the gates to be applied and the set of qubits to be reordered. The gates are reordered into sequences called levels. A level contains gates acting on local qubits. Within a level, the gates form subsequences called clusters. The gates of the same cluster are fused into a single multi-qubit gate that is simulated by a single matrix-vector multiplication. Between levels, qubit reordering is performed.

[0026] Figure 12 Such gate and level clustering is shown. Assuming qubits 0 - 2 are currently local and 3 - 4 are global, the first level consists of 2 gate clusters outlined by gray lines, and the second level consists of the clusters outlined by black lines. After applying the gates for the first-level qubits, reordering is performed: 3, 4 are swapped with 1, 2, and then the second level can be applied.

[0027] The main problem in implementing this method is the method of constructing clusters and levels. This method does not describe any algorithms and does not provide the source code of the scheduler.

[0028] In summary, although there is a set of main methods for quantum circuit simulation, including gate scheduling, construction of gate clusters, and qubit reordering, the problem of finding the optimal order of gates and qubits remains unsolved. None of the previous methods describe any method for calculating qubit and gate permutations according to clear optimality criteria. Summary of the Invention

[0029] In view of the above problems and drawbacks, the aim of embodiments of the present invention is to improve existing methods. The aim is to provide a method for calculating complex gates and qubit permutations for a quantum circuit simulator. This results in optimal data exchange and optimal quantum gate application scheduling in the quantum circuit simulator and, accordingly, reduces the amount of data transferred between nodes. The calculated permutation should provide the minimum number of matrix-vector multiplications and the minimum amount of data transfer. To this end, a device and a method should be provided that can be used in a distributed quantum circuit simulator for gate scheduling and qubit reordering scheduling.

[0030] This aim is achieved by embodiments of the present invention described in the appended independent claims. Advantageous implementations of the present invention are further defined in the dependent claims.

[0031] In particular, embodiments of the present invention provide a device and a method that calculate optimal data exchange and quantum gate application scheduling, thereby significantly reducing the amount of data transferred between nodes and the amount of arithmetic operations to be performed. All of this improves the performance of the quantum circuit simulator, even by several times.

[0032] Embodiments of the present invention are based on the associativity of tensor product operations, which allows relation (3) to be split into multiple factors in different ways, thereby constructing the factors in consideration of the performance or memory consumption of the calculation:

[0033]

[0034] The above relation (5) and the commutation property of quantum gates are applied, which forms the core of embodiments of the present invention and optimizes quantum circuit simulation through gate sequence permutation.

[0035] Based on the characteristics of individual gates and using a greedy algorithm, the device and the method specifically calculate the permutation of gates and the permutation of qubits, which results in the minimum number of clusters in a stage and the minimum number of stages during quantum circuit simulation.

[0036] A first aspect of the present invention provides a device for a quantum circuit simulator, the device being configured to: obtain a first quantum gate sequence; generate a second quantum gate sequence, which is a subsequence of the first quantum gate sequence, by using a greedy algorithm, in particular a greedy algorithm with backtracking; calculate a local qubit set and a global qubit set according to the second quantum gate sequence; generate a set of quantum gate clusters, where each cluster includes a subset of the quantum gates in the second quantum gate sequence merged together by using a greedy algorithm; generate a third quantum gate sequence according to the order of the clusters, the third quantum gate sequence including all the quantum gates in the second quantum gate sequence; provide the local qubit set and the global qubit set to the quantum circuit simulator; output the third quantum gate sequence to the quantum circuit simulator.

[0037] The calculated local and global qubit sets are especially the "optimal" local and global qubit sets. Thus, "optimal" means the best state that the algorithm can achieve. That is, the algorithm searches many variants of these qubit sets and then can select the qubit set with the largest number of gates in the second sequence. Before the device of the first aspect runs the algorithm, the local qubit set can be intentionally predefined. This means that the algorithm will include quantum gates that act on these qubits.

[0038] The device of the first aspect can be used for a distributed quantum circuit simulator and can provide gate scheduling and qubit reordering. In other words, the device can provide complex gate and qubit permutation calculations for the quantum circuit simulator. The calculated permutation can obtain an optimal data exchange and quantum gate application schedule in the quantum circuit simulator, thereby significantly reducing the amount of data transmitted between simulator nodes.

[0039] In one implementation of the first aspect, the device is further configured to, when generating the cluster set of quantum gates: arrange the clusters including more quantum gates before the clusters including fewer quantum gates in the order of the clusters.

[0040] In one implementation of the first aspect, the device is further configured to, when generating the cluster set of quantum gates: generate the clusters according to the maximum possible number of qubits in the clusters.

[0041] The above implementations can improve the efficiency of the algorithm executed by the device of the first aspect.

[0042] In one implementation of the first aspect, the device is further configured to, when generating the cluster set of quantum gates: select all possible combinations of qubits associated with the second quantum gate sequence one by one according to the maximum possible number of qubits in the clusters; construct clusters for each combination; and select the cluster with the largest number of gates among them.

[0043] In one implementation of the first aspect, the device is further configured to, when generating the cluster set of quantum gates: maintain a locked qubit set; if the matrix representation of the quantum gate is diagonal, include the quantum gate in the cluster; if at least one of the qubits on which the quantum gate acts does not belong to the selected qubit combination, skip the quantum gate; and / or if at least one of the qubits on which the quantum gate acts is in the locked qubit set, skip the quantum gate; if the quantum gate is skipped, add all the qubits on which the quantum gate acts to the locked qubit set; otherwise, include the quantum gate in the cluster.

[0044] In one implementation of the first aspect, the device is further configured to, when generating the quantum gate cluster set: determine the cluster including the maximum number of quantum gates; output the quantum gates of the determined cluster, specifically insert the output quantum gates into the third quantum gate sequence; and delete the output quantum gates from the second quantum gate sequence.

[0045] In one implementation of the first aspect, the device is further configured to, when calculating the local qubit set and the global qubit set: determine the local qubit set and / or the global qubit set respectively according to the maximum number of local qubits and / or the maximum number of global qubits.

[0046] In one implementation of the first aspect, the device is further configured to, when generating the second quantum gate sequence: fuse the quantum gates acting on a single qubit with the adjacent quantum gates in the first quantum gate sequence acting on a subset of qubits including the same single qubit.

[0047] In one implementation of the first aspect, the device is further configured to, when generating the second quantum gate sequence: include the quantum gates acting on at most the maximum number of local qubits into the second quantum gate sequence; if the first quantum gate sequence includes at least one quantum gate acting on a single qubit and another quantum gate acting on the same qubit and at least one other qubit, include the single qubit gate and the other multi - qubit gate together into the second quantum gate sequence.

[0048] In one implementation of the first aspect, the device is further configured to, when generating the second quantum gate sequence: create a branch of the greedy algorithm, where quantum gates are included into the second quantum gate sequence; and / or create a branch of the greedy algorithm, where the quantum gates in the first quantum gate sequence are skipped; if a quantum gate is included, add all the qubits on which the quantum gate acts to the local qubit set; or if a quantum gate is skipped, add all the qubits on which the quantum gate acts to the locked qubit set.

[0049] In one implementation of the first aspect, the device is further configured to, when generating the second quantum gate sequence: create at most the maximum number of branches of the greedy algorithm.

[0050] In one implementation of the first aspect, the device is further configured to, when applying a branch of the greedy algorithm: construct the second quantum gate sequence with as many gates as possible; test each gate in the first quantum gate sequence and, according to the test result, skip the gate or include the gate into the second quantum gate sequence.

[0051] In one implementation of the first aspect, the device is further configured to, when generating the second quantum gate sequence: maintain a locked qubit set; skip the quantum gate if applying the quantum gate requires more qubits than a predetermined threshold to be local qubits; and / or skip the quantum gate if at least one of the qubits on which the quantum gate acts is in the locked qubit set, and if the quantum gate is skipped, add all the qubits on which the quantum gate acts to the locked qubit set.

[0052] In one implementation of the first aspect, the device is further configured to, when generating the second quantum gate sequence: include the quantum gate in the second quantum gate sequence if the matrix representation of the quantum gate is diagonal, without adding the qubits on which the quantum gate acts to the local qubit set; and / or include the quantum gate in the second quantum gate sequence if all the qubits on which the quantum gate acts are already in the local qubit set.

[0053] In one implementation of the first aspect, the device is further configured to, when calculating the local qubit set and the global qubit set: construct a set of all the qubits on which the quantum gates in the first quantum gate sequence act; include all the qubits on which the quantum gates in the second quantum gate sequence act in the local qubit set; include all the qubits that are in the set of all the qubits and not in the local qubit set in the global qubit set.

[0054] The second aspect of the present invention provides a quantum circuit simulator, including a device according to the first aspect or any of its implementations.

[0055] The third aspect of the present invention provides a method for quantum gate and qubit scheduling for a quantum circuit simulator, the method including: obtaining a first quantum gate sequence; generating a second quantum gate sequence that is a subsequence of the first quantum gate sequence by using a greedy algorithm, particularly a greedy algorithm with backtracking; calculating a local qubit set and a global qubit set according to the second quantum gate sequence; generating quantum gate clusters, where each cluster includes a subset of the quantum gates in the second quantum gate sequence merged together by using a greedy algorithm; generating a third quantum gate sequence according to the order of the clusters, the third quantum gate sequence including all the quantum gates in the second quantum gate sequence; providing the local qubit set and the global qubit set to the quantum circuit simulator; outputting the third quantum gate sequence to the quantum circuit simulator.

[0056] A fourth aspect of the present invention provides a computer program product comprising program code for controlling a device according to the first aspect or any implementation thereof, or for performing the method according to the third aspect or any implementation thereof when implemented on a processor.

[0057] In one implementation of the fourth aspect, the method further includes, when generating a set of clusters of quantum gates: arranging the clusters including a larger number of quantum gates before the clusters including a smaller number of quantum gates in the order of the clusters.

[0058] In one implementation of the fourth aspect, the method further includes, when generating a set of clusters of quantum gates: generating the clusters according to the maximum possible number of qubits in the clusters.

[0059] In one implementation of the fourth aspect, the method further includes, when generating a set of clusters of quantum gates: selecting, one by one, all possible combinations of qubits associated with the second sequence of quantum gates according to the maximum possible number of qubits in the clusters; constructing a cluster for each combination; and selecting the cluster with the largest number of quantum gates among them.

[0060] In one implementation of the fourth aspect, the method further includes, when generating a set of clusters of quantum gates: maintaining a set of locked qubits; if the matrix representation of a quantum gate is diagonal, including the quantum gate in the cluster; if at least one of the qubits on which the quantum gate acts does not belong to the selected combination of qubits, skipping the quantum gate; and / or if at least one of the qubits on which the quantum gate acts is in the set of locked qubits, skipping the quantum gate; if a quantum gate is skipped, adding all the qubits on which the quantum gate acts to the set of locked qubits; otherwise, including the quantum gate in the cluster.

[0061] In one implementation of the fourth aspect, the method further includes, when generating a set of clusters of quantum gates: determining the cluster including the largest number of quantum gates; outputting the quantum gates of the determined cluster, in particular inserting the output quantum gates into the third sequence of quantum gates; and deleting the output quantum gates from the second sequence of quantum gates.

[0062] In one implementation of the fourth aspect, the method further includes, when calculating the set of local qubits and the set of global qubits: determining the set of local qubits and / or the set of global qubits respectively according to the maximum number of local qubits and / or the maximum number of global qubits.

[0063] In one implementation of the fourth aspect, the method further includes, when generating the second quantum gate sequence: fusing a quantum gate acting on a single qubit with adjacent quantum gates in the first quantum gate sequence acting on a subset of qubits including the same single qubit.

[0064] In one implementation of the fourth aspect, the method further includes, when generating the second quantum gate sequence: including a quantum gate acting on at most the maximum number of local qubits in the second quantum gate sequence; if the first quantum gate sequence includes at least one quantum gate acting on a single qubit and another quantum gate acting on the same qubit and at least one other qubit, including the single-qubit gate and the other multi-qubit gate together in the second quantum gate sequence.

[0065] In one implementation of the fourth aspect, the method further includes, when generating the second quantum gate sequence: creating a branch of the greedy algorithm where a quantum gate is included in the second quantum gate sequence; and / or creating a branch of the greedy algorithm where a quantum gate in the first quantum gate sequence is skipped; if a quantum gate is included, adding all the qubits on which the quantum gate acts to the local qubit set; or, if a quantum gate is skipped, adding all the qubits on which the quantum gate acts to the locked qubit set.

[0066] In one implementation of the fourth aspect, the method further includes, when generating the second quantum gate sequence: creating at most the maximum number of branches of the greedy algorithm.

[0067] In one implementation of the fourth aspect, the method further includes, when applying a branch of the greedy algorithm: constructing the second quantum gate sequence with as many gates as possible; testing each gate in the first quantum gate sequence and, based on the test results, skipping the gate or including the gate in the second quantum gate sequence.

[0068] In one implementation of the fourth aspect, the method further includes, when generating the second quantum gate sequence: maintaining a locked qubit set; skipping a quantum gate if applying the quantum gate requires more qubits than a predetermined threshold to be local qubits; and / or skipping a quantum gate if at least one of the qubits on which the quantum gate acts is in the locked qubit set; if a quantum gate is skipped, adding all the qubits on which the quantum gate acts to the locked qubit set.

[0069] In an implementation of the fourth aspect, the method further includes, when generating the second quantum gate sequence: if the matrix representation of a quantum gate is diagonal, including the quantum gate into the second quantum gate sequence and not adding the qubits on which the quantum gate acts to the local qubit set; and / or if all the qubits on which a quantum gate acts are already in the local qubit set, including the quantum gate into the second quantum gate sequence.

[0070] In an implementation of the fourth aspect, the method further includes, when calculating the local qubit set and the global qubit set: constructing a set of all the qubits on which the quantum gates in the first quantum gate sequence act; including all the qubits on which the quantum gates in the second quantum gate sequence act into the local qubit set; and including all the qubits that are in the set of all the qubits and not in the local qubit set into the global qubit set.

[0071] It should be noted that all the devices, elements, units, and modules described in this application can be implemented in software or hardware elements or any type of combination thereof. All the steps performed by the various entities described in this application and the functions described to be performed by the various entities are intended to indicate that each entity is suitable or used to perform its respective steps and functions. Although in the description of the following specific embodiments, the specific functions or steps performed by external entities are not reflected in the description of the specific detailed elements of the entity performing the specific steps or functions, those skilled in the art should clearly understand that these methods and functions can be implemented in the corresponding hardware or software elements or any combination thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] In combination with the accompanying drawings, the following description of specific embodiments elaborates on the various aspects and implementations of the present invention described above.

[0073] Figure 1 The device for a quantum circuit simulator according to an embodiment of the present invention is shown.

[0074] Figure 2 The pseudocode of the cluster scheduling method performed by the device for a quantum circuit simulator according to an embodiment of the present invention is shown.

[0075] Figure 3 The block diagram of the cluster scheduling method performed by the device for a quantum circuit simulator according to an embodiment of the present invention is shown.

[0076] Figure 4 The pseudocode of the stage scheduling method performed by the device for a quantum circuit simulator according to an embodiment of the present invention is shown.

[0077] Figure 5A block diagram of a stage scheduling method performed by a device for a quantum circuit simulator according to an embodiment of the present invention is shown.

[0078] The results of the scheduler on different top circuits are shown in Fig. 6(a), and the results of the 30-layer top circuit simulation of the quantum circuit simulator according to an embodiment of the present invention compared with the QuEST simulator on an 8-node cluster are shown in Fig. 6(b).

[0079] Figure 7 A method for quantum gate and qubit scheduling for a quantum circuit simulator according to an embodiment of the present invention is shown.

[0080] Figure 8 Typical quantum gates and their quantum circuit representations are shown.

[0081] Figure 9 A graphical representation of a quantum circuit for a quantum algorithm is shown.

[0082] Figure 10 A scheme of the state vector distribution is shown.

[0083] Figure 11 Qubit exchange is shown.

[0084] Figure 12 A cluster of gates and stages is shown. Detailed Description of the Invention

[0085] Figure 1 Device 100 according to an embodiment of the present invention is shown. Device 100 is applicable to quantum circuit simulator 110. Device 100 can be a part of quantum circuit simulator 110 or can be connected to quantum circuit simulator 110. Device 100 is specifically used for scheduling the quantum gates and qubits of quantum circuit simulator 110 to improve the performance of the quantum circuit simulator. Quantum circuit simulator 110 can be one or more classical computers or computer nodes that are used together to simulate the execution of quantum circuits on a quantum computer. Quantum circuit simulator 110 can include at least one device 100 or can work with at least one device 100.

[0086] Device 100 is used to obtain a first quantum gate sequence 101 according to a quantum circuit received as an input to device 100. The quantum circuit can be a quantum circuit to be simulated on / by quantum circuit simulator 110. Device 100 is also used to generate a second quantum gate sequence 102, and the second quantum gate sequence 102 is a subsequence of the first quantum gate sequence 101. Therefore, device 100 uses a greedy algorithm, especially a greedy algorithm with backtracking. That is, the second quantum gate sequence 102 is generated according to the first quantum gate sequence 101 using a greedy algorithm with backtracking.

[0087] In addition, the device 100 is used to calculate a local qubit set 103a and a global qubit set 103b respectively according to the generated second quantum gate sequence 102. These qubit sets can be referred to as optimal or final qubit sets. In addition, the device 100 is also used to generate a set of clusters 104 of quantum gates, where each cluster 104 includes a subset of the quantum gates in the second quantum gate sequence 102 merged together by using a greedy algorithm. The greedy algorithm can be essentially similar to the greedy algorithm used to generate the second sequence 102. Then, the device 100 is used to generate a third quantum gate sequence 105 according to the order of the clusters 104 of quantum gates, and the third quantum gate sequence 105 contains all the quantum gates in the second quantum gate sequence 102.

[0088] Finally, the device 100 is used to provide the local qubit set 103a and the global qubit set 103b to the quantum circuit simulator 110, and output the third quantum gate sequence 105 to the quantum circuit simulator. According to these inputs, the quantum circuit simulator 110 can simulate a quantum circuit, with less data to be transmitted between multiple nodes of the simulator 110 and fewer arithmetic operations performed.

[0089] It is worth noting that in Figure 1 the device 100, the generation of the clusters 104 of quantum gates and the generation of the third quantum gate sequence 105 can be referred to as a cluster scheduling algorithm. This algorithm enables the device 100 to perform quantum gate scheduling of the simulator 110. The calculation and output of the qubit sets 103a and 103b can be referred to as a level scheduling algorithm. This algorithm enables the device 100 to perform qubit scheduling of the simulator 110.

[0090] Figure 2 The pseudocode of the cluster scheduling algorithm is shown. The cluster scheduling algorithm can be executed by the device 100 according to an embodiment of the present invention, particularly Figure 1 the device 100, so as to generate the set of clusters 104 and output the third quantum gate sequence 105. Figure 3 The block diagram of the cluster scheduling algorithm is further shown.

[0091] The cluster scheduling algorithm has two parameters: "qubits", which is the set of all qubits involved in the input sequence of quantum gates; and k, which is the maximum possible number of qubits in the cluster 104. The algorithm also takes the sequence of quantum gates (specifically the second quantum gate sequence 102) as input.

[0092] This algorithm further combines quantum gates into quantum gate clusters 104. Thus, it attempts to minimize the total number of generated clusters 104. In addition, this algorithm uses a greedy method, which: (a) finds the cluster 104 that contains the largest number of quantum gates; (b) returns this cluster 104 as the result; and deletes this quantum gate cluster 104 from the input sequence of quantum gates; (c) repeats (a).

[0093] In step (a), the algorithm can select all possible combinations of k qubits one by one; a sequence of quantum gates containing only qubits can be generated from this combination, where these quantum gates can be combined in a cluster 104; the list with the largest size can be selected as the next cluster 104.

[0094] Device 100 can also immediately fuse single-qubit quantum gates. If there is at least one multi-qubit gate acting on qubit q, the single-qubit gate g acting on qubit q does not change the total number of levels. Therefore, this quantum gate g can be immediately fused (combined) into any adjacent quantum gate containing qubit q. This optimization helps to significantly speed up the level scheduling algorithm, which can be executed by device 100 and is described below.

[0095] Figure 4 The pseudo-code of the level scheduling algorithm is shown, and the level scheduling algorithm can be executed by device 100 according to an embodiment of the present invention, especially Figure 1 device 100, in order to schedule and output qubits. Figure 5 The block diagram of the level scheduling algorithm is shown.

[0096] The level scheduling algorithm has two parameters: L max , representing the maximum number of local qubits; B max , representing the maximum number of branches to be created. This algorithm takes a list of quantum gates as input. This algorithm returns a set of 103a qubits that must be local during the current level. Therefore, this algorithm attempts to minimize the total number of levels. In particular, this algorithm uses a greedy method, that is, this method constructs levels that contain as many quantum gates as possible.

[0097] This algorithm can also backtrack the sequence of quantum gates and can maintain: (a) locals, that is, a set of qubits that are expected to be local during the level; (b) locked, that is, a set of locked qubits (qubits that skip some operations); (c) B, that is, the maximum possible number of new branches in this branch of the backtracking; (d) N, that is, the number of quantum gates taken in this level.

[0098] The process of this algorithm can be specifically analyzed according to the following cases:

[0099] · If at least one of the gate qubits or the gate control qubits is locked, the quantum gate must be skipped.

[0100] · Otherwise, if the gate matrix is diagonal, it can also be applied to local and global qubits without imposing any requirements on the qubits.

[0101] · Otherwise, if the application of this quantum gate requires too many qubits to be local, the gate is skipped.

[0102] · Otherwise, if all the gate qubits are already required to be local, the quantum gate can also be applied without adding any requirements.

[0103] · Otherwise, if it is not possible to uniquely determine whether to apply / skip the gate, the algorithm branches on two branches: one branch skips this gate; the other branch applies this gate.

[0104] When the algorithm skips a gate, all its qubits can be locked. When the algorithm decides to apply a non - diagonal gate, all its qubits may need to be local. If all qubits are locked during backtracking, the algorithm may return to the previous level of recursion.

[0105] Some qubits can be intentionally kept local, for example, by pre - filling the set of local qubits before starting the algorithm. This enables other optimizations to be performed in the simulator 110 due to the memory placement layout that can adjust the amplitudes to be swapped.

[0106] The results of the method executed by the device 100 are shown in FIG. 6(a). The device 100 has been tested with 3 global qubits and different numbers of total qubits. According to the permutations between the obtained levels, all global qubits are applied with the exchange of the same number of local qubits.

[0107] In FIG. 6(b), the quantum circuit simulator 110 according to an embodiment of the present invention (i.e., including Figure 1 the shown device 100) is compared with the QuEST simulator, in particular, the QuEST simulator on an 8 - node cluster. The simulator 110 according to an embodiment of the present invention has a performance improvement of one order of magnitude due to the reduction in the number of matrix - vector multiplications. This is because the device 100 executes a cluster - level algorithm / method, and the level scheduling algorithm / method reduces the amount of data transmission.

[0108] Figure 7 The method 700 according to an embodiment of the present invention is shown. The method 700 is used for quantum gate and qubit scheduling in the quantum circuit simulator 110. The method 700 can be executed by Figure 1 the device 100 or by the quantum circuit simulator 110 including the device 100.

[0109] The method includes: Step 701, obtaining a first quantum gate sequence 101; Step 702, generating a second quantum gate sequence 102 by using a greedy algorithm, in particular a greedy algorithm with backtracking, where the second quantum gate sequence 102 is a subsequence of the first quantum gate sequence 101; Step 703, calculating a local qubit set 103a and a global qubit set 103b according to the second quantum gate sequence 102; Step 704, generating a set of quantum gate clusters 104, where each cluster 104 includes a subset of the quantum gates of the second quantum gate sequence 102 merged together by using a greedy algorithm; Step 705, generating a third quantum gate sequence 105 according to the order of the clusters 104, where the third quantum gate sequence 105 includes all the quantum gates in the second quantum gate sequence 102; Step 706, providing the local qubit set 103a and the global qubit set 103b to the quantum circuit simulator 110; Step 707, outputting the third quantum gate sequence 105 to the quantum circuit simulator 110.

[0110] The present invention has been described in connection with different embodiments and implementations taken as examples. However, upon study of the drawings, the present invention, and the independent claims, those skilled in the art will be able to understand and implement other variations when practicing the claimed invention. In the claims as well as in the description, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single element or other unit may fulfill the functions of several entities or items recited in the claims. Stating certain measures in mutually different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation.

Claims

1. A device (100) for a quantum circuit simulator (110), characterized in that, The device (100) is used for: Obtaining a first quantum gate sequence (101); wherein, when executed on the quantum circuit simulator (110), matrix-vector multiplication needs to be performed on the quantum circuit simulator (110); Generating a second quantum gate sequence (102) by using a first greedy algorithm, which is a greedy algorithm with backtracking, and the second quantum gate sequence (102) is a subsequence of the first quantum gate sequence (101); Calculating a local qubit set (103a) and a global qubit set (103b) according to the second quantum gate sequence (102); Generating a set of quantum gate clusters (104) based on the second quantum gate sequence (102), wherein each cluster (104) includes a subset of the quantum gates of the second quantum gate sequence (102) merged together by using a second greedy algorithm, and the second greedy algorithm is a greedy algorithm without backtracking; Generating a third quantum gate sequence (105) according to the order of the clusters (104), and the third quantum gate sequence (105) contains all the quantum gates in the second quantum gate sequence (102); Providing the local qubit set (103a) and the global qubit set (103b) to the quantum circuit simulator (110); Outputting the third quantum gate sequence (105) to the quantum circuit simulator.

2. The device (100) according to claim 1, characterized in that, It is also used for, when generating the set of quantum gate clusters (104): Arranging the clusters (104) including more quantum gates before the clusters (104) including fewer quantum gates in the order of the clusters (104).

3. The device (100) according to claim 1 or 2, characterized in that, It is also used for, when generating the set of quantum gate clusters (104): Generating the clusters (104) according to the maximum possible number of qubits in the clusters (104).

4. The device (100) according to claim 3, characterized in that, It is also used for, when generating the set of quantum gate clusters (104): Selecting all possible combinations of the qubits associated with the second quantum gate sequence (102) one by one according to the maximum possible number of qubits in the clusters (104); Constructing clusters (104) for each combination; Selecting the one with the largest number of quantum gates in the clusters (104).

5. The device (100) according to claim 4, characterized in that, It is also used for, when generating the set of quantum gate clusters (104): Maintaining a locked qubit set; If the matrix representation of the quantum gate is diagonal, including the quantum gate into the cluster (104); If at least one of the qubits on which the quantum gate acts does not belong to the selected qubit combination, skipping the quantum gate; and / or If at least one of the qubits on which the quantum gate acts is in the locked qubit set, skipping the quantum gate; If a quantum gate is skipped, adding all the qubits on which the quantum gate acts to the locked qubit set; Otherwise, including the quantum gate into the cluster (104).

6. The device (100) according to any one of claims 1 to 5, characterized in that, It is also used for, when generating the set of quantum gate clusters (104): Determining the cluster (104) including the largest number of quantum gates; Outputting the quantum gates of the determined cluster (104), in particular inserting the output quantum gates into the third quantum gate sequence (105); Delete the output quantum gate from the second quantum gate sequence (102).

7. The device (100) according to any one of claims 1 to 6, characterized in that, Also used for, when calculating the local qubit set (103a) and the global qubit set (103b): Determine the local qubit set (103a) and / or the global qubit set (103b) respectively according to the maximum number of local qubits and / or the maximum number of global qubits.

8. The device (100) according to any one of claims 1 to 7, characterized in that, Also used for, when generating the second quantum gate sequence (102): Fuse the quantum gates acting on a single qubit with the adjacent quantum gates in the first quantum gate sequence (101) acting on a subset of qubits including the same single qubit.

9. The device (100) according to any one of claims 1 to 8, characterized in that, Also used for, when generating the second quantum gate sequence (102): Include the quantum gates acting on at most the maximum number of local qubits into the second quantum gate sequence (102); If the first quantum gate sequence (101) includes at least one quantum gate acting on a single qubit and another quantum gate acting on the same qubit and at least one other qubit, then include the single-qubit gate and the other multi-qubit gate together into the second quantum gate sequence (102).

10. The device (100) according to any one of claims 1 to 9, characterized in that, Also used for, when generating the second quantum gate sequence (102): Create a branch of the first greedy algorithm, where quantum gates are included into the second quantum gate sequence (102); and / or Create a branch of the first greedy algorithm, where the quantum gates in the first quantum gate sequence (101) are skipped; If quantum gates are included, add all the qubits on which the quantum gates act to the local qubit set (103a); or, If quantum gates are skipped, add all the qubits on which the quantum gates act to the locked qubit set.

11. The device (100) according to claim 10, characterized in that, Also used for, when generating the second quantum gate sequence (102): Create at most the maximum number of branches of the first greedy algorithm.

12. The device (100) according to claim 10 or 11, characterized in that, Also used for, when applying the branches of the first greedy algorithm: Construct the second quantum gate sequence (102) with as many gates as possible, Test each gate in the first quantum gate sequence (101), and according to the test results, skip the gate or include the gate into the second quantum gate sequence (102).

13. The device (100) according to any one of claims 10 to 12, characterized in that, Also used for, when generating the second quantum gate sequence (102): Maintain the locked qubit set; If applying a quantum gate requires more qubits than a predetermined threshold to be local qubits, then skip the quantum gate; and / or If at least one of the qubits on which the quantum gate acts is in the locked qubit set, then skip the quantum gate, If a quantum gate is skipped, add all the qubits on which the quantum gate acts to the locked qubit set.

14. The device (100) according to any one of claims 1 to 13, characterized in that, Also used for, when generating the second quantum gate sequence (102): If the matrix representation of a quantum gate is diagonal, then include the quantum gate into the second quantum gate sequence (102) without adding the qubits on which the quantum gate acts to the local qubit set (103a); and / or If all qubits on which the quantum gates act are already in the local qubit set (103a), include the quantum gates into the second quantum gate sequence (102).

15. The device (100) according to any one of claims 1 to 14, characterized in that, Also used for, when calculating the local qubit set (103a) and the global qubit set (103b): Construct a set of all qubits on which the quantum gates in the first quantum gate sequence (101) act; Include all qubits on which the quantum gates in the second quantum gate sequence (102) act into the local qubit set (103a); Include all qubits that are in the set of all qubits and not in the local qubit set (103a) into the global qubit set (103b).

16. A quantum circuit simulator (110), characterized in that, Include the apparatus (100) according to any one of claims 1 to 15.

17. A method (700) for quantum gate and qubit scheduling in a quantum circuit simulator (110), characterized in that, The method includes: Obtain (701) a first quantum gate sequence (101), wherein, when executed on the quantum circuit simulator (110), matrix-vector multiplication needs to be performed on the quantum circuit simulator (110); Generate (702) a second quantum gate sequence (102) by using a first greedy algorithm, which is a greedy algorithm with backtracking, and the second quantum gate sequence (102) is a subsequence of the first quantum gate sequence (101); Calculate (703) a local qubit set (103a) and a global qubit set (103b) according to the second quantum gate sequence (102); Generate (704) a set of quantum gate clusters (104) based on the second quantum gate sequence (102), wherein each cluster (104) includes a subset of the quantum gates of the second quantum gate sequence (102) merged together by using a second greedy algorithm, and the second greedy algorithm is a greedy algorithm without backtracking; Generate (705) a third quantum gate sequence (105) according to the order of the clusters (104), and the third quantum gate sequence (105) contains all the quantum gates in the second quantum gate sequence (102); Provide (706) the local qubit set (103a) and the global qubit set (103b) to the quantum circuit simulator (110); Output (707) the third quantum gate sequence (105) to the quantum circuit simulator (110).

18. A computer program product, characterized in that, Include program code for controlling the apparatus (100) according to any one of claims 1 to 15, or for performing the method (700) according to claim 17 when implemented on a processor.