Emulating Quantum Circuits on a Computer Using Hierarchical Storage
By partitioning and simulating quantum circuits using tensor slicing and auxiliary memory, the method addresses the computational challenges of simulating large quantum circuits, achieving efficient storage and reduced resource usage.
Patent Information
- Application Number
- CN201980028977.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-05-08
- Filing Date
- 2019-04-15
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2039-04-15
AI Technical Summary
Simulating quantum circuits is computationally challenging due to the exponential growth of quantum state amplitudes with the number of qubits, making it impractical for even supercomputers to handle, and existing methods are limited to relatively small numbers of qubits and low circuit depth.
A method involving tensor slicing and partitioning of quantum circuits into sub-circuits, followed by staged simulation and storage of quantum state tensors in auxiliary memory, allowing for efficient utilization of storage resources and reduced read/write cycles.
Enables the simulation of larger quantum circuits with reduced memory requirements, overcoming the limitations of traditional methods by minimizing storage needs and optimizing read/write operations.
Smart Images

Figure CN112041859B_ABST
Abstract
Description
[0001] Background
[0002] The present invention generally relates to quantum computing and, more particularly, to simulating quantum circuits. Quantum information processing (quantum computing) has the potential to solve certain specific types of mathematical problems that are difficult for conventional machine computing to solve. Quantum computers use qubits (quantum bits) to encode information, where a qubit is the fundamental unit of quantum information. Quantum circuits are based on quantum mechanical phenomena such as qubit superposition and qubit entanglement.
[0003] Although quantum computing has potential benefits, the fabrication of quantum circuits is difficult and expensive, and suffers from various problems such as scaling and quantum decoherence. Thus, quantum circuit simulation, which has been performed using commercially available conventional (non-quantum) computers, including the ability to quantify circuit fidelity and evaluate correctness, performance, and scalability, is based on the computational measurement of the quantum state amplitudes of the circuit. However, since the number of amplitudes grows exponentially with the number of qubits, it becomes very tricky to compute the quantum state amplitudes of the measurement results using the prior art, which is beyond the reach of even powerful supercomputers. In fact, a recent publication, "Characterizing Quantum Supremacy in Near-Term Devices," by Boixo et al., Articles, Nature Physics (2018), doi:10.1038 / s41567-018-0124-x, states that: "The output amplitudes of circuits with 7×7 qubits and a depth of approximately 40 cycles are currently infeasible." Thus, quantum circuit simulation is limited to circuits with a relatively small number of qubits and a low circuit depth (circuit depth refers to the number of layers into which the circuit can be divided at most with one gate acting on a given qubit at any given layer).
[0004] Summary
[0005] The following provides an overview to provide a basic understanding of one or more embodiments of the present invention. This overview is not intended to identify key or important elements nor to delineate any scope of a particular embodiment of the present invention or any scope of the claims. Its sole purpose is to present concepts in a simplified form as a prelude to the more detailed description that is presented later.
[0006] According to an embodiment of the present invention, a system includes a partitioning component that partitions an input quantum circuit including a machine-readable specification into sub-circuits based on at least two sets of determined tensor slice qubit partitions, wherein the sub-circuits have associated qubit sets for tensor slicing. A simulation component simulates the input quantum circuit in stages into a simulated quantum state tensor based on the sub-circuits, one stage for each sub-circuit, wherein the qubit sets associated with the sub-circuits are used to partition the simulated quantum state tensor of the input quantum circuit into quantum state tensor slices, and the quantum gates in the sub-circuits are used to update the quantum state tensor slices into updated tensor slices. A read / write component stores the updated quantum state tensor slices of the simulated quantum state tensor as micro-slices in an auxiliary memory.
[0007] The read / write component may write the micro-slices in a size that spans at least two disk sectors of the auxiliary memory. The read / write component may further retrieve the updated quantum state tensor slices from the auxiliary memory, and the simulation component may process another sub-circuit as other sub-circuit tensors and update the other sub-circuit tensors with the updated quantum state tensor slices retrieved from the auxiliary memory into further updated quantum state tensor slices.
[0008] According to an embodiment of the present invention, a computer-implemented method includes: processing an input quantum circuit including a machine-readable quantum circuit specification. The processing includes partitioning the input quantum circuit into a set of sub-circuits based on at least two sets of qubits identified for tensor slicing, wherein the set of sub-circuits has associated qubit sets for tensor slicing. Simulating the input quantum circuit in stages into a simulated quantum state tensor based on the set of sub-circuits, one stage for each sub-circuit, wherein the qubit sets associated with the sub-circuits are used to partition the simulated quantum state tensor of the input quantum circuit into quantum state tensor slices and the quantum gates in the sub-circuits are used to update the quantum state tensor slices into updated quantum state tensor slices. The method includes storing the updated quantum state tensor slices of the simulated quantum state tensor as micro-slices in an auxiliary memory.
[0009] According to one or more embodiments, a computer program product for simulating a quantum circuit is provided. The computer program product includes an input quantum circuit that includes a machine-readable specification of the quantum circuit. The product includes a computer-readable storage medium and program instructions stored on the storage medium. The program instructions are executable by a processing component to cause the processor to partition the input quantum circuit into a set of sub-circuits based on at least two sets of qubits identified for tensor slicing, where the set of sub-circuits has associated qubit sets for tensor slicing. Further instructions stage-simulate a simulated quantum state tensor into a simulated quantum state tensor based on the set of sub-circuits, one stage per sub-circuit, where the qubit set associated with a sub-circuit is used to partition the simulated quantum state tensor of the input quantum circuit into quantum state tensor slices and the quantum gates in the sub-circuit are used to update the quantum state tensor slices into updated quantum state tensor slices. The method includes storing the updated quantum state tensor slices of the simulated quantum state tensor as micro-slices in an auxiliary memory.
[0010] Further instructions can include processing another sub-circuit into other sub-circuit tensors, retrieving the updated quantum state tensor slices from the auxiliary memory, and updating the other sub-circuit tensors with the retrieved updated quantum state tensor slices from the auxiliary memory into further updated quantum state tensor slices. Storing the updated quantum state tensor slices of the simulated quantum state tensor as micro-slices in the auxiliary memory can include storing the micro-slices in at least two disk sectors of the auxiliary memory.
[0011] According to one or more embodiments, a computer-implemented method, the method being performed by a device operatively coupled to at least two processors, includes: simulating an input quantum circuit that includes a machine-readable specification of a quantum circuit. The simulation includes partitioning the input quantum circuit into meta sub-circuits based on at least two sets of qubits identified for tensor slicing, where at least one of the meta sub-circuits exceeds the memory of a single processor. The method includes sub-partitioning the meta sub-circuits into sub-sub-circuits that fit the memory of a single processor, computing tensors of the sub-sub-circuits, and contracting the tensors into tensor slices by tensor slicing.
[0012] The method may further include using the gates of another sub-circuit for the tensor slice to obtain an updated tensor slice representing the quantum state data and storing the updated tensor slice in an auxiliary memory. The auxiliary memory may include one or more disk devices, and the method may further include organizing the quantum state data into micro-slices, where the size of the micro-slices spans at least two disk sectors of the auxiliary memory.
[0013] According to one or more embodiments, there is provided a quantum computer program product for simulating a quantum circuit of an input quantum circuit including a machine-readable specification. The product includes a computer-readable storage medium and program instructions stored on the storage medium. The program instructions are executable by a processing component to cause a processor to partition the input quantum circuit into meta-sub-circuits based on at least two sets of qubits identified for tensor slicing, where at least one of the meta-sub-circuits exceeds the memory of a single processor. The program instructions sub-partition the meta-sub-circuits into sub-sub-circuits that fit the memory of a single processor, calculate tensors for the sub-sub-circuits, and contract the tensors into tensor slices by tensor slicing. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a block diagram of components for implementing simulation of a quantum circuit based on the techniques described herein, according to an embodiment of the present invention.
[0015] Figure 2 is a flowchart for simulating a quantum circuit based on the techniques described herein, according to an embodiment of the present invention.
[0016] Figure 3 is a schematic diagram of a meta-slice according to an embodiment of the present invention, the meta-slice including a tensor slice of quantum state data that spans the memory of a single processor and having a corresponding meta-sub-circuit that is divided into sub-sub-circuits suitable for the memory of a single processor for further tensor slicing and processing.
[0017] Figure 4 is a schematic diagram of a 49-qubit quantum circuit with a depth of 55, which generally belongs to a class of randomly generated quantum circuits, according to an embodiment of the present invention.
[0018] Figure 5 is a schematic diagram of a 49-qubit quantum circuit with a depth of 55, which is partitioned for quantum circuit simulation, according to an embodiment of the present invention.
[0019] Figure 6is a schematic diagram of a 49-qubit quantum circuit of depth 83, partitioned for quantum circuit simulation, according to an embodiment of the present invention.
[0020] Figure 7 is a schematic diagram of a 49-qubit quantum circuit of depth 111 according to an embodiment of the present invention, which is partitioned for quantum circuit simulation.
[0021] Figure 8 is a flow chart of a quantum circuit partitioning operation according to an embodiment of the present invention.
[0022] Figure 9 is a flow chart of a set of quantum circuit partitioning operations according to an embodiment of the present invention.
[0023] Figure 10 According to an embodiment of the present invention Figure 9 Flowchart of additional details of the quantum circuit partitioning operation.
[0024] Figures 11A - 11C is a diagram of pre-partition optimization of a quantum circuit according to an embodiment of the present invention.
[0025] Figure 12A is a flow chart of quantum circuit execution operations according to an embodiment of the present invention.
[0026] Figure 12B According to an embodiment of the present invention Figure 12A A flowchart with additional details of the operations performed by the quantum circuit.
[0027] Figure 13 is a block diagram of a system for implementing various aspects of the techniques described herein.
[0028] Figure 14 is a block diagram of a computer-implemented method for implementing various aspects of the techniques described herein.
[0029] Figure 15 is a block diagram of another computer-implemented method for implementing various aspects of the techniques described herein.
[0030] Figure 16 is a block diagram of an operating environment in which one or more embodiments of the invention described herein may be facilitated. DETAILED DESCRIPTION
[0031] The following detailed description is illustrative only and is not intended to limit the embodiments of the invention and / or the application or uses of the embodiments of the invention. In addition, there is no intention to be bound by any express or implied information provided in the previous sections or the "Detailed Description" section.
[0032] Reference is now made to the accompanying drawings to describe one or more embodiments of the present invention, wherein like reference numerals are used throughout to refer to like elements. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a more thorough understanding of the embodiments of the present invention. However, it is evident that the embodiments of the present invention may be practiced without these specific details in various instances.
[0033] In addition, it should be understood that the present invention will be described in accordance with a given illustrative architecture; however, other architectures, structures, and / or operations may be varied within the scope of the present invention.
[0034] It should be understood that, as described herein, an "input quantum circuit" as actually referred to by traditional computer simulation is a machine-readable specification of a quantum circuit. It will also be understood that when an element such as a qubit is said to be coupled to or connected to another element, it may be directly coupled to or connected to the other element, or indirectly coupled to or connected to the other element by means of: there may be one or more intermediate elements. Conversely, when and only when an element is said to be "directly" coupled or connected, there are no intermediate elements, that is, when and only when an element is said to be "directly connected" or "directly coupled to" another element, there are no intermediate elements.
[0035] In at least one embodiment of the present invention, references in the specification to "one embodiment" or "an embodiment" of the present principle and other variations thereof mean including the specific features, structures, characteristics, etc. described in connection with that embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" and any other variations thereof occurring throughout the specification do not necessarily all refer to the same embodiment.
[0036] One problem with existing quantum circuit simulations is that simulating a quantum circuit requires a large amount of memory. Due to memory limitations, it is generally considered impossible or at least impractical for more than 49 qubits. More specifically, a 7x7 qubit array requires 249 stored values to represent the possible states of 49 qubits. When using two eight-byte floating-point values to store (complex) qubit state information, 8 petabytes of storage space are required; a 56-qubit circuit requires 1 Exabyte of storage space.
[0037] Described herein is the use of tensors to represent quantum states in quantum simulation, and a tensor slicing method that enables the quantum state of a circuit to be computed in slices, rather than having to implement the entire quantum state in main memory. Further described herein is the organization of computations into relatively large tensor slices. Still further, as will be appreciated, auxiliary memory is used after computing the quantum state for a reasonable circuit depth in-memory, where the use of auxiliary memory is effective because of the relatively large tensor slices. In addition, the present disclosure describes the effective utilization of the auxiliary memory based on a combination of choices in quantum circuit partitioning and data organization, which simultaneously optimizes the number of global read / write cycles to the auxiliary memory and the efficiency of performing reads and writes.
[0038] Note that the "auxiliary memory" (or hierarchical storage) as described herein includes at least one level of auxiliary memory. It will be appreciated that the framework described herein can be generalized to address one (second) or more levels of storage. Any such storage level can have its own capacity and access characteristics, which can be incorporated into the simulation computations and storage allocations. Thus, the techniques described herein are not limited to any physical hardware or type of storage device, but can use any device or combination of devices to organize data into micro-slices and allow data to be transferred in blocks between hardware devices.
[0039] As used herein, the terms "simulation" and "execution" (e.g., with respect to a quantum sub-circuit) are synonymous in nature, because (contrary to many simulations that are only approximations of real things) the simulation methods disclosed herein produce real, computationally effective results. As used herein, phrases such as "index variable" refer to any type of memory addressing or indexing that results in the access of desired data and / or memory locations.
[0040] It is recognized from at least some embodiments of the present invention disclosed herein that less memory can be used to simulate a quantum circuit than is typically assumed. For example, at least some embodiments of the present invention disclosed herein reduce the storage space required to represent and simulate a quantum circuit including X qubits from 2 X complex values to a fraction of that amount. More specifically, at least some embodiments of the present invention disclosed herein recognize that a quantum circuit can be partitioned into sub-circuits by having entangled quantum circuit elements in two or more sub-circuits. At least some embodiments of the present invention disclosed herein further recognize that the simulation results of some sub-circuits can be computed in slices, and by doing so, contrary to conventional expectations, the working memory required to simulate the entire quantum circuit can be significantly reduced.
[0041] Reference is now made to the accompanying drawings, in which like numerals represent like or similar elements, Figure 1FIG. 100 shows a general system in which the techniques described herein may be implemented. In Figure 1 this example, a processor 102, a memory 104, and a secondary memory 106 are provided in the form of a conventional (non-quantum) computer, which may be a supercomputer. As described herein, a partitioning component 108 processes an input quantum circuit (i.e., a machine-readable specification for simulating a quantum circuit) to partition the input quantum circuit into sub-circuits that are represented as tensors (arrays) in the memory 104.
[0042] A tensor slicing component 110 slices the tensors in the memory 104, and a simulation component 112 simulates the input quantum circuit in stages, with each sub-circuit being a stage, e.g., applying gates, to provide a simulated quantum state tensor for the input quantum state circuit. A read / write (out / in) secondary memory component 114 stores slices of the simulated quantum state tensor for the input quantum state circuit into the secondary memory, retrieves slices of the simulated quantum state tensor for the input quantum state circuit from the secondary memory 106. And organizes the secondary memory as micro-slices according to the combined qubits for tensor slicing, where slices of a given set of qubits are stored and retrieved as multiple micro-slices.
[0043] Figure 2 summarizes Figure 1 an example operation of the components. Operation 202 represents receiving an input quantum circuit (of a machine-readable specification). Operation 204 analyzes the input quantum circuit to identify at least two sets of qubits that will be used for tensor slicing. Operation 206 partitions the input quantum circuit into multiple sub-circuit sets based on the sets of qubits, where any resulting sub-circuit thus has a set of associated qubits that will be used for tensor slicing.
[0044] Operation 208 represents simulating the input quantum circuit in stages, with each sub-circuit being a stage. A set of qubits associated with a sub-circuit is used to partition the simulated quantum state tensor of the input quantum state circuit into slices, and quantum gates in the sub-circuit are applied to update the slices.
[0045] Operation 210 represents storing slices of the simulated quantum state tensor of the input quantum state circuit into the secondary memory 106 ( Figure 1 ). Operation 212 represents retrieving slices of the simulated quantum state tensor for the input quantum state circuit from the secondary memory 106. Operation 214 represents organizing the secondary memory into micro-slices according to the combined set of qubits for tensor slicing, where slices for a given set of qubits are stored and retrieved as a set of micro-slices.
[0046] As described herein, to minimize the cost of auxiliary storage, computations are organized into fairly large tensor slices, called "meta" slices. Generally, the larger the tensor slice, the fewer read / write cycles to auxiliary storage. As Figure 3 shown, a "meta" slice 330 needs to be able to fit into the aggregate main memory 104, but does not necessarily need to fit into the memory of a single processor (processing node), e.g., the memory 304(1) of processor 1 for processor 1 302(1). When using the gates in each "meta" sub-circuit, each sub-sub-circuit is used for the inter-processor slice to minimize communication applications. The corresponding "meta" sub-circuit 332 is then sub-partitioned (e.g., by Figure 1 the partitioning component 108) into sub-sub-circuits 334(1) to 344(N) such that a known tensor slicing strategy ( Figure 1 a part 334(1) to 344(N) of the tensor slicing component 110 in an embodiment of
[0047] can be applied. For example, a disclosed method (Thomas and Damian S Steiger, "0.5 Petabyte emulation of 45 qubit quantum circuits", arXiv preprint, arXiv:1704.01127, (2017)) can be applied in the computation. In this reference, the circuit is partitioned so that all gates in a sub-circuit can be applied to update the quantum state tensor based on each slice without passing quantum state information between processing nodes. In this reference, with respect to circuit partitioning and tensor slicing, "global" qubits are used to index across processing nodes, which correspond to the tensor indices to be sliced, while "local" qubits correspond to the tensor indices used to index into the tensor slices stored on each processing node. Zero communication updates are possible when all non-diagonal gates in a sub-circuit are applied only to "local" qubits. In fact, the circuit is partitioned by choosing different subsets of "local" qubits and analyzing which gates can be applied to zero bits to produce the corresponding sub-circuits. During the simulation, communication between processing nodes occurs only when the simulation switches from one sub-circuit to another. During these communication phases, the memory layout of the quantum state tensor is reorganized so that a new subset of indices is used to index "globally" across processing nodes according to the requirements for simulating the next sub-circuit, rather than "locally" indexing within the memory of a single node.
[0048] In contrast, the present description is of circuit partitioning, where the resulting tensor is stored in auxiliary memory and the quantum state is updated on a per sub-circuit basis by loading "meta" slices of the resulting tensor into available aggregate memory, applying gates of the sub-circuit to these meta slices, and then writing the updated meta slices back to the auxiliary storage. Nevertheless, by further partitioning each sub-circuit into sub-sub-circuits, the methods disclosed above can be utilized to some extent to fit corresponding sub-slices of the meta slices to the memory of individual processors..
[0049] Reference is made herein Figures 8 to 12B to the circuit decomposition method described, which can be used to bootstrap the entire process by calculating the quantum state to a reasonable depth entirely in memory before the first write to the auxiliary memory. For example, Figure 5 shows how the partitioning can be extended from simulating a 49-qubit circuit depth of 27 (e.g., as described in U.S. Patent Application Serial No. 15 / 713,323 filed on September 22, 2017) to a circuit depth of 55 by one memory read / write operation. As described below Figure 6 shows how the circuit depth of a simulated 49-qubit can be extended to a circuit depth of 83 by two memory read / write cycles, Figure 7 shows how the partitioning of a simulated 49-qubit can be extended to a circuit depth of 27, to a circuit depth of 111, and so on up to any depth by three memory read / write cycles.
[0050] Figure 4 The example of depicts a 49-qubit, depth 55 quantum circuit, which generally belongs to a class of randomly generated quantum circuits. In this example, the qubits are arranged in a 7x7 array. A given 7x7 square represents a time slice, and gates are applied to the qubits, e.g., the gates applied at each level (the levels are shown on the left boundary, e.g., 0, 9, 17…49). Level 0 is the initialization level. Pairs of black dots (each pair surrounded by a gray background) represent controlled-Z (CZ) gates. Other gray dots with a white background represent various single-qubit gates.
[0051] Figure 5 depicts partitioning and dividing the Figure 4 circuit shown into multiple sub-circuits that enable the quantum circuit to be simulated on a currently available computing system. It will be understood that multiple levels of granularity (e.g., tensor meta slices to micro slices) allow simulation to any circuit depth described herein.
[0052] In Figure 5In the partitioning scheme, the given square represents a time slice of a 7x7 qubit array over time, and the gates applied to the qubits. For example, the gates applied at each level (these levels are shown on the left border, e.g., 0, 9, 17…49). Level 0 is the initialization level.
[0053] In this example, the qubits (corresponding to the gates) of the quantum circuit are labeled with the numbers 1-4 Figure 5 . The qubit identified with the number "1" (above the dark gray background) represents a gate belonging to one sub-circuit, the qubit identified with the number "2" (with a very light gray background) represents a gate belonging to another sub-circuit, the qubit identified with the number "3" (with a medium to dark gray background) represents a gate belonging to a third sub-circuit, and the qubit identified with the number "4" (with a light gray to medium gray background) represents a gate belonging to a fourth sub-circuit. Sub-circuits (e.g., identified by 1 and 2) can be entangled, where the bridging gate is a controlled-Z (CZ) gate (where one qubit determines whether the Pauli-Z operation is applied to another qubit). For example, as Figure 5 shown, these CZ gates can be assigned to either sub-circuit without affecting the ability of the simulation circuit, and these CZ gates can be arbitrarily assigned to the bottom sub-circuit.
[0054] The numbering of the qubits (combined with the shading) shows a way to distribute the gates between the sub-circuits for tensor contraction and slicing. Generally, the process first operates to a depth of 27, and then cyclically applies the values of the last seven rows of the qubits being computed. In Figure 5 the example, to extend beyond level 27, auxiliary memory is used. That is, the partitioning is applied to the qubits / gates identified as 3, and then to the qubits / gates identified as 4.
[0055] More specifically, in Figure 5 the example, the computations for sub-tensor-circuits 1 and 2 are done in a known way. By slicing the bottom 7 qubits, the two resulting tensors are contracted into 64TB slices. In the first stage, i.e., stage 1 (for emphasis, the last row of the qubits corresponding to the qubits labeled 2 is labeled row 541), the bottom row of the qubits corresponding to the qubits labeled 2 is sliced. In the second stage, i.e., stage 2, the top row of the qubits corresponding to the qubits labeled 3 is sliced (for emphasis, the top row of the qubits labeled 3 is labeled row 542).
[0056] The gates in sub-circuit 3 are applied to each slice, except that now sub-circuit 3 extends beyond depth 27, up to depth 55, and the updated slices are sent to the memory. For sub-circuit 4, the 64TB slices are retrieved from the memory, the top 7 qubits are sliced, and the updated slices are sent back to the memory.
[0057] Regarding Figure 6 and 7 , in the case of sub - circuit 5, the bottom 7 qubits are sliced again to retrieve a 64TB slice from the memory, and then the updated slice is sent back to the memory. In Figure 6 (circuit depth 83), in the third stage (Stage 3), the qubits in the bottom row are also sliced (the last row is labeled 643). For sub - circuit 6, the top 7 qubits are sliced, a 64TB slice is retrieved from the storage device, and the updated slice is either processed or sent back to the storage again. In Figure 7 (circuit depth 111), in the third stage, Stage 3, the bottom row of qubits is sliced, and in the fourth stage (Stage 4), the top row of the fourth row of qubits is sliced (the top row is labeled 744).
[0058] Thus, the techniques described herein provide an expansion of the partitioning of a 7×7 qubit, depth - 27 random circuit to a greater depth by utilizing an auxiliary memory while minimizing the number of read / write cycles. Using this partitioning scheme, a 7×7 qubit, depth - 55 random circuit generated in a known manner can be emulated using only one read - write cycle ( Figure 5 ), a 7×7 qubit, depth - 83 random circuit can be emulated. Using only two read - write cycles ( Figure 6 ), and so on, up to at least depth 111 ( Figure 7 ).
[0059] In the case of a depth - 55 circuit, first the tensors of sub - circuits 1 and 2 are computed, and then the resulting tensors are contracted once, one slice at a time. Before performing the contraction of the tensor of sub - circuit 1, slicing is performed by iterating over the possible values of qubits 43 - 49 and slicing the tensor of sub - circuit 2 into these values. Then, the gates belonging to Figure 5 in sub - circuit 3 are applied to the contraction result, and the resulting updated slice is sent to the auxiliary memory. In this example, the process of slicing, contracting, applying gates, and sending the result to the auxiliary memory is repeated 128 times, once for each of the 128 possible values of qubits 43 - 49.
[0060] To complete the depth - 55 emulation, the gates belonging to Figure 5The gates of sub-circuit 4 up to depth 55 are applied to the intermediate results transferred to the auxiliary memory. These gate applications can also be executed in slices, this time slicing by qubits 1-7. To ensure an efficient retrieval from the auxiliary memory, the data can be organized in the auxiliary memory as 214 micro-slices indexed by the values of qubits 1-7 and 43-49, where each micro-slice contains 235 complex amplitudes corresponding to qubits 8-42. Thus, during the stage of applying gates in sub-circuit 3 discussed above, for the 128 values of the sliced qubits 43-49, 128 micro-slices are written to the auxiliary memory corresponding to the 128 possible values of qubits 1-7. During the current stage of applying gates in sub-circuit 4, for each of the 128 values of the sliced qubits 1-7, 128 micro-slices are read from the auxiliary memory, corresponding to the 128 possible values of qubits 43-49. Once the amplitudes of these 128 micro-slices are loaded into the memory, the gates in sub-circuit 4 can then be applied. Similarly, once these gates are applied, each updated slice can be written back to the memory. Alternatively, the final amplitudes can be processed in memory on a per-segment basis. This can be achieved by considering the physical distribution of the auxiliary storage elements within the computer cluster employed and its characteristics Figure 1 The read / write (out / in) auxiliary storage component 114 shown in Figure 1 is utilized to optimize the reading and writing of micro-slices. The communication network of the cluster is exploited, and the opportunities for parallelism during the microcomputer read / write process and during the microcomputer communication between the processing nodes and the storage elements are utilized. For example, each micro-slice can be stored in a distributed manner across multiple physical disk drives to benefit from parallelism, and the data is striped across sequential sectors on each disk to optimize the data transfer rate.
[0061] To continue the simulation to depth 83 ( Figure 6 ), the gates in sub-circuit 4 are applied to depth 83, and the updated slices are written back to the auxiliary memory. Once this process is completed for all slices in the simulation of sub-circuit 4, the simulation continues to Figure 6 the sub-circuit 5 shown in Figure 6 to simulate the remaining gates at depth 83. Except for this time, the process proceeds in the same manner as for sub-circuit 4, where for each of the 128 values of the qubits 43-49 that are now being sliced, 128 micro-slices are loaded from the auxiliary memory, corresponding to the 128 possible values of qubits 1-7.
[0062] Similarly, to continue the simulation to depth 111 ( Figure 7 ), the gates in sub-circuit 5 are applied to depth 111, and the updated slices are written back to the auxiliary memory. Once this process is completed for all slices in the simulation of sub-circuit 5, the simulation continues to Figure 7The subcircuit 6 shown emulates the remaining gates with a simulation depth of 111. This time, except that the process proceeds in the same manner as for subcircuit 5 this time, for each of the 128 values of qubits 1-7 that are now being sliced, 128 micro-slice logic files corresponding to the 128 possible values of qubits 43-49 are loaded from the auxiliary memory.
[0063] According to one or more embodiments of the present invention, generally, minimizing the number of global read / write cycles does not guarantee efficient utilization of the auxiliary memory. More specifically, data is typically read and written in blocks (disk sectors) when using the auxiliary memory, so updating a single data item in a block requires reading and rewriting the entire block. Therefore, the total amount of reading and writing is determined by the number of blocks that must be updated in the global read / write cycle. The selection in the circuit partitioning can be combined with the selection in the data organization to optimize both the number of global read / write cycles and the efficiency of performing read / writes simultaneously.
[0064] The quantum state data stored on the disk is organized into "micro" slices. The union of all qubits in a slice at a given stage constructs a "meta" slice for slicing during read and write operations on the disk, thereby organizing the data into "micro" slices. For example, in the example given above, qubits 1-7 are sliced at some stages, while qubits 43-49 are sliced at other stages, resulting in different "meta" slices at each stage. Then, the union of the qubits in these slices (i.e., 1 qubit, 7 qubits, 43 qubits to 49 qubits) is used to create "micro" slices. Thus, each "meta" slice in the computing stage corresponds to multiple "micro" slices. Therefore, the reading and writing of "meta" slices become atomic at the "micro" slice level, and thus redundant reading and writing of "micro" slices do not occur.
[0065] Non-redundant disk access is achieved by ensuring the size of each "micro" slice spans multiple disk sectors. The characteristics of the file system also affect the optimal size of the "micro" slice. Therefore, the optimization process for selecting which qubits to slice for circuit partitioning can be considered to produce appropriately sized "micro" slices at each stage.
[0066] Refer to Figure 9 , the quantum circuit partitioning method 900 starts from initialization (operation 902) with a set of global qubits for which the micro-slice index is an empty set. Subsequent operations will add to this set of global qubits. The number of such qubits determines the size of the micro-slices on the auxiliary storage. Therefore, a limit is set on the number of global qubits used to index the micro-slices so that the micro-slices are large enough to enable efficient auxiliary storage access.
[0067] Method 900 proceeds to identify and select (operation 904) a sub - circuit that is consistent with the current set of global qubits used to index the micro - slices. Consistency in this context means that the global qubits associated with the selected sub - circuit, when added to the current set of global qubits used to index the micro - slices, do not cause the resulting set of global qubits used to index the micro - slices to exceed the maximum number of such qubits required to achieve efficient auxiliary memory access. After operation 904, the global qubits of the selected sub - circuit are added (operation 906) to the set of global qubits that will be used to index the micro - slices, and then the gates in the selected sub - circuit are removed (operation 908) from the input circuit. Then the updated input circuit is examined (operation 910) to determine if any remaining gates are present in the input circuit, and if so, the quantum circuit partitioning method 900 returns to operation 904 to identify and select the next sub - circuit. If no gates remain, the circuit partitioning terminates by determining (operation 912) the execution order of the identified and selected sub - circuits.
[0068] Figure 10 An example of operation 904 is depicted. The example includes computing (operation 1002) the set of local qubits on which each gate in the current input circuit depends (for application), identifying (operation 1004) the local qubits that are consistent with the current set of global qubits used to index the micro - slices and maximize the number of applicable gates in the available set memory, and identifying and selecting (operation 1006) the given identified local qubits. In operation 1004, consistency means that adding the global qubits that are not included in the identified local qubits to the current set of global qubits used to index the micro - slices does not cause the resulting set of global qubits used to index the micro - slices to exceed the maximum number of such qubits required to achieve efficient auxiliary memory access.
[0069] Figure 9 The example of operation 904 can also include a circuit decomposition method to guide the whole process by fully computing the quantum state to a reasonable depth - before writing to secondary storage for the first time, which is described herein with reference to Figures 8 to 12B and in U.S. Patent Application Serial No. 15 / 713323.
[0070] Referring to Figure 8 Quantum circuit partitioning method 800 includes creating (operation 802) sub - circuits for each initial - stage qubit, adding (operation 804) subsequent - stage gates to the sub - circuits, and determining (operation 806) if all relevant gates have been assigned, and if so, determining (operation 808) the execution order.
[0071] Operation 810 represents selecting a bridging gate before all relevant gates are allocated. Operation 812 represents determining whether to entangle the bridged sub-circuit and, if not, closing (operation 814) the bridged sub-circuit, or if so, adding (operation 816) the bridging gate to one of the entangled sub-circuits. Quantum circuit operation 800 enables partitioning of a quantum circuit in a way that can reduce or minimize the resources required to simulate the quantum circuit using a conventional computer and associated memory.
[0072] Creating (operation 802) sub-circuits for each initial-stage qubit can include: determining the number of qubits required for the quantum circuit (at least initially); and initializing a data structure for defining sub-circuits for each qubit required for the quantum circuit. Adding (operation 804) subsequent-stage gates to a sub-circuit can include adding unallocated next-stage gates that require inputs only from qubits currently assigned to that sub-circuit. The addition operation can also include determining whether all remaining unallocated gates assigned to a sub-circuit are diagonal-unitary gates, and if so, closing the sub-circuit, creating a new sub-circuit with the same assigned qubits as the closed sub-circuit, and marking this new sub-circuit as created for the purpose of slicing qubits for which all remaining unallocated gates are diagonal-unitary gates. Closing a sub-circuit can include marking the sub-circuit as complete to prevent adding other gates to the sub-circuit.
[0073] Determining (operation 806) whether all gates have been allocated can include determining whether an unallocated gate count or some other indicator indicates that all gates have been allocated to sub-circuits. Selecting (operation 810) a bridging gate can include selecting an unallocated next-stage gate that requires inputs from at least one qubit currently assigned to a sub-circuit and one or more qubits currently not assigned to the sub-circuit. In other words, a bridging gate is a gate that requires inputs from multiple currently created sub-circuits.
[0074] Determining (operation 812) whether to entangle the bridged sub-circuit can include estimating the resource costs of alternatives in which a decision is made whether to entangle or not entangle the bridged sub-circuit, comparing the resource costs associated with each decision, and then selecting the alternative with the lowest cost. Closing (operation 814) the bridged sub-circuit can include marking the sub-circuit as complete to prevent adding other gates to the bridged sub-circuit. Assigning the bridging gate to a new sub-circuit (also part of operation 814) can include creating a new sub-circuit, assigning the qubits of the bridged sub-circuit that are closed for this new sub-circuit, and assigning the bridging gate to this new sub-circuit.
[0075] Adding a bridging gate to one of the entangled subcircuits (operation 816) may include adding the bridging gate to the list of gates included in the subcircuit. This addition operation may also include first replacing the bridging gate with an equivalent combination of gates, where the new bridging gate becomes a diagonal singleton. An example is replacing a CNOT gate ( Figure 11A ) with a combination of a CZ gate and a Hadamard gate, as Figure 11B shown. As Figure 11C shown, this may replace the CNOT gate with such an equivalent combination of gates. After making such a replacement, the new bridging gate will be assigned to a bridging subcircuit, and any single-qubit gates that may be introduced in this rewrite may be assigned to the subcircuits according to the qubits assigned to it. These subcircuits. For example, in the specific case of replacing a CNOT gate with a CZ gate, the introduced Hadamard gate is assigned to the subcircuit to which the corresponding qubit is assigned, in accordance with the rules for adding subsequent stage gates to it. The subcircuits described above in connection with operation 804. On the other hand, the CZ gate may be assigned to any one of the entangled subcircuits.
[0076] Note that Figures 11A - 11C is an exemplary schematic diagram depicting a specific example of optimizing a quantum circuit. In some embodiments of the present invention, gate replacement (also known as circuit rewriting) may be performed at various points within the techniques disclosed herein. For example, circuit rewrites such as the rewrite shown as Figure 8 may be performed in conjunction with step 816 ( Figure 11B ). More specifically, replacing a bridging CNOT gate with an equivalent configuration of a CZ gate and a Hadamard gate may reduce the number of entanglement indices introduced when entangling the corresponding subcircuits, and thus has the effect of reducing the amount of memory required to simulate the generated entangled subcircuits.
[0077] Thus, Figure 8 shows an example flowchart that may be used, for example, to find a circuit partitioning that minimizes the number of floating-point operations for each of the above-described usage cases and other possible usage cases of the technology described herein. Examples include depth-first search, breadth-first search, iterative deepening depth-first search, Dijkstra's algorithm, and A * search.
[0078] For example, implementing the operation shown as Figure 8 as a depth-first recursive optimization process may involve introducing loops at decision points to loop through possible decision choices, and then recursively calling forward from these points within these loops Figure 8The operations shown are performed to achieve possible selections and return, at the end of the loop, the selection that optimizes the required cost metric. These decision points can include operation 810 for selecting assignable bridging gates, operation 812 for selecting whether to entangle subcircuits, operation 816 for assigning bridging gates to subcircuits, and operation 808 for determining the execution order of subcircuits. Instead, some decisions can be made at these points by applying rules of thumb, while other decisions can be included as part of a depth-first search. The required resource cost metric to be minimized can include the maximum memory requirement, the total number of floating-point operations for calculating amplitudes, and / or the total number of floating-point operations for calculating a single amplitude. Conditional tests can also be introduced to discard a selection if a desired constraint is violated. The required constraints can include keeping the total memory requirement of the simulation within a specified limit. They can also include a limit on the total running time consumed by the depth-first process itself. Depth-first processing can be implemented to record the current best set of selections found so far according to the required resource cost metric, so that the depth-first processing can be terminated before a full search is performed (e.g., at runtime). Then, benefits can also be obtained by continuing from the depth-first processing that has been performed.
[0079] The depth-first search process effectively generates a tree of possible sequences of decision selections and the resulting circuit partitioning. The breadth-first search explores the tree one stage at a time. The breadth-first search can be implemented as an iterative deepening depth-first search, where a limit is imposed on the number of decisions made, and once that limit is exceeded, a branch of the search tree is discarded. A limit can be set on the depth of the search, or the number of times an entanglement selection can be made can be chosen at operation 812.
[0080] Figure 12A is a flowchart of the execution of a quantum circuit according to at least one embodiment of the invention set forth herein. The quantum circuit execution operations include: receiving (operation 1202) an ordered set of quantum subcircuits; assigning (operation 1204) different index variables; propagating (operation 1206) the index variables; and executing (operation 1208) each quantum subcircuit. The illustrated quantum circuit operations can enable the execution of a quantum circuit partitioned into subcircuits.
[0081] Receiving (operation 1202) an ordered set of quantum subcircuits can include receiving an ordered list of pointers to objects or data structures that define each quantum subcircuit including the gates contained therein. In one embodiment of the invention, the definition of each quantum subcircuit is substantially a subcircuit execution plan similar to Figure 8 the execution plan shown. Assigning (operation 1204) different index variables can include assigning different index variables to initial values. The state of each qubit and the output of each non-diagonal single gate.
[0082] The propagation (operation 1206) index variable may include iteratively propagating from input to output the index variables of each diagonal single gate. Executing (operation 1208) each quantum sub-circuit may include executing each quantum sub-circuit and combining the generated results in a specified order.
[0083] Figure 12B is depicted Figure 12A is a flowchart of an example of executing operation 1208 depicted therein, where the sub-circuit and combined results may include constructing (operation 1210) the product of tensors and executing (operation 1212) a sum.
[0084] Constructing (operation 1210) the product of tensors may include: in Figure 12A operations 1204 and 1206, identifying the gates belonging to the sub-circuit and the index variables assigned to those gates, assigning these index variables as subscripts to the corresponding tensors for the gates in the sub-circuit, and assembling the tensors of the sub-circuit into a product arranged in input-to-output order. Constructing (operation 1210) the product of tensors may also include assembling the tensors corresponding to the simulation results of the sub-circuit into the product of tensors according to the execution order arrived at by a determination operation (e.g., Figure 8 operation 808).
[0085] Executing (operation 1212) the sum may include calculating the product of the tensors of the sub-circuit in the input-to-output order determined in the above construction operation 1210 and performing a sum on the index variables inside these sub-circuits as they are encountered in the determined input-to-output order. Performing a sum (operation 1212) on the index variables inside the entire circuit may include: when calculating the product of the tensors for combining the simulation results of the sub-circuits determined in the above construction operation 1210, performing a sum on these index variables.
[0086] If there are no gates to simulate in a qubit's circuit or all the remaining gates of that qubit are diagonal single gates, a for loop may be introduced before combining the simulation results of the sub-circuits, looping over the possible values of one or more such qubits. Then the subsequent tensor products and their sums may be calculated for slices of the affected tensors to reduce the storage requirements for subsequent calculations.
[0087] Return Figure 12A operation 1208 and its in Figure 12BIn the example shown, in at least one embodiment, a sub-circuit can be effectively emulated from an input state to an output state, starting from an initial sub-circuit created according to the initial states of individual qubits. Using this method, subsequent sub-circuits are not emulated until their input-dependent previous sub-circuits have been emulated. The emulation result of each sub-circuit can correspond to an m-dimensional tensor, which can be represented as an m-dimensional array in computer memory. Other data structures, such as linear arrays, can also be used to provide an equivalent representation.
[0088] Accordingly, circuit partitioning is described herein, where the resulting tensor as a whole fits into the available collective memory, or slices of the resulting tensor can be calculated using the available collective memory based on other tensors that have been computed and stored in the collective memory. The resulting tensor and / or their slices are generally larger than the memory of individual processing nodes. The memory of each existing processing node, in combination with the techniques described herein, minimizes communication by further partitioning the sub-circuit into sub-sub-circuits so that the corresponding sub-slices can fit into the memory of individual nodes.
[0089] In addition, when the quantum state is too large to fit into the collective memory, a second storage technique can be used in combination. Since auxiliary storage is typically several orders of magnitude slower than main memory, the feasibility of using auxiliary storage depends on the extent to which the number of read / write cycles can be minimized. To achieve this minimization, an initial portion of the circuit is described herein to be partitioned to attempt to maximize the number of gates that can be emulated using the available collective memory, and then the resulting quantum state is sliced and written to the auxiliary memory. Then, depending on the size of the collective memory, known partitioning methods can be applied to the remaining gates in the circuit, setting the number of "local" qubits higher rather than being limited to the memory size available on individual processing nodes. The resulting tensor slices ("meta" slices) are relatively large, allowing more gates to be emulated before additional auxiliary storage read / write cycles are required. As discussed herein, the resulting sub-circuit can then be further partitioned into sub-sub-circuits to minimize inter-node communication in the overall computation.
[0090] As set forth herein, for a 7×7 qubit, depth 27 random circuit such as Figures 4 to 7 the partitioning scheme shown can be extended so that a 7×7 qubit, depth 55 circuit ( Figure 5 ) can also be emulated using only one auxiliary memory read / write cycle, and a circuit of depth 83 ( Figure 6 ), depth 111 ( Figure 7 ), etc. The flow chart depicted in Figures 8 - 10 can be applied to construct the partitioning scheme shown in Figures 5 to 7 . Specifically, Figure 8 the flow chart 800 in Figure 9in the first iteration of the loop shown in Figure 9 and used in the example of operation 904 in Figure 10 The flowchart 1000 in can be used as this operation 904 in subsequent iterations. In Figure 9 in the first iteration of the loop shown in, an example of operation 904 can be applied in Figure 8 the flowchart 800 with an increased depth limit in multiple times to identify that the quantum state can be completely within a computationally reasonable depth - memory has a first - write - to - auxiliary - storage before. In this way, in Figures 4 to 7 the case of, 27 can be determined to be such a reasonable depth. In the case where the depth limit is set to 27, when considering the CZ gates bridging the third and fourth rows in stages 7 and 8 in the circuit, the sub - circuits 1 and 2 shown in Figure 8 can be identified by following the "yes" branch of operation 812 in Figures 5 to 7 and for the CZ gates strictly internal to sub - circuits 1 and 2, follow the "no" branch of operation 812. Doing so results in two multiple ordered sub - circuits corresponding to sub - circuits 1 and 2 respectively, as shown in Figures 5 to 7 . As a by - product of this process, since all remaining unassigned gates in the circuit up to depth 27 are diagonal single gates, qubits 43 - 49 can also be identified as slice objects. Qubits 43 - 49 can thus become the identified global qubits for constructing slices of the meta - slice. When considering any one of the CZ gates bridging the third and fourth rows in stages 15 and 16 in the quantum circuit, the sub - circuit 3 shown in Figure 8 can be identified by subsequently following the "no" branch of operation 812 in Figures 5 to 7 . When operation 804 is subsequently applied, the remaining gates in sub - circuit 3 up to depth 27 can then be identified. To extend the partitioning beyond depth 27, operation 804 can be continued beyond depth 27 under the constraint that qubits 43 - 49 are global, and thus, when extending sub - circuit 3 beyond depth, only the gates applied to local qubits 1 - 42 are considered. Doing so can produce a partitioning corresponding to sub - circuit 3 as shown in Figures 5 to 7 . In Figure 9 in the first iteration of the loop shown in, use the above - mentioned process as an embodiment of operation 904 in Figure 9 and then, by using Figure 10The flowchart 1000 in is used as operation 904 in subsequent iterations to identify sub - circuits 4, 5, and 6. It should be noted that it is not necessary to use flowchart 800 in operation 904. However, using flowchart 800 in the above - described manner can increase the depth of the quantum circuit that can be simulated before auxiliary memory must be written for the first time. It should also be noted that known optimization techniques (such as depth - first search, breadth - first search, iterative - deepening depth - first search, Dijkstra's algorithm, and A * * search) can be used in combination with flowchart 900 to optimize embodiments of operation 904; for example, to minimize the total execution time while taking into account computational cost, communication cost, and secondary - storage access cost.
[0091] A published method (Riling Li, Bujiao Wu, Mingsheng Ying, Xiaoming Sun, and Guangwen Yang, “Quantum Supremacy Circuit Simulation on Sunway TaihuLight,” arXiv preprint arXiv:1804.04797, (2018)) demonstrated the feasibility of estimating the individual amplitudes of a universal random circuit with ≈50 qubits and depth >40, which is considered beyond the reach of current technology. In contrast, as described above, a universal random circuit with more than 50 qubits and super depth can be fully simulated for all amplitudes computable on a supercomputer with auxiliary memory. Due to the high cost of disk read / write operations, this results in longer execution times, but for the instances under study, the slowdown is less than a factor of two. Additionally, recent system advancements, such as NVRAM-based burst buffers, can have a very beneficial impact on these runtimes. The larger storage pool available through auxiliary memory allows for further expansion of the boundaries of quantum circuit simulation. To this end, a quantum circuit can have a set of qubits on the boundary of the grid that do not interact with the other qubits of several layers of gates periodically. In particular, for a universal random circuit on a 7×7 grid, a set of 7 qubits on the boundary has two layers of two-qubit interactions with other rows or columns, followed by six layers with no further interactions (only single-qubit gates or two-qubit gates in a group). Therefore, slicing these qubits is very effective. One of the boundary rows or columns is chosen, and then the corresponding qubits of the six layers are sliced. The circuit applied to the remaining qubits is simulated independently. More specifically, the technique slices the qubits at the boundary of the grid (e.g., the bottom row of the grid) and uses the Schrödinger method to simulate the rest of the circuit in a way that the state vector is as large as the memory. As many gate layers as possible are simulated without introducing additional entanglement exponents. For qubits on the opposite row / column (e.g., the top row of the grid), this allows for the application of more than thirty layers of gates, while for qubits closer to the qubits that have been sliced, the process stops at fewer layers before the memory footprint increases. This results in several slices of the same size, which are stored to disk. To apply more gates, a new subcircuit is launched, slicing the qubits in the row / column opposite to the row / column sliced in the previous subcircuit (e.g., the top row of the grid) and simulating the rest of the circuit using the Schrödinger method. The initial state can be loaded from the slices stored on disk, and then the process can be repeated. This results in a “waveform” pattern in the subcircuits, where a given subcircuit starts from the top or bottom row of the qubits and then expands to the rest of the circuit until it includes the row opposite to the starting set of qubits, at which point it contracts back. One can reduce the memory requirements (rather than the disk storage requirements) by slicing additional qubits, e.g., two rows.Even without disk operations, the techniques described herein allow for the computation of a single amplitude of a depth-46 circuit on a 7×7 qubit grid.
[0092] As described herein, the efficiency of using auxiliary memory depends on the extent to which the number of read / write cycles can be minimized. The data transfer time required to write 249 quantum amplitudes to auxiliary memory and read them back is approximately 2 to 5 hours, with a transfer rate of approximately 1.0 to 2.2 TB / second. The quantum amplitudes can be stored in single-precision and / or double-precision storage formats; (when the amplitudes are stored in single precision, in-memory computations can still be performed in double precision to minimize the accumulation of rounding errors). Thus, auxiliary memory can be used as long as the number of read / write cycles can be kept to a minimum.
[0093] Figure 13 A system is represented that includes a partitioning component (block 1302) that partitions an input quantum circuit including a machine-readable specification of a quantum circuit into subcircuits for tensor slicing based on at least two identified sets of qubits, where the subcircuit groups have associated sets of qubits for tensor slicing. A simulation component (block 1304) simulates the input quantum circuit in stages into a simulated quantum state tensor based on the subcircuits, one stage per subcircuit, where the set of qubits associated with a subcircuit is used to partition the simulated quantum state tensor of the input quantum circuit into quantum state tensor slices, and the quantum gates in the subcircuit are used to update the quantum state tensor slices to updated quantum state tensor slices. A read / write component (block 1306) stores the updated quantum state tensor slices of the simulated quantum state tensor as micro-slices in auxiliary memory.
[0094] The input quantum circuit can include at least forty-nine qubits and have a circuit depth rating of at least fifty.
[0095] The read / write component can write the micro-slices in a size that spans at least two disk sectors of the auxiliary memory. The read / write component can retrieve the updated quantum state tensor slices from the auxiliary memory, while the simulation component can process another subcircuit into another subcircuit tensor and use the retrieved updated quantum state tensor slices to update the other subcircuit tensor. To a further updated quantum state tensor slice. The read / write component can store the further updated quantum state tensor slices in the auxiliary memory.
[0096] The partitioning component can partition an input quantum circuit into groups of qubit partitions corresponding to sub-circuits, including a first qubit partition group, a second qubit partition group, a third qubit partition group, and a fourth qubit partition group. The simulation component can apply the quantum gates of the third group to the tensor slices of the first and second groups to obtain an updated quantum state tensor, and the read / write component can read the updated quantum state tensor slices from the auxiliary memory into the memory. The simulation component can apply the quantum gates of the fourth group to the updated quantum state tensor slices to obtain a further updated quantum state tensor.
[0097] Figure 14 Represents an exemplary computer-implemented method, where operation 1402 represents processing an input quantum circuit that includes a machine-readable specification of a quantum circuit. The processing can include partitioning (e.g., by operation 1404 performed by partitioning component 108) the input quantum circuit into groups of sub-circuits based on at least two groups of qubits identified for tensor slicing, where the sub-circuits have associated sets of qubits for tensor slicing. Operation 1406 represents simulating (e.g., by simulation component 112 of) the input quantum circuit in stages based on the groups of sub-circuits as a simulated quantum state tensor, one stage per sub-circuit, where the set of qubits associated with the sub-circuit is used to partition the simulated quantum state tensor of the input quantum circuit into quantum state tensor slices, and the quantum gates in the sub-circuit are used to update the quantum state tensor slices to updated quantum state tensor slices. Operation 1408 represents storing (e.g., by read / write component 114 of) the updated quantum state tensor slices of the simulated quantum state tensor as micro-slices in auxiliary memory. Figure 1 The partitioning component can partition an input quantum circuit into groups of qubit partitions corresponding to sub-circuits, including a first qubit partition group, a second qubit partition group, a third qubit partition group, and a fourth qubit partition group. The simulation component can apply the quantum gates of the third group to the tensor slices of the first and second groups to obtain an updated quantum state tensor, and the read / write component can read the updated quantum state tensor slices from the auxiliary memory into the memory. The simulation component can apply the quantum gates of the fourth group to the updated quantum state tensor slices to obtain a further updated quantum state tensor. Figure 1 The partitioning component can partition an input quantum circuit into groups of qubit partitions corresponding to sub-circuits, including a first qubit partition group, a second qubit partition group, a third qubit partition group, and a fourth qubit partition group. The simulation component can apply the quantum gates of the third group to the tensor slices of the first and second groups to obtain an updated quantum state tensor, and the read / write component can read the updated quantum state tensor slices from the auxiliary memory into the memory. The simulation component can apply the quantum gates of the fourth group to the updated quantum state tensor slices to obtain a further updated quantum state tensor. Figure 1 The partitioning component can partition an input quantum circuit into groups of qubit partitions corresponding to sub-circuits, including a first qubit partition group, a second qubit partition group, a third qubit partition group, and a fourth qubit partition group. The simulation component can apply the quantum gates of the third group to the tensor slices of the first and second groups to obtain an updated quantum state tensor, and the read / write component can read the updated quantum state tensor slices from the auxiliary memory into the memory. The simulation component can apply the quantum gates of the fourth group to the updated quantum state tensor slices to obtain a further updated quantum state tensor.
[0098] Storing the updated quantum state tensor slices of the simulated quantum state tensor as secondary slices can include storing the micro-slices in at least two disk sectors of the auxiliary memory.
[0099] Aspects can include processing another sub-circuit into other sub-circuit tensors, retrieving the updated quantum state tensor slices from the auxiliary memory, and updating the other sub-circuit tensors with the updated quantum state tensor slices retrieved from the auxiliary memory into further updated quantum state tensor slices. Aspects can include the processing of storing the further updated quantum state tensor slices in the auxiliary memory.
[0100] Partitioning an input quantum circuit into sub-circuit groups may include partitioning the input quantum circuit into a first sub-circuit partition group, a second sub-circuit partition group, and a third sub-circuit partition group, and storing the updated quantum state tensor slices may include applying the gates of the third sub-circuit partition group to the tensor slices corresponding to the first sub-circuit partition group and the second sub-circuit partition group to obtain the updated quantum state tensor slices. Partitioning the input quantum circuit into sub-circuit groups may further include: partitioning the input quantum circuit into a fourth sub-circuit partition group; and retrieving the updated quantum state tensor slices from an auxiliary memory; updating other sub-circuit tensors with the updated quantum state tensor slices retrieved from the auxiliary memory to further updated quantum state tensor slices may include applying the gates of the fourth sub-circuit partition group.
[0101] As is understood, Figure 14 Exemplary operations generally can be implemented in a computer program product for simulating a quantum circuit, including an input quantum circuit that includes a machine-readable specification of the quantum circuit. To this end, the computer program product may include one or more computer-readable storage media and program instructions stored on the one or more computer-readable storage media. The program instructions, when executed, may correspond to Figure 14 at least some of the operations illustrated in
[0102] Figure 15 An exemplary computer-implemented method represented, the operations of which may be performed by a device operably coupled to at least two processors, includes operation 1502, which represents simulating an input quantum circuit that includes a machine-readable specification of a quantum circuit. Operation 1504 represents partitioning the input quantum circuit into a plurality of sub-sub-circuits based on at least two sets of qubits identified for tensor slices, where at least one sub-sub-circuit exceeds the memory of a single processor. Aspects may include sub-partitioning a meta sub-circuit into sub-sub-circuits that fit within the memory of a single processor (operation 1506), computing tensors for the sub-sub-circuits (operation 1508), and contracting the tensors into tensor slices by tensor slicing (operation 1510). Aspects may include applying the gates of another sub-circuit to the tensor slices to obtain updated tensor slices representing quantum state data and storing them in an auxiliary memory. Aspects may include retrieving the updated tensor slices from the auxiliary memory and applying the gates of additional other sub-circuits to the updated tensor slices to obtain further updated tensor slices.
[0103] The auxiliary storage device may include one or more disk devices, and aspects may include organizing meta - slices into micro - slices, where the size of a micro - slice spans at least two disk sectors of the auxiliary storage device, and storing a slice of an updated tensor into the auxiliary storage may include storing the micro - slice. Alternatively, or in addition to disk devices, the auxiliary storage may include NVRAM, phase - change memory, flash memory, etc. Further, the auxiliary storage may refer to “remote” memory at least to some extent. For example, Remote Direct Memory Access allows for efficient data transfer without CPU intervention and thus can be used with micro - slices. Note that the operations performed need not be “memory - symmetric” as, for example, a node (e.g., node 0) may store a slice on another node (e.g., node 158) because node 158 has more memory or has the same amount of memory but requires fewer resources.
[0104] As is understood, Figure 15 Exemplary operations, generally speaking, may be implemented as a computer program product for simulating a quantum circuit, including an input quantum circuit that includes a machine - readable specification of the quantum circuit. To this end, the computer program product may include one or more computer - readable storage media and program instructions stored on the one or more computer - readable storage media. The program instructions, when executed, may correspond to Figure 15 at least some of the operations illustrated in
[0105] To provide background for various aspects of the present invention, Figure 16 and the following discussion is intended to provide a general description of a suitable environment in which various aspects of the present invention may be implemented. Figure 16 A block diagram of an example operating environment is shown in which one or more embodiments of the present invention described herein may be facilitated.
[0106] Referring to Figure 16, the environment 1600 may also include a computer 1612. The computer 1612 may also include a processing unit 1614, a system memory 1616, and a system bus 1618. The system bus 1618 couples system components, including: The processing unit 1614 may include, but is not limited to, the system memory 1616. The processing unit 1614 may be any of a variety of available processors. Dual microprocessors and other multi-processor architectures may also be used as the processing unit 1614. The system bus 1618 may be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus or external bus, and / or a local bus using any available bus architecture, including but not limited to Industry Standard Architecture (ISA), Micro Channel Architecture (MSA), Extended ISA (EISA), Intelligent Drive Electronics (IDE), VESA Local Bus (VLB), Peripheral Component Interconnect (PCI), Card Bus, Universal Serial Bus (USB), Advanced Graphics Port (AGP), FireWire (IEEE1394), and Small Computer System Interface (SCSI).
[0107] The system memory 1616 may also include volatile memory 1620 and non-volatile memory 1622. The Basic Input / Output System (BIOS) contains basic routines such as those that transfer information between elements within the computer 1612 during startup. It is stored in the non-volatile memory 1622. The computer 1612 may also include removable / non-removable, volatile / non-volatile computer storage media. Figure 16 The disk storage device 1624 is illustrated. The disk storage device 1624 may also include, but is not limited to, disk drives, floppy disk drives, tape drives, Jaz drives, Zip drives, LS-100 drives, flash cards, or memory sticks. The disk storage device 1624 may also include storage media, either alone or in combination with other storage media. To facilitate the connection of the disk storage device 1624 to the system bus 1618, a removable or non-removable interface, such as interface 1626, is typically used. Figure 16 Software that acts as an intermediary between a user and the basic computer resources described in the user interface in the appropriate operating environment 1600 is also depicted. Such software may also include, for example, an operating system 1628. The operating system 1628, which may be stored on the disk memory 1624, is used to control and allocate the resources of the computer 1612.
[0108] The system application 1630 manages resources using the operating system 1628 via program modules 1632 and program data 1634 (such as program data stored in the system memory 1616 or disk storage 1624). It should be understood that the present disclosure can be implemented using various operating systems or combinations of operating systems. A user inputs commands or information to the computer 1612 through one or more input devices 1636. The input devices 1636 include, but are not limited to, pointing devices such as a mouse, trackball, stylus, touchpad, keyboard, microphone, joystick, gamepad, satellite dish, scanner, TV tuner card, digital camera, digital video camera, web camera, etc. These and other input devices are connected to the processing unit 1614 via the interface port 1638 through the system bus 1618. The interface port 1638 includes, for example, serial ports, parallel ports, game ports, and universal serial bus (USB) ports. The output device 1640 uses some of the same types of ports as the input device 1636. Thus, for example, a USB port can be used to provide input to the computer 1612 and to output information from the computer 1612 to the output device 1640. An output adapter 1642 is provided to account for the fact that some output devices 1640, such as monitors, speakers, and printers, require special adapters among other output devices 1640. By way of illustration and not limitation, the output adapter 1642 includes video and sound cards, which provide a means of connection between the output device 1640 and the system bus 1618. It should be noted that other devices and / or device systems provide both of these input and output functions simultaneously, such as the remote computer 1644.
[0109] Computer 1612 can operate in a networked environment using logical connections to one or more remote computers, such as remote computer 1644. Remote computer 1644 can be a computer, server, router, network PC, workstation, microprocessor-based device, peer device, or other common network node, etc., and typically can also include many or all of the elements described relative to computer 1612. For simplicity, only memory storage device 1646 of remote computer 1644 is illustrated. Remote computer 1644 is logically connected to computer 1612 via network interface 1648 and then physically connected via communication connection 1650. Network interface 1648 includes wired and / or wireless communication networks, such as local area network (LAN), wide area network (WAN), cellular network, etc. LAN technologies include Fiber Distributed Data Interface (FDDI), Copper Distributed Data Interface (CDDI), Ethernet, Token Ring, etc. WAN technologies include, but are not limited to, point-to-point links, circuit-switched networks (such as Integrated Services Digital Network (ISDN)) and variants thereof, packet-switched networks, and Digital Subscriber Line (DSL). Communication connection 1650 refers to the hardware / software used to connect network interface 1648 to system bus 1618. Although communication connection 1650 is shown inside computer 1612 for clarity of illustration, it can also be external to computer 1612. The hardware / software used to connect to network interface 1648 can also include internal and external technologies for exemplary purposes only, such as modems including conventional telephone-grade modems, cable modems, and DSL modems, ISDN adapters, and Ethernet cards.
[0110] The present invention can be a system, method, apparatus, and / or computer program product at any possible level of integration of technical details. The computer program product can include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to execute aspects of the present invention. The computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium can be, by way of example and not limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium can also include the following: a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions recorded thereon, and any appropriate combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable) or an electrical signal transmitted through a wire.
[0111] The computer-readable program instructions described herein can be downloaded to a corresponding computing / processing device from a computer-readable storage medium or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network). A regional network and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device. The computer-readable program instructions for performing the operations of the present invention can be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for an integrated circuit, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions can be executed entirely on the user's computer, partially as a stand-alone software package on the user's computer, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can establish a connection with an external computer (for example, via the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), can perform the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit to perform aspects of the present invention.
[0112] Aspects of the present invention are described herein with reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions. These computer-readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner. The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, e.g., the instructions, which execute on the computer, other programmable apparatus, or other devices, implement the functions / acts specified in the flowchart and / or block diagram block.
[0113] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, depending on the functionality involved, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order.
[0114] Although the subject matter has been described above in the general context of computer-executable instructions of a computer program product running on one or more computers, those skilled in the art will recognize that the present disclosure may also or may be implemented in conjunction with other program modules. In general, program modules include routines, programs, components, data structures, etc. that perform particular tasks and / or implement particular abstract data types. Additionally, those skilled in the art will understand that the computer-implemented methods of the present invention may be practiced using other computer system configurations, including single-processor or multi-processor computer systems, small computing devices, mainframe computers, and hand-held computers, computing devices (e.g., PDAs, telephones), microprocessor-based or programmable consumer electronics or industrial electronics, etc. The aspects shown may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network. However, some aspects of the present invention, if not all, may be practiced on a stand-alone computer. In a distributed computing environment, program modules may be located in local and remote storage devices.
[0115] As used in this application, the terms "component", "system", "platform", "interface", etc. may refer to and / or may include a computer-related entity or an operable machine entity with one or more specific functions related to a computer. The entities disclosed herein may be hardware, a combination of hardware and software, software, or software in execution. For example, a component may be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, an application running on a server and the server can both be components. One or more components may reside within an execution process and / or thread, and a component may be located on one computer and / or distributed between two or more computers. In another example, various components may execute from various computer-readable media on which various data structures are stored. A component may communicate, for example, via signals having one or more data packets through local and / or remote procedures (e.g., data from one component interacts with another component in a local system, a distributed system, and / or across a network connected to other systems via signals as the Internet). As another example, a component may be a device having a particular function provided by a mechanical component operated by an electrical or electronic circuit, where the mechanical component is operated by a software or firmware application executed by a processor. In such a case, the processor may be inside or outside the device and may execute at least a portion of the software or firmware application. As yet another example, a component may be a device having a particular function provided by an electronic component without a mechanical component, where the electronic component may include a processor or other means to execute software or firmware that at least partially imparts the function to the electronic component. In one aspect, a component may emulate an electronic component via a virtual machine, e.g., within a cloud computing system.
[0116] Additionally, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise stated or clear from the context, "X uses A or B" is intended to mean any natural inclusive permutation. That is, if X uses A; X uses B; or X uses both A and B, then "X uses A or B" is satisfied in any of the foregoing instances. Further, the articles "a" and "an" as used in the specification and drawings herein are generally to be construed to mean "one or more" unless otherwise stated or clearly understood from the context to be in the singular form. As used herein, the terms "example" and / or "exemplary" are used to mean serving as an example, instance, or illustration. To avoid doubt, the subject matter disclosed herein is not limited by these examples. Additionally, any aspect or design described herein as "example" and / or "exemplary" need not be construed as preferred or advantageous over other aspects or designs, nor does it imply exclusion of equivalent exemplary structures and techniques known to those of ordinary skill in the art.
[0117] As used in this specification, the term "processor" can generally refer to any computing processing unit or device; including but not limited to a single-core processor; a single processor with software multi-threading execution capabilities; a multi-core processor; a multi-core processor with software multi-threading execution capabilities; a multi-core processor with hardware multi-threading technology; a parallel platform; and a parallel platform with distributed shared memory. Additionally, a processor can refer to an integrated circuit, an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic controller (PLC), a complex programmable logic device (CPLD), discrete gate or transistor logic, discrete hardware components. Further, a processor can utilize nanoscale architectures, such as but not limited to transistors, switches, and gates based on molecules and quantum dots, in order to optimize space usage or enhance the performance of a user device. A processor can also be implemented as a combination of computing processing units. In this disclosure, terms such as "storage", "storage device", "data storage", "data storage device", "database", and substantially any other information storage component related to the operation and function of a component are used to refer to "memory", an entity contained in "memory" or a component that includes memory. It should be understood that the memory and / or memory components described herein can be volatile memory or non-volatile memory, or can include both volatile and non-volatile memory. By way of illustration and not limitation, non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), flash memory, or non-volatile random access memory (RAM) (such as ferroelectric RAM (FeRAM)). Volatile memory can include RAM, for example, RAM can be used as an external cache. By way of illustration and not limitation, RAM has various forms, such as synchronous RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), direct Rambus RAM (DRRAM), direct Rambus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM). Additionally, the memory components of the systems or computer-implemented methods disclosed herein are intended to include but not limited to include these and any other suitable types of memory.
[0118] The foregoing description has included only examples of systems and computer-implemented methods. Of course, for purposes of describing the present invention, it is not possible to describe every possible combination of components or computer-implemented methods, but one of ordinary skill in the art will recognize that many other combinations and permutations of the present invention are possible. In addition, to the extent that the terms "including," "having," "comprising," and the like are used in the specification, claims, appendices, and drawings, these terms are intended to be construed in a manner similar to the term "comprising" as interpreted when used as a transitional word in a claim.
[0119] The description of the various embodiments of the present invention has been presented for purposes of illustration but is not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to one of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein were chosen to best explain the principles of the embodiments of the present invention, the practical application to techniques found in the marketplace, or the improvement of such techniques, or to enable other one of ordinary skill in the art to understand the embodiments of the present invention disclosed herein.
Claims
1. A system, comprising Partitioning component that partitions an input quantum circuit of a machine-readable specification including a quantum circuit into sub-circuits represented as tensors in an aggregated main memory based on at least two sets of qubits identified for tensor slicing, where The sub - circuit has a relevant qubit set for tensor slicing; wherein the partitioning component initializes a global qubit set for micro - slice indexing, identifies and selects a sub - circuit that is consistent with the global qubit set for micro - slice indexing, adds the global qubits of the selected sub - circuit to the global qubit set for micro - slice indexing, and removes the gates in the selected sub - circuit from the input quantum circuit, and confirms the execution order of the sub - circuits; wherein identifying and selecting a sub - circuit that is consistent with the global qubit set for micro - slice indexing includes: for each gate in the input quantum circuit, calculating the local qubit set on which the gate depends; identifying the local qubit set that is consistent with the current global qubit set for micro - slice indexing and maximizes the number of applied gates; and given the identified local qubits, identifying and selecting the sub - circuits that can be applied, wherein identifying the local qubit set that is consistent with the current global qubit set for micro - slice indexing is such that adding the global qubits not included in the identified local qubit set to the current global qubit set for indexing micro - slices does not cause the resulting global qubit set for micro - slice indexing to exceed the maximum number of qubits required to achieve efficient auxiliary memory access; a simulation component that, based on the sub - circuits, simulates the input quantum circuit in stages as a simulated quantum state tensor, one stage for each sub - circuit, wherein the qubit set associated with a sub - circuit is used to partition the simulated quantum state tensor of the input quantum circuit into quantum state tensor slices, and the quantum gates in the sub - circuit are used to update the quantum state tensor slices into updated quantum state tensor slices; and a read - write component that stores the updated quantum state tensor slices of the simulated quantum state tensor as micro - slices in an auxiliary memory (106) comprising at least two disk sectors, the micro - slices being stored across the at least two disk sectors, wherein the auxiliary memory comprises one or more disk devices; wherein the read - write component also retrieves the updated quantum state tensor slices from the auxiliary memory, and wherein the simulation component processes another sub - circuit into another sub - circuit tensor and updates the another sub - circuit tensor with the updated quantum state tensor slices retrieved from the auxiliary memory into a further updated quantum state tensor slice.
2. The system according to claim 1, wherein The partitioning component partitions the input quantum circuit into qubit sets corresponding to the sub-circuits, including a first qubit partitioning group, a second qubit partitioning group, a third qubit partitioning group, and a fourth qubit partitioning group. Wherein, the simulation component applies the quantum gates of the third qubit partitioning group to the tensor slices of the first qubit partitioning group and the second qubit partitioning group to obtain the updated quantum state tensor. Wherein, the read / write component reads the updated quantum state tensor slices from the auxiliary memory (106) into the memory, and wherein the simulation component applies the quantum gates of the fourth qubit partitioning group to the updated quantum state tensor slices to obtain a further updated quantum state tensor.
3. A computer-implemented method, comprising: Processing an input quantum circuit including a machine-readable specification of a quantum circuit by a device operably coupled to a processor, the processing comprising, Partitioning the input quantum circuit into a set of sub-circuits based on at least two sets of qubits identified for tensor slices, wherein the set of sub-circuits has a corresponding set of qubits for tensor slices; Wherein partitioning the input quantum circuit into a set of sub-circuits includes: initializing a global set of qubits for micro-slice indexing, identifying and selecting sub-circuits consistent with the global set of qubits for micro-slice indexing, adding the global qubits of the selected sub-circuits to the global set of qubits for micro-slice indexing, and removing the gates in the selected sub-circuits from the input quantum circuit, and confirming the execution order of the sub-circuits; Wherein identifying and selecting sub-circuits consistent with the global set of qubits for micro-slice indexing includes: for each gate in the input quantum circuit, calculating the local set of qubits on which the gate depends; identifying the local set of qubits consistent with the current global set of qubits for micro-slice indexing and maximizing the number of applied gates; and given the identified local qubits, identifying and selecting the sub-circuits that can be applied. Wherein, identifying the local set of qubits consistent with the current global set of qubits for micro-slice indexing such that adding the global qubits not included in the identified local set of qubits to the current global set of qubits for indexing micro-slices does not cause the resulting global set of qubits for micro-slice indexing to exceed the maximum number of qubits required to achieve efficient auxiliary memory access; Phased simulating the input quantum circuit as a simulated quantum state tensor based on the set of sub-circuits, one phase for each sub-circuit, wherein the set of qubits associated with the sub-circuit is used to partition the simulated quantum state tensor of the input quantum circuit into quantum state tensor slices, and the quantum gates in the sub-circuit are used to update the quantum state tensor slices to updated quantum state tensor slices; Storing the updated quantum state tensor slices of the simulated quantum state tensor as micro-slices in an auxiliary memory (106) including at least two disk sectors, the micro-slices being stored across the at least two disk sectors; and Processing another sub-circuit into other sub-circuit tensors, retrieving the updated quantum state tensor slice from the auxiliary memory (106), and updating the other sub-circuit tensors with the retrieved updated quantum state tensor slice from the auxiliary memory (106) to become further updated quantum state tensor slices.
4. The computer-implemented method according to claim 3, further comprising: Storing the further updated quantum state tensor slices in the auxiliary memory (106) by the device.
5. The computer-implemented method according to claim 3, wherein partitioning the input quantum circuit into the groups of sub-circuits includes partitioning the input quantum circuit into a first sub-circuit partition group, a second sub-circuit partition group, and a third sub-circuit partition group, and wherein storing the updated quantum state tensor slices includes: Applying the gates of the third sub-circuit partition group to the tensor slices corresponding to the first sub-circuit partition group and the second sub-circuit partition group to obtain the updated quantum state tensor slices.
6. The computer-implemented method according to claim 5, wherein partitioning the input quantum circuit into a set of sub-circuits further comprises: Partitioning the input quantum circuit into a fourth sub-circuit partition group, and further comprising: retrieving the updated quantum state tensor slices from the auxiliary memory (106), and updating other sub-circuit tensors with the retrieved updated quantum state tensor slices from the auxiliary memory (106) to become further updated quantum state tensor slices, including applying the gates of the fourth sub-circuit partition group.
7. A computer-implemented method, comprising: Emulating an input quantum circuit of a machine-readable specification including a quantum circuit by a device operatively coupled to a processor, the emulation comprising: Partitioning the input quantum circuit into a plurality of meta-sub-circuits represented as tensors in an aggregate main memory (104) based on at least two sets of qubits identified for tensor slicing, wherein at least one of the meta-sub-circuits exceeds the memory (304) of a single processor; Wherein partitioning the input quantum circuit into a plurality of meta-sub-circuit groups includes: initializing a global set of qubits for micro-slice indexing, identifying and selecting meta-sub-circuits consistent with the global set of qubits for micro-slice indexing, adding the global qubits of the selected meta-sub-circuits to the global set of qubits for micro-slice indexing, and removing the gates in the selected meta-sub-circuits from the input quantum circuit, and confirming the execution order of the meta-sub-circuits. Wherein identifying and selecting sub-circuits consistent with the global set of qubits for micro-slice indexing includes: for each gate in the input quantum circuit, calculating the local set of qubits on which the gate depends; identifying the local set of qubits consistent with the current global set of qubits for micro-slice indexing and maximizing the number of applied gates; and given the identified local qubits, identifying and selecting sub-circuits that can be applied, wherein identifying the local set of qubits consistent with the current global set of qubits for micro-slice indexing such that adding the global qubits not included in the identified local set of qubits to the current global set of qubits for indexing micro-slices does not cause the resulting global set of qubits for micro-slice indexing to exceed the maximum number of qubits required for efficient auxiliary memory access. Sub-partitioning the meta-sub-circuits into sub-sub-circuits suitable for the memory (304) of the single processor; Calculating the tensors of the sub-sub-circuits; Contracting the tensors into tensor slices by tensor slicing; Apply the gates of another sub-circuit to the tensor slice to obtain an updated tensor slice representing quantum state data, and store the updated tensor slice as a micro-slice in an auxiliary memory (106) comprising at least two disk sectors, the micro-slice being stored across the at least two disk sectors, wherein the auxiliary memory (106) comprises one or more disk devices; and Retrieve the updated tensor slice from the auxiliary memory (106) and apply the gates of another other sub-circuit to the updated tensor slice to obtain a further updated tensor slice.
8. A computer program product for simulating a quantum circuit, comprising an input quantum circuit that includes a machine-readable specification of the quantum circuit, the computer program product comprising one or more computer-readable storage media having program instructions embodied thereon, the program instructions being executable by a processing component to cause the processor to perform the method according to any one of claims 3 to 7.
Citation Information
Patent Citations
Simulating quantum circuits
US11250190B2
Cited By
Simulating quantum circuits
US20220164506A1