Processor and method for tensor network contraction in quantum simulators

By optimizing the tensor network contraction through the quantum simulator processor and utilizing local search algorithms and convolution operations, the problem of high complexity of tensor network contraction in traditional technologies is solved, and efficient quantum circuit simulation and verification is achieved.

CN117203647BActive Publication Date: 2025-09-23HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180096422.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-30
Publication Date
2025-09-23
Estimated Expiration
2041-03-30

AI Technical Summary

Technical Problem

Traditional quantum simulators need to process a large number of random amplitudes or batch amplitudes when simulating and verifying large quantum systems. Existing technologies make it difficult to efficiently shrink tensor networks, resulting in excessively high computational complexity and memory requirements.

Method used

A quantum simulator processor is provided to optimize the contraction of a tensor network through a local search algorithm, select the contraction expression and contraction tree with the lowest cost, reduce memory requirements and computational complexity, and improve efficiency by using convolution operations and slicing operations.

Benefits of technology

It enables efficient simulation and verification of quantum circuits on large quantum systems, reduces memory and computational complexity, and improves the efficiency of processing large numbers of amplitudes or batches of amplitudes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117203647B_ABST
    Figure CN117203647B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of quantum computing, and more particularly to simulating quantum circuits using a quantum simulator. The present disclosure provides a processor for a quantum simulator. The processor is configured to execute a local search algorithm to determine a plurality of contraction expressions suitable for contracting a corresponding tensor network into a determined contraction tensor network. The processor is further configured to, for each contraction expression, determine a contraction cost for contracting the corresponding tensor network based on a cost function, and select the contraction expression with the lowest contraction cost to contract each tensor network into the determined contraction tensor network. The cost function is based on three parameters, which respectively indicate the amount of memory required to contract the corresponding tensor network into the determined contraction tensor network, the computational complexity, and the number of read and write operations required.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of quantum computing, and more particularly to the task of simulating quantum circuits using a quantum simulator. The present disclosure provides a processor for such a quantum simulator, wherein the processor is configured to collapse one or more tensor networks into a deterministic collapsed tensor network. The present disclosure also provides a quantum simulator comprising the processor, and wherein the tensor network collapse performed by the processor can be used during simulation and / or verification of a quantum circuit. Background Art

[0002] Conventional quantum simulators are used to simulate and / or verify quantum circuits. Specifically, they can be used to find the individual amplitudes or batch amplitudes of samples generated when a quantum circuit is run on a quantum system. A batch amplitude is a collection of bit strings that share a fixed number of bits.

[0003] However, many quantum simulation tasks require looking up very large sets of random amplitudes or batches of amplitudes. For example, for the verification task of Google's Sycamore 53-qubit random quantum circuit (RQC) experiment used to prove quantum supremacy, each quantum circuit needs to look up about 10 6 Specifically, the verification task requires finding the exact amplitudes of random samples generated by multiple quantum circuits running on a quantum system with a 53-qubit chip (e.g. Figure 1 shown).

[0004] Given an n-qubit quantum circuit C represented by its unitary operator, the expression p C (s)=| <s|C|0 n >| 2 Represents the bit string s∈{0,1} n The probability of the bit string found can be used to estimate the fidelity of the quantum system. In Google's quantum computing supremacy experiment, the linear cross entropy benchmark (linear XEB) was proposed as a tool for estimating fidelity. The bit string (called sample s1,…,s N ) sequence of linear XEB (also expressed as ) is defined as follows:

[0005]

[0006] Since the number of amplitudes used in the verification task is very large (about 10 in the Google experiment), 6 ), and the corresponding bit string is random, so a fairly powerful quantum simulator is required. Given an n-qubit quantum circuit C and a sample sequence s1,…,s N ∈{0,1} n, the quantum simulator should look for the corresponding N complex amplitudes i |C|0 n >, i=1,…,N.

[0007] Sometimes, you may also need to find N batches of amplitudes instead of N amplitudes, where the batch size is 2 m The set of units strings Where 1≤i1<… m ≤n,a i ∈{0,1}. Therefore, within a batch, nm components are fixed, and the remaining m components can be any binary value. Each batch can be described by a mask of length n over the alphabet {0,1,x}, where 0 or 1 is placed in a fixed nm positions and x is placed in the remaining m free positions. For example, 01x0x corresponds to a batch B = {01000, 01001, 01100, 01101}.

[0008] For example, one may have to find the amplitude of N random batches of size 64, and one may have to calculate the amplitude of N random batches of size 64 based on the 64 probabilities p obtained. C (s), s∈B samples a bit string s∈B from each batch B.

[0009] In summary, there is a need for a powerful quantum simulator that can perform the tasks described above. Specifically, there is a need for a quantum simulator that can simulate and / or verify quantum circuits running on large quantum systems. Summary of the Invention

[0010] The present disclosure and its embodiments are further based on the following considerations made by the inventors.

[0011] Tensor networks are an important tool for simulating large quantum systems, i.e., the processor of a quantum simulator can use such tensor networks to perform simulations and / or verifications of quantum circuits. Tensor networks can represent quantum circuits (e.g., see Figure 9 ), and can therefore be used to simulate and / or verify quantum circuits.

[0012] Tensors are multidimensional arrays It can also be viewed as a multivariable function A(i1,…,i n ), where variables i1,…,i n Also called the "index" or "tensor leg", the number n is often called the "rank" or "degree" of A. The range of each tensor leg as a variable is a finite set, the size of which is often called the "bond dimension". Therefore, vectors can be considered as 1-degree tensors and matrices can be considered as 2-degree tensors. The shape of a tensor A is the vector (d1, ..., d n ), where d i ​​is the bond dimension of the i-th tensor leg.

[0013] Furthermore, tensor networks can be represented as graphs (see Figure 2 (a)), where tensors are assigned to vertices 202, tensor legs (i.e., indices) are assigned to edges 203, and tensors are connected by edges 203 if they share a common tensor leg.

[0014] Tensor networks are often used as a graph representation of sums involving multiple tensors. For example, Figure 2 The tensor network shown in (a) represents the following sum:

[0015] ∑ i,j,k,l,m T i,j S i,k U j,k,m Q m,l R l,f .

[0016] It should be noted that the sum depends on the index f, which does not participate in the summation. This is because, by definition, in a tensor network, the sum is only established over the tensor legs that connect two tensors. All other tensor legs do not participate in the sum and can be called "open legs". In the more general case, open tensor legs that do not need to be connected to only one tensor can be arbitrarily selected. In addition, sometimes tensor legs can be connected to more than two tensors, in which case these tensor legs can be represented by hyperedges instead of edges 203.

[0017] The present disclosure can be well applied to this more general case without any major modifications. However, for the sake of brevity, the present disclosure assumes in the following that all tensor legs are connected to no more than two tensors, and that a tensor leg connected to only one tensor is open. The sum of is called "the result of its contraction" and is expressed as result is a tensor, where the tensor leg is a tensor network Open legs in.

[0018] To optimize the memory used by contractions, some tensor legs can be fixed in the tensor network (this is often also called "slicing" or "variable projection"). This requires finding the contractions for all possible values ​​of the fixed tensor legs (this can be done in parallel) and then summing all the results (see Figure 2 (b)).

[0019] Therefore, in a tensor network, several tensor legs can be fixed to reduce the maximum size of intermediate tensors during a contraction of the tensor network so that all intermediate tensors are placed in the memory of the device performing the contraction.

[0020] Finally, the result of shrinking the entire network is obtained.

[0021] However, there is a problem with this approach. Typically, the overall complexity of the contraction increases. However, with good optimization, the loss in complexity is not that significant.

[0022] If the tensor network There are two tensors and and their common non-open tensor legs are exactly k1,…,k q , then they are expressed as The convolution of can be found as follows:

[0023]

[0024] If the tensor network is known from the context, the index is usually omitted Therefore, convolution can be simply expressed as T*S. The computational cost C of the convolution operation is easy to estimate. It involves d q·r multiplications and almost the same number of additions, where d is the bond dimension of the tensor legs (for simplicity, we can assume that all bond dimensions are equal) and r is the number of open tensor legs in the result T*S. In fact, we can do q Sum the terms and calculate all possible d r In fact, on modern processors, the complexity of the operation fma(a,b,c)=a+b*c is usually the same as the multiplication b*c, so the total complexity of the convolution can be calculated as d q·r operations. The number of memory operations RW is also an important parameter, because sometimes the required read / write operations are the bottleneck of the algorithm. As an estimate of the memory operations, it can be assumed that all operations only consist of reading tensors T, S and writing the results T*S. Therefore, the number of such operations can be estimated as size(T)+size(S)+size(T*S), where size(X) is the size of tensor X, that is, the product of all key dimensions of its tensor legs.

[0025] In terms of tensor networks, the convolution of two tensors corresponds to the merging of corresponding vertices in the tensor network (see Figure 3 (a), where the vertices of tensors T and S are merged into vertex T*S).

[0026] It can be determined that the convolution operation T*S (as a binary operation on tensors) satisfies the following associativity and commutativity conditions:

[0027] T*(S*R)=(T*S)*R (correlation)

[0028] T*S=S*T (commutativity)

[0029] Therefore, several tensors T1,…,T n The convolution result does not depend on the order of the tensors and the way brackets are used, and can be expressed as T1*…*T n .

[0030] It is also easy to see that if the convolution operations are applied to the tensor network in a certain order For all tensors involved in the tensor, tensor contraction can be obtained For example, for Figure 2 The tensor network shown in (a) can use the following convolution expression. The convolution expression is the contraction expression in this disclosure:

[0031]

[0032] It should be noted that the same result can be obtained by computing any other contraction expression of ΣN, such as (Q*T)*(S*U)*R. However, from a practical perspective, different contraction expressions usually have different computational costs, which are usually measured in terms of the number of arithmetic floating-point operations (such as addition and multiplication, FLOPs) and the number of tensor elements read and written. Different convolution expressions also have different memory budgets. The memory budget can be estimated by the maximum size of the intermediate results during the contraction of the convolution expression (i.e., the size of the intermediate convolution).

[0033] Each contraction expression can be naturally represented by a binary tree, which is usually called a convolution tree. In this disclosure, a convolution tree is a contraction tree. In this contraction tree, the leaves correspond to tensors and the internal nodes correspond to convolutions. For example, Figure 3 The contraction tree shown in (b) corresponds to the convolution expression (1).

[0034] if is a tensor with T1,…,T n The contraction tree of the tensor network, then the result of the corresponding contraction expression (i.e. contraction The result) is expressed as

[0035] In view of the above, a general object of the present disclosure is to provide a processor for a quantum simulator that is capable of simulating and / or verifying quantum circuits running on large quantum systems. A specific object is to provide a processor that is capable of performing a tensor network contraction operation that finds a good contraction expression (and corresponding contraction tree) for one or more tensor networks. This operation can then be used to simulate and / or verify quantum circuits. Specifically, the goal is to find a good contraction expression and a good contraction tree for a quantum circuit. , while making the following features as small as possible (specifically, the resulting contraction expression is considered "good" when these features are taken into account):

[0036] The amount of memory required Includes cache size and memory for intermediate shrink results.

[0037] Computational complexity That is, the number of floating-point operations (Flops), calculated as a contraction tree The sum of the complexities of all convolutions in .

[0038] ·parameter Equal to the contraction tree The number of read and write operations in memory for all convolutions in .

[0039] These and other objects are achieved by the embodiments of the present disclosure, as set out in the appended independent claims. Advantageous implementations of these embodiments are further defined in the dependent claims.

[0040] A first aspect of the present disclosure provides a processor for a quantum simulator, wherein the processor is configured to: execute a local search algorithm to determine a plurality of contraction expressions, wherein each contraction expression is suitable for contracting a corresponding tensor network in one or more tensor networks into a determined contraction tensor network; for each contraction expression, determine a contraction cost for contracting the corresponding tensor network into the determined contraction tensor network based on a cost function; select a contraction expression with the lowest contraction cost; contract each tensor network in the one or more tensor networks into the determined contraction tensor network; wherein the cost function is based at least on: a first parameter indicating an amount of memory required to contract the corresponding tensor network into the determined contraction tensor network; a second parameter indicating a computational complexity required to contract the corresponding tensor network into the determined contraction tensor network based on the selected contraction expression; and a third parameter indicating a number of read and write operations required to contract the corresponding tensor network into the determined contraction tensor network.

[0041] As described above, the contraction expression can be a convolution expression. In addition, the contraction expression corresponds to a contraction tree, so the contraction tree can be a convolution tree.

[0042] The processor of the first aspect is capable of finding (good) contraction expressions and corresponding contraction trees for contracting one or more tensor networks into a defined contracted tensor network, i.e., a contracted tensor network with a defined shape or structure. Specifically, the processor can find a contraction expression that minimizes the three aforementioned characteristics. Since the processor is configured to also contract multiple tensor networks (e.g., in parallel when multiple tensor networks have the same shape), it is suitable for MA quantum simulation. Therefore, an improved quantum simulator, in particular a MA quantum simulator, can be provided based on the processor of the first aspect, wherein the quantum simulator is capable of simulating and / or verifying quantum circuits running on large quantum systems.

[0043] In an implementation of the first aspect, the one or more tensor networks have the same graph representation.

[0044] The processor of the first aspect can efficiently find extraction expressions for multiple tensor networks (specifically, it finds a contraction expression / contraction tree of the same shape or structure for each of the multiple tensor networks).

[0045] In an implementation of the first aspect, the cost function is based on arithmetic intensity, and the arithmetic intensity is defined by a ratio of the second parameter to the third parameter.

[0046] Taking arithmetic intensity into account can achieve the best overall performance of the tensor network contraction operation performed by the processor of the first aspect because this takes into account the different memory speeds and computation speeds on specific hardware.

[0047] As mentioned above, arithmetic intensity is defined as the ratio of computational complexity (number of elementary floating-point operations) to the number of memory read / write operations during tensor contraction. The arithmetic intensity parameter cannot be less than the ratio of computational speed (Flops, floating-point operations per second) to the maximum speed of device memory (B / s). For example, in the Tesla V100 GPU, this value is approximately equal to 16. Therefore, arithmetic intensity is a device-dependent parameter and may vary for different hardware.

[0048] In an implementation of the first aspect, the cost function Determined by:

[0049]

[0050] Wherein, M is the first parameter, C is the second parameter, RW is the third parameter, M max is the upper limit of memory amount, α is the arithmetic intensity, and β is a penalty factor for controlling the weight of the first parameter.

[0051] Specifically, the penalty factor β controls the weight of memory size in the cost function. If memory budget is more important, the value of β may be increased.

[0052] In one implementation of the first aspect, executing the local search algorithm to determine the plurality of contraction expressions includes applying a plurality of local transformations, the plurality of local transformations including at least one of the following local transformations from a first contraction expression to a second contraction expression:

[0053] - from (a*b)*c to (c*b)*a;

[0054] - from (a*b)*c to (a*c)*b;

[0055] - from a*(b*c) to b*(a*c);

[0056] - from a*(b*c) to c*(b*a);

[0057] Where a, b, and c are tensors of the corresponding tensor network, and * represents a convolution operation.

[0058] In an implementation manner of the first aspect, the local search algorithm includes a hill climbing algorithm or a simulated annealing algorithm.

[0059] In one implementation of the first aspect, the corresponding tensor network has a graph representation, which includes a vertex corresponding to each tensor of the corresponding tensor network and at least one edge corresponding to each tensor leg of each tensor, and the processor is further configured to perform a slicing operation during the local search algorithm by fixing a set of tensor legs in the corresponding tensor network.

[0060] Therefore, the above advantages of slicing during the execution of the tensor network contraction operation can be achieved.

[0061] In one implementation of the first aspect, the processor is configured to update the fixed set of tensor legs after executing a determined number of steps of the local search algorithm by removing random tensor legs from the fixed set and / or adding tensor legs that minimize the first parameter to the fixed set.

[0062] Therefore, the first parameter may be advantageously reduced, ie, ideally minimized for lowest memory requirements.

[0063] In an implementation of the first aspect, the processor is further configured to select the same contraction expression with the lowest contraction cost for parallel contraction of each tensor network in multiple tensor networks with the same graph representation into the determined contracted tensor network.

[0064] The processor is therefore suitable for use in a MA quantum simulator, ie for finding multiple amplitudes or batches of amplitudes in parallel.

[0065] In one implementation of the first aspect, each contraction expression includes at least one convolution of two tensors of the corresponding tensor network in the one or more tensor networks; and / or determining the contraction cost of the contraction expression includes performing at least one convolution of two tensors of the corresponding tensor network in the one or more tensor networks.

[0066] That is, the contraction expression is a convolution expression. Due to the above characteristics of the convolution expression, an efficient tensor network contraction operation based on the local search algorithm can be performed.

[0067] In an implementation of the first aspect, the processor is configured to store a convolution result of determining the convolution of two tensors of the corresponding tensor network in a cache.

[0068] In an implementation of the first aspect, the processor is configured to reuse the convolution result saved when shrinking the first tensor network when shrinking the second tensor network.

[0069] Therefore, the computational speed of processing (contracting) multiple tensor networks in parallel can be improved.

[0070] A second aspect of the present disclosure provides a quantum simulator for simulating quantum circuits, wherein the quantum simulator includes the processor according to the first aspect or any implementation manner thereof.

[0071] In an implementation of the second aspect, the quantum simulator is configured to simulate a quantum circuit, wherein simulating the quantum circuit includes shrinking one or more tensor networks into the determined shrunken tensor network.

[0072] In an implementation of the second aspect, the quantum simulator is configured to perform verification of the quantum circuit by finding the amplitude of each of a plurality of samples generated by the quantum circuit.

[0073] In an implementation of the second aspect, the quantum simulator is configured to determine a linear cross entropy benchmark for a bit string comprising a plurality of samples generated by the quantum circuit.

[0074] A third aspect of the present disclosure provides a processing method for a quantum simulator, wherein the method includes: executing a local search algorithm to determine multiple contraction expressions, wherein each contraction expression is suitable for contracting a corresponding tensor network in one or more tensor networks into a determined contraction tensor network; for each contraction expression, determining the contraction cost of contracting the corresponding tensor network into the determined contraction tensor network based on a cost function; selecting the contraction expression with the lowest contraction cost; contracting each tensor network in the one or more tensor networks into the determined contraction tensor network; wherein the cost function is based on at least: a first parameter indicating the amount of memory required to contract the corresponding tensor network into the determined contraction tensor network based on the selected contraction expression; a second parameter indicating the computational complexity required to contract the corresponding tensor network into the determined contraction tensor network; and a third parameter indicating the number of read and write operations required to contract the corresponding tensor network into the determined contraction tensor network.

[0075] In an implementation of the third aspect, the one or more tensor networks have the same graph representation.

[0076] In an implementation of the third aspect, the cost function is based on arithmetic intensity, where the arithmetic intensity is defined by a ratio of the second parameter to the third parameter.

[0077] In an implementation of the third aspect, the cost function Determined by:

[0078]

[0079] Wherein, M is the first parameter, C is the second parameter, RW is the third parameter, M max is the upper limit of memory amount, α is the arithmetic intensity, and β is a penalty factor for controlling the weight of the first parameter.

[0080] In an implementation of the third aspect, executing the local search algorithm to determine the plurality of contraction expressions includes applying a plurality of local transformations, the plurality of local transformations including at least one of the following local transformations from a first contraction expression to a second contraction expression:

[0081] - from (a*b)*c to (c*b)*a;

[0082] - from (a*b)*c to (a*c)*b;

[0083] - from a*(b*c) to b*(a*c);

[0084] - from a*(b*c) to c*(b*a);

[0085] Where a, b, and c are tensors of the corresponding tensor network, and * represents a convolution operation.

[0086] In an implementation manner of the third aspect, the local search algorithm includes a hill climbing algorithm or a simulated annealing algorithm.

[0087] In an implementation of the third aspect, the corresponding tensor network has a graph representation, which includes a vertex corresponding to each tensor of the corresponding tensor network and at least one edge corresponding to each tensor leg of each tensor, and the method also includes performing a slicing operation during the local search algorithm by fixing a set of tensor legs in the corresponding tensor network.

[0088] In one implementation of the third aspect, the method includes: after performing a certain number of steps of the local search algorithm, updating the fixed set of tensor legs by deleting random tensor legs from the fixed set and / or adding tensor legs that minimize the first parameter to the fixed set.

[0089] In an implementation of the third aspect, the method includes: selecting the same contraction expression with the lowest contraction cost, for parallel contraction of each tensor network in multiple tensor networks with the same graph representation into the determined contracted tensor network.

[0090] In one implementation of the third aspect, each contraction expression includes at least one convolution of two tensors of the corresponding tensor network in the one or more tensor networks; and / or determining the contraction cost of the contraction expression includes performing at least one convolution of two tensors of the corresponding tensor network in the one or more tensor networks.

[0091] In an implementation of the third aspect, the method includes: storing a convolution result of determining the convolution of two tensors of the corresponding tensor network in a cache.

[0092] In an implementation of the third aspect, the method includes: when shrinking the second tensor network, reusing the convolution result saved when shrinking the first tensor network.

[0093] The method of the third aspect and its implementation achieves the same advantages and effects as the device of the first aspect and its corresponding implementation.

[0094] A fourth aspect of the present disclosure provides a computer program, comprising code, wherein when the code is executed by a processor, the processor is caused to execute the method described in the third aspect or any implementation manner thereof.

[0095] A fifth aspect of the present disclosure provides a non-transitory storage medium storing executable program code. When the executable program code is executed by a processor, the method according to the third aspect or any implementation manner thereof is executed.

[0096] It should be noted that all devices, elements, units and modules described in this application can be implemented in software or hardware elements or any combination thereof. All steps performed by the various entities described in this application and the functions described to be performed by the various entities are intended to indicate that the corresponding entities are suitable for or configured to perform the corresponding steps and functions. Although in the description of the following specific embodiments, the specific functions or steps performed by the external entity are not reflected in the description of the specific detailed elements of the entity that performs the specific steps or functions, it should be clear to the technician that these methods and functions can be implemented in the corresponding software or hardware elements or any combination thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0097] The following description of specific embodiments in conjunction with the accompanying drawings will illustrate the above aspects and implementation methods, wherein:

[0098] Figure 1 shows a simulation of a random quantum circuit running on a quantum system (Google's Sycamore 53-qubit chip);

[0099] Figure 2 (a) shows an example of a tensor network with open tensor legs;

[0100] Figure 2 (b) shows a slice during the contraction of the tensor network;

[0101] Figure 3 (a) shows the convolution operation of two tensors;

[0102] Figure 3 (b) shows Figure 2 (b) An exemplary contraction tree of a tensor network;

[0103] Figure 4 A processor according to an embodiment of the present disclosure is shown;

[0104] Figure 5 A tensor network, a contraction expression for the tensor network, and a corresponding contraction tree for the tensor network are shown;

[0105] Figure 6 A plurality of tensor networks represented by the same graph, contraction expressions of the same shape found for the tensor networks, and corresponding contraction trees of the same shape found for the tensor networks are shown;

[0106] Figure 7(a) shows the local transformation used in the local search algorithm from the perspective of the contraction tree;

[0107] Figure 7 (b) shows an example of possible steps of a local search algorithm;

[0108] Figure 8 (a) shows the tensor network and the corresponding tensor network graph;

[0109] Figure 8 (b) shows an exemplary contraction tree;

[0110] Figure 8 c shows an exemplary contraction tree subsequently obtained;

[0111] Figure 9 (a) shows an exemplary quantum circuit;

[0112] Figure 9 (b) shows Figure 9 (a) An exemplary tensor network diagram of an exemplary quantum circuit;

[0113] Figure 10 An exemplary contraction tree is shown in detail;

[0114] Figure 11 Results of single-amplitude and multi-amplitude simulations are shown;

[0115] Figure 12 A quantum simulator according to an embodiment of the present disclosure is shown;

[0116] Figure 13 A method for a quantum simulator according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0117] Figure 4 FIG4 shows a processor 400 according to an embodiment of the present disclosure. The processor 400 is configured for use in a quantum simulator (e.g., see Figure 12 The processor 400 may be specifically configured to perform a contraction of one or more tensor networks 401. Thus, the one or more tensor networks 401 may all have the same graph representation (i.e., they may have the same shape or structure). The quantum simulator may use the processor 400 to simulate and / or verify one or more quantum circuits (e.g., see Figure 9 ).

[0118] The processor 400 is configured to execute a local search algorithm 402 (e.g., a hill climbing algorithm or a simulated annealing algorithm) to determine a plurality of contraction expressions 403, such as convolution expressions. Each of the contraction expressions 403 is suitable for contracting a corresponding tensor network 401 in the one or more tensor networks 401 into a determined contracted tensor network 409. The determined contracted tensor network 409 may be a predefined target contracted tensor network. Each contraction expression 403 enables it to contract any one of the tensor networks into a determined contracted tensor network 409, i.e., into a contracted tensor network having a specific shape, form, or configuration.

[0119] The processor 400 is further configured (e.g., at optional block 404) to determine a contraction cost 406 for each contraction expression 403, wherein the contraction cost 406 is the cost required to contract the corresponding tensor network 401 into the determined contracted tensor network 409. The contraction cost 406 is determined based on a cost function 405. The cost function 405 is based on at least three parameters: a first parameter indicating the amount of memory required to contract the corresponding tensor network 401 into the determined contracted tensor network 409. a second parameter indicating the computational complexity required to contract the corresponding tensor network 401 into the determined contracted tensor network 409. and a third parameter indicating the number of read and write operations required to contract the corresponding tensor network 401 into the determined contracted tensor network 409.

[0120] The processor 400 is further configured to (e.g., at optional block 407) select a contraction expression 403 with the lowest contraction cost 406 and contract 408 each of the one or more tensor networks 401 into a determined contracted tensor network 409 based on the selected contraction expression 403. That is, the processor 400 may use the selected contraction expression 403 to contract each of the tensor networks 401. Thus, the "selected contraction expression" has the same form or shape for each tensor network. Of course, different tensors of the tensor network are also reflected as different tensors in the contraction expression because it is used to contract the corresponding tensor network 401.

[0121] The processor 400 may include hardware (processing circuitry) and / or may be controlled by software. The hardware may include analog circuitry or digital circuitry, or both. The digital circuitry may include components such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or a multi-purpose processor. The processor 400 may also include a memory circuit that stores one or more instructions that can be executed by the processing circuitry (specifically, executed under the control of software). For example, the memory circuitry may include a non-transitory storage medium that stores executable software code that, when executed by the processing circuitry, causes the processor 400 to perform various operations.

[0122] Figure 5 A single-amplitude (SA) tensor network contraction operation that may be performed by the processor 400 is shown. Specifically, an exemplary corresponding tensor network 401 having a particular graph representation 501 is shown. The graph representation 501 includes a vertex 502 corresponding to each tensor 504 of the corresponding tensor network 401 and at least one edge 503 corresponding to each tensor leg of each tensor 504. The processor 400 is configured to find and then select a "good" contraction expression 403 as described above. The contraction expression 403 may include at least one convolution 505 of two tensors 504 of the corresponding tensor network 401. Determining the contraction cost 406 of the contraction expression 403 may include performing at least one convolution 505 of the two tensors 504. The convolution 505 may be as described above with respect to Figure 3 The contraction expression 403 can be represented by the corresponding contraction tree 506 .

[0123] Figure 6 A (MA) tensor network contraction operation that can be performed by the processor 400 is shown. For the MA tensor network contraction operation, m tensor networks 401 having the same graph representation 501 (i.e., the same shape and / or the same tensor graph or the same structure) but having different specific tensors 504 can be contracted. The processor 400 is configured to select the same contraction expression 403 (i.e., the shape or structure of the contraction expression is the same, but of course, with different tensors 504 reflecting the specific tensors 504 of different tensor networks 401) with the lowest contraction cost 406 when contracting each tensor network in the tensor network 401 in parallel into a determined contracted tensor network 409. The contraction expression 403 results in the corresponding contraction tree 506 of each tensor network 401 being the same (the shape or structure is the same, but with different specific tensors 504).

[0124] More details regarding embodiments of the present disclosure, particularly the processor 400 and the quantum simulator 1200, are described below.

[0125] In order to find a near-optimal contraction tree 506 (and therefore a "good" contraction expression 503), any local search algorithm 402, such as hill climbing or simulated annealing, can be used with a cost function 405 (also referred to in this disclosure as the "objective function"). ”). The local search algorithm 402 is a combinatorial optimization method, given an objective function f in the search space X of all possible states: Try to apply a small fixed number of local transformations L = {l1,…,l n}(Each transformation l i is the function l i :X→X). Therefore, we get the state sequence x0,x1,…,x N , where each next state x i+1 is to use one of the local transformations from the set L from the previous state x i Obtained, that is, for a certain l∈L, x i+1 =l(x i ).

[0126] The choice of the local transformation l∈L at each step is usually controlled by the gain,

[0127] Δf(x i ,l):=f(l(x i ))-f(x i )

[0128] The gain is obtained according to the objective (cost) function f. The local search algorithm 402 usually stops when it reaches a point where it cannot improve locally (i.e., for all l∈L, Δf(x i ,l)<0) N Or stop when the maximum number of steps is reached. There are many other details about how to implement the local search algorithm 402. For example, in the simulated annealing algorithm, the local transformation can be randomly selected with a probability that depends on the gain Δf(x i , 1). Meanwhile, in the hill climbing algorithm, local transformations can be deterministically selected in a greedy manner (i.e., the local transformation that provides the best possible gain can be selected). However, in this disclosure, only the objective function 406 and the local transformations are described. Therefore, embodiments of the present disclosure can be used with any local search algorithm 402 (such as hill climbing or simulated annealing).

[0129] The above shows that the convolution operation T*S of two or more tensors 504 satisfies the associativity and commutativity conditions. Using these two conditions, the following identities can be derived. These can be used as local transformations in the local search algorithm 402:

[0130] (a*b)*c→(c*b)*a

[0131] (a*b)*c→(a*c)*b

[0132] a*(b*c)→b*(a*c)

[0133] a*(b*c)→c*(b*a)

[0134] 403 (left side) is locally transformed (→) into a second contracted expression 403 (right side). Thus, executing the local search algorithm 402 (by the processor 400) includes applying a plurality of local transformations, including at least one of the above-described local transformations, where a, b, and c are tensors 504 of the respective tensor networks 401, and * represents a convolution operation.

[0135] Figure 7 (a) shows the four local transformations from the perspective of the contraction tree 506 (where the triangles correspond to subtrees of the contraction tree 506). The state set in the local search (contraction tree optimization) algorithm 502 is a tensor network The set of all possible contraction trees 506 can also be interpreted as a set of trees for finding contractions The resulting convolution expression 403. It can be assumed that at each step of the local search algorithm 402, one of these four local transformations can be applied to any sub-expression (i.e., to the contraction tree Subtree ). Figure 7 (b) shows an example of some possible steps of the local search algorithm 402.

[0136] It should be noted that some additional slicing-related operations can also be added to the local transformation list. As described above, slicing (also known as "variable projection") can be used to reduce the memory budget at the expense of a moderate increase in computational complexity. That is, the processor 400 can also be configured to perform slicing operations during the execution of the local search algorithm 402, specifically by fixing the tensor leg set 503 in the corresponding tensor network 401.

[0137] Therefore, in an embodiment of the present disclosure, the following slight modification of the above-mentioned local search algorithm 402 is proposed. Every K steps of the local search algorithm 402, the list S of legs for slicing can be updated. With probability 1 / 2, one of the following two additional steps can be performed: adding the legs that lead to a reduction in the optimal memory budget M to the list S; deleting random legs from S. This means that after executing a certain number of steps of the local search algorithm 402, the processor 400 can update the fixed set of tensor legs by deleting random tensor legs from the fixed set and / or adding the tensor legs that maximize the reduction of the first parameter (indicating the memory budget) to the fixed set. If K is relatively large (for example, K=10 5 ), the two additional steps are not often applied, and the overall runtime of the local search method hardly changes with the slight modification.

[0138] Neither the computational complexity nor the memory budget depends on the contents of the tensors 502 in the contraction tree 506. In fact, they only depend on the tensors T1, ..., T n Therefore, it is helpful to consider the formal expression, where, in the convolution expression 403, there exists a representation with tensors T1,…,T n Any tensor 504 variables of the same shape X1,…,X n , rather than some fixed tensor. Each convolutional tree has n leaves It can also be viewed as a convolution expression of the form Therefore, it can be seen that the shrinkage tree Just express the convolution expression In addition, it can be seen that the convolution expression The contraction tree Subtree Represents its subexpression Therefore, in the following, a formal contraction expression 403 and a corresponding contraction tree are identified.

[0139] like Figure 8 As shown in (a), a tensor network graph (graph representation 501 of tensor network 401) is simply a tensor network D, where there are variables X1, ..., X corresponding to arbitrary tensors of the same shape. n , rather than fixed tensors of some shape T1,…,T n To emphasize its variables, the tensor network graph can be represented as D(X1,…,X n ). If the tensors T1,…,T n Assigned to variables X1,…,X n , we can obtain the value of D(T1,…,T n ). The contraction result of this tensor network is expressed as ΣD(T1,…,Tn ).

[0140] In a MA (or multi-batch) quantum simulator, one typically looks for N different amplitudes (or batches). A simple approach is to run the single amplitude or single batch contraction algorithm N times. However, when we need to look for a large number (on the order of 10 6 This simple approach is not efficient when the number of cells is large or large. The following describes how to use processor 400 to accomplish this in a more efficient manner.

[0141] Given an exemplary quantum circuit C, it can be converted into a tensor network in a standard way The tensor network There are s open legs, where s is the number of qubits in the exemplary quantum circuit (each open leg corresponds to an output qubit). Further let D = D(X), X = (X1, ..., X n ) is a tensor variable X1,…,X n of As mentioned above, in the MA simulator, we need to find N bit strings s1,…,s N ∈{0,1} n N complex amplitudes i |C|0 n >. This can be used as a network of N tensors D(T 1 ),…,D(T N ) is obtained by the contraction result, where each tensor set Corresponding to a bit string s i (distribute Output leg of ). Given a contraction tree (e.g., found by the previously described local search algorithm 402), which can be used to execute N tensor networks D(T 1 ),…,D(T N ) and obtain: Therefore, it can be seen that MA contraction is equivalent to the following task. Let D be the tensor network graph, is formed by some contraction expression (e.g., The corresponding contraction tree given (see Figure 8 (b)).

[0142] The goal is to create multiple collections of tensors Find A key observation is that if these N contractions are performed sequentially for i = 1, 2, ..., N, and if Some sub-expressions of has shrunk, then in the variable ​If the value of is the same, the result can be reused in the next contraction (see Figure 8 c). In other words, the processor 400 may be configured to reuse the convolution results saved when shrinking the first tensor network when shrinking the second tensor network.

[0143] Therefore, we can consider a recursive algorithm for computing these N contractions (it can be called a multi-tensor contraction process because it produces N tensors) that stores the intermediate results of all its recursive calls in some global cache K. We can assume that K can be updated during the work of the algorithm. This cache K can be implemented as a key-lookup data structure, where the key is a set of tensors and the value is a tensor. If there is an entry T for a tensor set v, we can write K(v) = T. Otherwise, if there is no available entry for the key v, we can write K(v) = null. Therefore, the following process can be used for multiple tensor sets Find

[0144] K = empty; / / Start with an empty global cache

[0145] for i=1,…,N do: (for i=1,…,N execute)

[0146] (calculate )

[0147] (return )

[0148] It can be seen that in the above process, given the contraction tree Tensor set T i , calling the subroutine times, while the intermediate results of the previous calls give That is, the expression The result of evaluating the set of tensors T. A description of this process is provided below.

[0149] (process )

[0150] (like Then return )

[0151] (make for variables)

[0152] (like but)

[0153] (make in and for The left and right subtrees of

[0154] (Calculate tensor

[0155] Calculate U: =U L *U R ;(Calculate U:=U L *U R ) / / Perform convolution operation;

[0156] / / Store the result U in cache K;

[0157] (Otherwise return / / The value K is already in K.

[0158] The idea of ​​the above algorithm can be demonstrated on a simple example considering a quantum circuit 900 with 3 qubits and the corresponding tensor graph (see Figure 9 ). In this tensor graph D(X0,…,X8), for simplicity, tensor variables X0,…,X8 are represented by numbers 0-8, their tensor legs are represented by numbers 0-7, and open tensor legs are represented by numbers 8-10.

[0159] If the three binary values ​​s1, s2, s3 of the output qubits are fixed (i.e., the values ​​of open legs 8, 9, 10 are fixed), then the values ​​T0, ..., T8 of all tensor variables X0, ..., X8 in Figure D are fixed, and the complex amplitude of the bit string s1s2s3 is <s1s2s3|C|0 n >Equal to the result of contraction: ΣD(T0,…,T8)=ΣD(T), where T=(T0,…,T8).

[0160] The goal is to find the complex amplitude of the following three bit strings: 000, 100, 111. First, a contraction tree can be found (such a tree can be found using the local search algorithm 402 described above). For example, for Figure 9 The quantum circuit C shown in (a) Figure 9 (b) shows the corresponding tensor network 401 of the quantum circuit 900 (in a graph representation) which can be expressed using the following tree given by the convolution expression

[0161]

[0162] To find the N=3 complex amplitudes of the bit string s1s2s3∈{000,100,111} <s1s2s3|C|0 n >, it is necessary to find and Among them, the tensor Each vector of corresponds to three bit strings 000, 100, and 111 respectively. Figure 10 Shows an annotated convolutional tree There, for each inner tree node corresponding to a convolution, legs 0-7 are shown (they are summed in this convolution), and open tensor legs 8-10 are shown (the values ​​of these legs are fixed when the bit strings s1s2s3 are fixed).

[0163] The possible values ​​of the open tensor legs are also shown. The number of these values ​​shows how many times the contraction must be performed for this subtree. For example, for the subtree (X0*X3)*(X1*X5), there are no open legs, so when looking for It only needs to be calculated once, and the result can be found in and At the same time, for the subtree (X2*X4)*(X6*X8), open legs 9 and 10 have two possible values ​​(00 and 11), so this subtree needs to be contracted twice. However, if a 3-times single amplitude simulator is used, each subtree needs to be contracted three times.

[0164] The embodiments presented in this disclosure, specifically the algorithms executed by processor 400, can be applied to the verification of large random quantum circuits, such as those used in Google's recent quantum supremacy experiment. They can also be used for other tasks that require finding many amplitudes that are not arranged in batches. Figure 11 The figure shows the complexity (measured in FLOPs) of N runs of the SA simulator and one run of the MA simulator for Google's Sycamore circuit according to an embodiment of the present disclosure. It can be seen that the complexity of the MA simulator is thousands of times smaller than that of the SA simulator.

[0165] Figure 12 A quantum simulator 1200 according to an embodiment of the present disclosure is shown. The quantum simulator 1200 may be the aforementioned MA simulator. The quantum simulator includes a processor 400 as described in the present disclosure (e.g., see Figure 4 ). The quantum simulator 1200 is configured to simulate the quantum circuit 900, for example Figure 9During the simulation of quantum circuit 900, quantum simulator 1200 is configured to use processor 400 to collapse one or more tensor networks 401, as described above. Specifically, quantum simulator 1200 can use processor 400 to perform verification of quantum circuit 900 according to the aforementioned MA simulation, i.e., by finding the amplitude of each of the plurality of samples generated by quantum circuit 900.

[0166] Figure 13 1 shows a method 1300 according to an embodiment of the present disclosure. The method 1300 may be executed by the processor 400 or the quantum simulator 1200 using the processor 400.

[0167] The method includes step 1301: executing a local search algorithm 402 to determine a plurality of contraction expressions 403. Each contraction expression 403 is suitable for contracting a corresponding tensor network 401 in the one or more tensor networks 401 into a determined contracted tensor network 409. In addition, the method 1300 includes step 1302: for each contraction expression 403, determining a contraction cost 406 for contracting the corresponding tensor network 401 into the determined contracted tensor network 409 based on a cost function 405. Then, the method 1300 includes step 1303: selecting the contraction expression 403 with the lowest contraction cost 406; and step 1304: contracting each tensor network 401 in the one or more tensor networks 401 into the determined contracted tensor network 409 based on the selected contraction expression 403. The cost function 405 used in step 1302 is based on at least: a first parameter indicating the amount of memory required to shrink 1304 the corresponding tensor network 401 into the determined shrunken tensor network 409; a second parameter indicating the computational complexity required to shrink 1304 the corresponding tensor network 401 into the determined shrunken tensor network 409; and a third parameter indicating the number of read and write operations required to shrink 1304 the corresponding tensor network 401 into the determined shrunken tensor network 409.

[0168] The present disclosure has been described with reference to various embodiments and implementations as examples. However, other variations will be apparent to and will be realized by those skilled in the art in practicing the claimed subject matter, based on a study of the drawings, the present disclosure, and the independent claims. In the claims and the specification, the word "comprising" does not exclude other elements or steps, and "a" or "an" does not exclude a plurality. A single element or other unit may fulfil the functions of several entities or items recited in the claims. The recitation of certain measures in mutually different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation.

Claims

1. A processor (400) for a quantum simulator (1200), characterized in that The processor (400) is configured to: executing a local search algorithm (402) to determine a plurality of contraction expressions (403), wherein each contraction expression (403) is suitable for contracting a corresponding tensor network (401) of the one or more tensor networks (401) into a determined contracted tensor network (409); For each contraction expression (403), determining (404) a contraction cost (406) for contracting the corresponding tensor network (401) to the determined contracted tensor network (409) based on a cost function (405); Select (407) the contraction expression (403) with the lowest contraction cost (406); Based on the selected contraction expression (403), contracting (408) each of the one or more tensor networks (401) into the determined contracted tensor network (409); Wherein, the cost function (405) is based at least on: − a first parameter indicating the amount of memory required to shrink the corresponding tensor network ( 401 ) into the determined shrunken tensor network ( 409 ); − a second parameter indicating the computational complexity required to shrink the corresponding tensor network ( 401 ) to the determined shrunken tensor network ( 409 ); − a third parameter indicating the number of read and write operations required to shrink the corresponding tensor network ( 401 ) into the determined shrunken tensor network ( 409 ); Wherein, the cost function (405) Determined by: Wherein, M is the first parameter, C is the second parameter, RW is the third parameter, M max Is the upper limit of memory, is the arithmetic strength, β is a penalty factor for controlling the weight of the first parameter; and the arithmetic strength is defined by the ratio of the second parameter to the third parameter.

2. The processor (400) according to claim 1, characterized in that The one or more tensor networks have the same graph representation (501).

3. The processor (400) according to claim 1, characterized in that Executing the local search algorithm (402) to determine the plurality of contraction expressions (403) includes applying a plurality of local transformations, the plurality of local transformations including at least one of the following local transformations from a first contraction expression (403) to a second contraction expression (403): −From (a b) c to (c b) a; −From (a b) c to (a c) b; −From a (b c) to b) (a c); −From a (b c) to c) (b a); Wherein, a, b and c are tensors of the corresponding tensor network (401), Represents the convolution operation.

4. The processor (400) according to claim 1, characterized in that The local search algorithm (402) includes a hill climbing algorithm or a simulated annealing algorithm.

5. The processor (400) according to claim 1, characterized in that: The corresponding tensor network (401) has a graph representation (501) comprising a vertex (502) corresponding to each tensor (504) of the corresponding tensor network (401) and at least one edge (503) corresponding to each tensor leg of each tensor (504), and The processor (400) is further configured to perform a slicing operation during the local search algorithm (402) by fixing a set of tensor legs in the corresponding tensor network (401).

6. The processor (400) according to claim 5, characterized in that The processor (400) is configured to update the fixed set of tensor legs after executing the determined number of steps of the local search algorithm (402) by removing random tensor legs from the fixed set and / or adding tensor legs that minimize the first parameter to the fixed set.

7. The processor (400) according to any one of claims 1 to 6, characterized in that The processor (400) is further configured to select the same contraction expression (403) with the lowest contraction cost (406) for contracting each tensor network in the plurality of tensor networks (401) having the same graph representation (500) into the determined contracted tensor network (409) in parallel.

8. The processor (400) according to any one of claims 1 to 6, characterized in that: Each contraction expression (403) comprises at least one convolution (505) of two tensors (504) of said corresponding tensor network (401) of said one or more tensor networks (401); and / or Determining a contraction cost (406) of the contraction expression (403) includes performing at least one convolution (505) of two tensors (504) of the corresponding tensor network (401) of the one or more tensor networks (401).

9. The processor (400) according to claim 8, characterized in that The processor (400) is configured to store in a cache a convolution result of a convolution (505) of two tensors (504) determining the corresponding tensor network (401).

10. The processor (400) according to claim 9, characterized in that The processor (400) is configured to reuse the convolution results saved when shrinking the first tensor network when shrinking the second tensor network.

11. A quantum simulator (1200) for simulating a quantum circuit (900), characterized in that The quantum simulator (1200) comprises a processor (400) according to any one of claims 1 to 10.

12. The quantum simulator (1200) according to claim 11, characterized in that The quantum simulator (1200) is configured to simulate a quantum circuit (900), wherein simulating the quantum circuit (900) comprises contracting one or more tensor networks (401) into the determined contracted tensor network (409).

13. The quantum simulator (1200) according to claim 11, characterized in that The quantum simulator (1200) is configured to perform verification of the quantum circuit (900) by finding the amplitude of each of a plurality of samples generated by the quantum circuit (900).

14. The quantum simulator (1200) according to claim 11, characterized in that The quantum simulator (1200) is configured to determine a linear cross entropy benchmark for a bit string comprising a plurality of samples generated by the quantum circuit (900).

15. A processing method (1300) for a quantum simulator (1200), characterized in that The method comprises: executing ( 1301 ) a local search algorithm ( 402 ) to determine a plurality of contraction expressions ( 403 ), wherein each contraction expression ( 403 ) is suitable for contracting a corresponding tensor network ( 401 ) of the one or more tensor networks ( 401 ) into a determined contracted tensor network ( 409 ); For each contraction expression (403), determining (1302) a contraction cost (406) for contracting the corresponding tensor network (401) to the determined contracted tensor network (409) based on a cost function (405); Select (1303) the contraction expression (403) with the lowest contraction cost (406); Based on the selected contraction expression (403), contracting (1304) each of the one or more tensor networks (401) into the determined contracted tensor network (409); Wherein, the cost function (405) is based at least on: − a first parameter indicating the amount of memory required to shrink ( 1304 ) the corresponding tensor network ( 401 ) into the determined shrunken tensor network ( 409 ); − a second parameter indicating the computational complexity required to shrink ( 1304 ) the corresponding tensor network ( 401 ) into the determined shrunken tensor network ( 409 ); − a third parameter indicating the number of read and write operations required to shrink ( 1304 ) the corresponding tensor network ( 401 ) into the determined shrunken tensor network ( 409 ); Wherein, the cost function (405) Determined by: Wherein, M is the first parameter, C is the second parameter, RW is the third parameter, M max Is the upper limit of memory, is the arithmetic strength, β is a penalty factor for controlling the weight of the first parameter; and the arithmetic strength is defined by the ratio of the second parameter to the third parameter.

16. A computer program comprising code, characterized in that When the code is executed by a processor (400), the processor (400) is caused to perform the method (1300) according to claim 15.

Citation Information

Patent Citations

  • Simulating quantum circuits

    CN111052122A

  • Quantum circuit simulation method, device and equipment and storage medium

    CN111738448A