Quantum Error Mitigation Acceleration Method and Accelerator Utilizing Sparsity in Tensor Product

By identifying the sparsity and parallelism in the process of quantum error mitigation, the accelerator uses sparsity and parallelism technologies to optimize the computing path, solving the problems of high computational complexity and insufficient fidelity in the existing technology, and achieving efficient and accurate quantum computing acceleration.

CN119831067BActive Publication Date: 2025-07-11ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510329503.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-11
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

Existing quantum computers face the problems of high computational complexity, large resource consumption and insufficient fidelity in relieving read errors. Especially in large-scale quantum systems, tensor product-based methods have bottlenecks of delay and increased memory requirements, and the existing technology is difficult to achieve efficient and accurate end-to-end acceleration.

Method used

By identifying the sparseness in the process of quantum error mitigation, using the sparseness of Hilbert space and ground states, a quantum error mitigation acceleration method and accelerator is designed, using sparse acceleration, parallel acceleration and operation fusion technology, optimizing the computing path, reducing zero-value calculations, and designing a special hardware architecture to achieve end-to-end acceleration.

Benefits of technology

It significantly improves computing efficiency, reduces computing load, improves resource utilization and overall computing efficiency, while maintaining high error mitigation fidelity, achieving high-performance quantum computing acceleration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119831067B_ABST
    Figure CN119831067B_ABST
Patent Text Reader

Abstract

The present invention discloses a quantum error mitigation acceleration method and an accelerator utilizing sparsity in the tensor product, belonging to the technical field of quantum computing. The method includes: establishing a quantum error mitigation transformation equation; dividing the quantum error mitigation transformation equation into three-step operations of matrix-vector multiplication, tensor product, and multiply-accumulate; achieving sparsity acceleration based on output sparsity and input sparsity; performing probabilistic-level and state-level parallel computations on the tensor product operation to achieve sparsity-based parallel acceleration; and alternately performing row-major and column-major generation operations on the tensor product operation in iterative rounds to achieve operation fusion acceleration based on sparsity and parallelism. The present invention can improve the computational efficiency and fidelity of the quantum readout error mitigation method, thereby further promoting the development and application of quantum computing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of quantum computing, and particularly relates to a quantum error mitigation acceleration method and an accelerator that utilize sparsity in tensor products. Background Art

[0002] As a revolutionary advancement in the field of information technology, quantum computers have demonstrated significant advantages over classical computers in solving specific types of problems, such as combinatorial optimization and machine learning. However, current quantum computers are mainly responsible for the computational part, and the remaining operations such as encoding, decoding, and post-processing still rely on classical computers. Many so-called "quantum acceleration" claims often overlook the latency issues of classical computers, which may lead to the inability to achieve end-to-end acceleration. To implement an end-to-end quantum accelerator, not only the encoding and compilation processes of qubits need to be concerned, but also many peripheral processes that rely on classical computing must be considered. Facing the heterogeneous and decoupled quantum computer architecture, peripheral accelerators are crucial for achieving true quantum advantage. Among these processes, quantum readout error mitigation (QEM) is widely regarded as a major bottleneck in realizing quantum accelerators. To overcome this challenge, precise calibration of noise errors is necessary to ensure high-fidelity output results.

[0003] To address this challenge, academia and industry are actively exploring solutions at the software and hardware levels. At the software level, researchers have developed various algorithms and techniques to improve the efficiency and accuracy of readout error mitigation. For example, by measuring subsets of qubits to reduce readout crosstalk errors, using Hamming weight graphs and iterative Bayesian networks for modeling, and adopting bit-flip techniques to compensate for the probability of reading "1". Although these methods have alleviated the readout error problem to a certain extent, they still face the problems of high computational complexity and large resource consumption. At the hardware level, researchers are committed to designing more efficient readout circuits and discriminators to improve the accuracy of single-shot qubit state discrimination. The hardware discriminator in the HERQULES project is a typical example, which uses the trajectory information of readout frequencies to improve the discrimination accuracy. In addition, the SpREM project proposed a hardware-software co-design method aimed at reducing computational complexity by utilizing Hamming sparsity in the mitigation matrix. Although these methods have made certain progress, they still need to be further optimized and improved to achieve more efficient and accurate readout error mitigation.

[0004] In addition to efforts at the software and hardware levels, improving the fidelity after error mitigation is also a major challenge in current research. This includes structure-based error mitigation methods, deep learning-based methods, and algebraic methods, etc. Among them, algebra-based error mitigation is a white-box and formal method that has proven effective in superconducting and optical quantum hardware. The technologies in this direction can be further divided into matrix-based error mitigation and iterative tensor product-based methods. Matrix-based error mitigation is achieved by multiplying a noisy probability vector by a mitigation matrix, which makes the probability distribution closer to the ideal distribution. However, as the number of qubits increases, the scale of the matrix grows exponentially, resulting in a sharp increase in computational workload and a sharp rise in memory requirements. The tensor product-based method is essentially an approximation of the matrix-based method, aiming to reduce the computational workload caused by the exponential growth of the matrix scale. Although the tensor product-based mitigation method is software-friendly and effective in experiments, it still inevitably faces the dilemma between latency and fidelity. Especially in large-scale quantum systems, the extension of mitigation time and the increase in memory requirements have become key factors restricting its application.

[0005] In summary, currently facing the problem of quantum readout errors, leveraging the sparsity in the tensor product provides a new acceleration approach for quantum error mitigation. Nevertheless, existing techniques mainly rely on pruning intermediate data or threshold pruning strategies, and these methods are doubly restricted by limited quantum accelerator hardware resources and fidelity loss in practical applications. Therefore, it is necessary to further improve the computational efficiency and fidelity of mitigation methods to further promote the development and application of quantum computing technology. Summary of the Invention

[0006] In view of the above, the object of the present invention is to provide a quantum error mitigation acceleration method and accelerator that utilize the sparsity in the tensor product, and use the inherent sparsity of the Hilbert space to achieve quantum error mitigation acceleration. The key lies in that the superposition state generated by the quantum algorithm only covers a subset of the Hilbert space (output sparsity), while the ground state naturally exhibits bit-level sparsity (input sparsity). The sparse data stream characterized by probability-level and state-level parallelism effectively skips the calculations related to zero values, and realizes end-to-end acceleration by designing a hardware architecture that supports various tensor product calculation paradigms to reduce readout errors.

[0007] To achieve the above object of the invention, the technical solutions provided by the present invention are as follows:

[0008] A quantum error mitigation acceleration method that utilizes the sparsity in the tensor product provided by an embodiment of the present invention includes the following steps:

[0009] Multiply the noise probability vector of the quantum readout by the mitigation matrix to establish a quantum error mitigation transformation equation. Divide the mitigation matrix into several groups of sub-mitigation matrices and divide the basis vectors of the noise probability vector into qubits corresponding to the several groups of sub-mitigation matrices;

[0010] Divide the quantum error mitigation transformation equation into three steps: matrix-vector multiplication, tensor product, and multiply-accumulate. Matrix-vector multiplication performs the multiplication calculation between the mitigation matrix and the basis vector and flattens it to obtain an intermediate value. The tensor product performs the tensor product calculation of the intermediate value to obtain an intermediate vector and flattens it into an intermediate matrix. Multiply-accumulate performs the multiplication and accumulation of the intermediate matrix and the noise probability vector to obtain the mitigation probability vector as the output;

[0011] Pre-obtain the non-zero elements in the mitigation probability vector based on its output sparsity to capture the elements that need to be multiplied and accumulated and trace back to the tensor product operation and the multiply-accumulate operation. Based on the input sparsity of the basis vector, convert the matrix-vector multiplication into directly extracting the columns in the mitigation matrix to achieve sparsity acceleration;

[0012] Perform row-wise parallel calculation and column-wise parallel calculation on the intermediate vector in the tensor product operation to generate the rows and columns of the intermediate matrix respectively, achieving parallelism acceleration based on sparsity;

[0013] Alternately perform row-major order and column-major order in the iteration rounds of the tensor product operation to generate the calculation of the intermediate matrix, so as to reduce the waiting time of the multiply-accumulate operation between adjacent iteration rounds, achieving operation fusion acceleration based on sparsity and parallelism.

[0014] Preferably, the step of multiplying the noise probability vector of the quantum readout by the mitigation matrix to establish a quantum error mitigation transformation equation includes:

[0015] Multiply the noise probability vector of the quantum readout by the mitigation matrix through a matrix-based quantum error mitigation method to obtain a mitigation probability vector, so that the noise probability vector approaches the ideal probability vector. The formula of the quantum error mitigation transformation equation is:

[0016] ,

[0017] where, represents the mitigation probability vector, represents the mitigation matrix, represents the noise probability vector.

[0018] Preferably, the step of dividing the mitigation matrix into several groups of sub-mitigation matrices and dividing the basis vectors of the noise probability vector into qubits corresponding to the several groups of sub-mitigation matrices includes:

[0019] With a quantum error mitigation method based on the tensor product, the mitigation matrix in the quantum error mitigation transformation equation is partitioned into groups of sub-mitigation matrices , then the quantum error mitigation transformation equation formula is transformed into:

[0020] ,

[0021] wherein, represents the mitigation probability vector, represents the noise probability vector, represents the tensor product;

[0022] The basis vectors of the noise probability vector are further sliced into qubits corresponding to the groups of sub-mitigation matrices , then the quantum error mitigation transformation equation formula is transformed into:

[0023] ,

[0024] wherein, represents the noise probability vector in the state of the basis vector , represents the set of non-zero probability basis vectors in after qubit measurement.

[0025] Preferably, the output sparsity based on the mitigation probability vector pre-obtains the non-zero elements therein to capture the elements that need to perform multiplication and accumulation and traces back to the tensor product operation and the multiply-accumulate operation, including:

[0026] Based on the output sparsity of the mitigation probability vector, pre-obtain the non-zero elements in the mitigation probability vector, trace back to the multiply-accumulate operation according to each non-zero element to determine the elements that need to perform multiplication and accumulation, and further trace back to the tensor product operation step to determine the elements that need to perform the tensor product operation.

[0027] Preferably, the input sparsity based on the basis vector converts the matrix-vector multiplication into directly extracting the columns in the mitigation matrix, including:

[0028] The basis vectors in the quantum system are represented as unit vectors. By identifying the non-zero indices in the basis vectors, the matrix-vector multiplication operation is converted into directly selecting the corresponding columns from the mitigation matrix according to the non-zero indices.

[0029] Preferably, performing row-wise parallel calculation and column-wise parallel calculation on the intermediate vectors in the tensor product operation to generate the rows and columns of the intermediate matrix respectively, including:

[0030] Identify orthogonal probabilistic parallelism and state-level parallelism in the tensor product operation. Probabilistic parallelism means parallelly computing multiple intermediate vectors, and state-level parallelism means parallelly computing multiple elements in each intermediate vector. Perform row-wise parallel computing on the elements of each row of the intermediate vectors in the tensor product operation according to probabilistic parallelism, and at the same time perform column-wise parallel computing on the elements of each column of the intermediate vectors in the tensor product operation according to state-level parallelism, so as to generate a block in the intermediate matrix in each computing cycle, and finally output the entire intermediate matrix.

[0031] Preferably, for the values in the same row of each block, their operands and row indices are the same, and accordingly, operand reuse and index reuse are performed when calculating the values in the same row of each block.

[0032] Preferably, the tensor product operation alternately performs row-major order and column-major order in iterative rounds to generate the calculation of the intermediate matrix so as to reduce the waiting time of the multiply-accumulate operation between adjacent iterative rounds, including:

[0033] In the current iterative round, the tensor product operation follows the row-major order to generate the intermediate matrix row by row, and in the next iterative round, the tensor product operation follows the column-major order to generate the intermediate matrix column by column. Then, after the tensor product operation in the next iterative round, the multiply-accumulate operation can be independently executed without waiting for all the multiply-accumulate operations in the current iterative round to be completed.

[0034] To achieve the above invention purpose, an embodiment of the present invention also provides a quantum error mitigation accelerator that utilizes the sparsity in the tensor product, which is implemented based on the above-mentioned quantum error mitigation acceleration method that utilizes the sparsity in the tensor product, including: a matrix-vector multiplication calculation unit, a tensor product calculation unit, a multiply-accumulate calculation unit, and a memory;

[0035] The matrix-vector multiplication calculation unit is used to perform a column selection operator on the mitigation matrix according to the strings in the basis vectors, obtain the columns of the mitigation matrix, and flatten them into intermediate values;

[0036] The tensor product calculation unit is used to process the tensor product data stream, in a probabilistic parallel and state-level parallel manner, calculate to obtain intermediate vectors and flatten them into an intermediate matrix;

[0037] The multiply-accumulate calculation unit is used to receive the results from the tensor product calculation unit in a block-by-block manner. Each block of the intermediate matrix is generated in a row-major order or column-major order and multiplied by the noise probability vector, and the results of different blocks are accumulated and aggregated to obtain a mitigation probability vector as the output;

[0038] The memory is used to store the input mitigation matrix and the output mitigation probability vector respectively.

[0039] Preferably, a cascaded multiplexer is designed in the matrix-vector multiplication calculation unit to separately select strings in the basis vectors and work in a cascaded structure; a processing unit including a multiplier and a demultiplexer is designed in the tensor product calculation unit, where the multiplier is used to perform tensor product calculations, and the demultiplexer is used to select input operands or intermediate results of the tensor product operation and configure the output path to directly output the calculation result or transfer it to the next processing unit through a register.

[0040] Compared with the prior art, the beneficial effects of the present invention at least include:

[0041] (1) By dividing the calculation into three operators: matrix-vector multiplication, tensor product, and multiply-accumulate, the present invention further identifies opportunities for accelerating the inherent sparsity. By constructing an efficient sparse mitigation data stream, it can achieve sparsity acceleration based on output sparsity and input sparsity. On this basis, parallel acceleration based on probability level and state level is realized, and further iterative acceleration based on the cooperation of row-major order and column-major order is realized, achieving high data reuse and high throughput, thereby effectively reducing the calculation load and improving the calculation efficiency.

[0042] (2) By designing a special accelerator hardware architecture, the present invention can connect the sparse mitigation data stream in the acceleration method. By designing a cascaded multiplexer in the matrix-vector multiplication calculation unit, it can optimize the processing efficiency of the sparse data stream, reduce invalid calculations, improve the resource utilization rate of the calculation unit, and by designing a reconfigurable processing unit in the tensor product calculation unit, it can dynamically handle various parameter configurations in the iterative tensor product, realize the efficient utilization of hardware resources and the flexible allocation of calculation tasks, and improve the overall calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0044] Figure 1 is a schematic flowchart of a quantum error mitigation acceleration method using sparsity in a tensor product provided by an embodiment of the present invention;

[0045] Figure 2 is a schematic flowchart of a quantum readout error mitigation calculation based on iterative tensor product provided by an embodiment of the present invention;

[0046] Figure 3 is a schematic diagram of the delay decomposition based on the BV algorithm and the sparsity comparison of quantum algorithms provided by an embodiment of the present invention;

[0047] Figure 4 It is a schematic diagram of the sparse mitigation data flow of the parallelized tensor product operation provided by the embodiments of the present invention;

[0048] Figure 5 It is a schematic diagram of the simultaneous use process of two parallelization strategies involving four non-zero probability basic states provided by the embodiments of the present invention;

[0049] Figure 6 It is a schematic diagram of the fusion of MVM, TP, and MAC operations provided by the embodiments of the present invention;

[0050] Figure 7 It is a schematic diagram of the overall architecture design of the accelerator SiTA provided by the embodiments of the present invention;

[0051] Figure 8 It is a schematic diagram of the structural details of the cascaded multiplexer provided by the embodiments of the present invention;

[0052] Figure 9 It is a schematic diagram of the dynamic multiplier chain of different tensor product configurations provided by the embodiments of the present invention;

[0053] Figure 10 It is a schematic diagram of the acceleration effect of the accelerator SiTA relative to the baseline provided by the embodiments of the present invention;

[0054] Figure 11 It is a schematic diagram of the results of the accelerated ablation study provided by the embodiments of the present invention;

[0055] Figure 12 It is a schematic diagram of the comparison results of the throughput and efficiency of the accelerator SiTA relative to the baseline provided by the embodiments of the present invention. Detailed implementation manners

[0056] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation manners described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0057] The inventive concept of the present invention is as follows: Aiming at the problems of insufficient computational efficiency and fidelity in the quantum readout error mitigation method in the prior art, embodiments of the present invention provide a quantum error mitigation acceleration method and an accelerator that utilize sparsity in the tensor product. By dividing the general mitigation process into three operations: matrix-vector multiplication (MVM), tensor product (TP), and multiply-accumulate (MAC), two acceleration sparsity opportunities, namely output sparsity and input sparsity, are identified. Furthermore, based on these two sparse patterns, a sparse mitigation data stream with probabilistic-level and state-level parallelism is proposed, and the computational path is optimized using these sparse characteristics. Finally, these operators are fused to achieve overall acceleration, and an accelerator architecture is designed to implement this data stream, which has significant parallelism and data reuse characteristics, effectively skips zero-value calculations, thereby greatly improving computational efficiency and obtaining high performance while maintaining good error mitigation fidelity.

[0058] Figure 1 is a schematic flow chart of the quantum error mitigation acceleration method provided by the embodiments of the present invention. As Figure 1 shown, the embodiment provides a quantum error mitigation acceleration method that utilizes sparsity in the tensor product, including the following steps:

[0059] S1, Multiply the noise probability vector of the quantum readout by the mitigation matrix to establish a quantum error mitigation transformation equation, divide the mitigation matrix into several groups of sub-mitigation matrices, and divide the basis vectors of the noise probability vector into qubits corresponding to the several groups of sub-mitigation matrices.

[0060] When measuring the qubits, the readout noise maps the ideal probability vector (i.e., the ideal probability distribution) to the noise probability vector (i.e., the noise probability distribution) , and the noise matrix is a matrix that requires measurements to characterize it. Conversely, in the embodiment, the matrix-based quantum error mitigation method is used to reduce the noise error, multiply the noise probability vector of the quantum readout by the mitigation matrix to obtain the mitigation probability vector, so that the noise probability vector approaches the ideal probability vector. The formula representation of the quantum error mitigation transformation equation is:

[0061] ,

[0062] where represents the mitigation probability vector, represents the mitigation matrix, , represents the noise probability vector.

[0063] In the embodiment, further through the quantum error mitigation method based on the tensor product, the mitigation matrix in the quantum error mitigation transformation equation Divided according to qubits into groups of sub - mitigation matrices , then the quantum error mitigation transformation equation formula is transformed into:

[0064] ,

[0065] wherein, represents the tensor product.

[0066] Further divide the qubits, that is, divide the basis vectors of the noise probability vector into qubits corresponding to the groups of sub - mitigation matrices, then the quantum error mitigation transformation equation formula is transformed into:

[0067] ,

[0068] wherein, represents the noise probability vector in the state of the basis vector , represents the set of non - zero probability basis vectors in after qubit measurement.

[0069] Taking 4 qubits as an example, is expressed as:

[0070] ,

[0071] Next, divide the 4 qubits into three groups , and the mitigation process is shown as follows:

[0072] ,

[0073] wherein, refers to the first two qubits, while and refer to the 3rd and 4th qubits respectively. According to the above process, in the embodiment, further operations are performed based on the iterative tensor product.

[0074] S2. Divide the quantum error mitigation transformation equation into three - step operations of matrix - vector multiplication, tensor product, and multiply - accumulate. Matrix - vector multiplication performs the multiplication calculation between the mitigation matrix and the basis vector and flattens to obtain an intermediate value. Tensor product performs the tensor product calculation of the intermediate value to obtain an intermediate vector and flattens it into an intermediate matrix. Multiply - accumulate performs the multiplication and accumulation of the intermediate matrix and the noise probability vector to obtain the mitigation probability vector as the output.

[0075] As Figure 2As shown, still taking 4 qubits as an example, in the embodiment, the input basis vector is first The mitigation operation is decomposed into three tensor operations: matrix-vector multiplication (MVM), tensor product (TP), and multiply-accumulate (MAC). The matrix-vector multiplication (MVM) operation refers to the multiplication between the mitigation matrix and the basis vector. For example, in the above 4-qubit mitigation process, , , where represents all elements. Considering the case of the static grouping scheme, the overall computational complexity is determined by the non-zero probability . The tensor product (TP) operation performs a tensor product calculation on the vector after the MVM step, generating an intermediate vector that can be flattened into an intermediate matrix (a combination of non-zero probability ground states (basis vectors) after measurement). Each non-zero probability state is responsible for a series of tensor products, such as Figure 2 in . Each non-zero probability requires multiplication operations, which occupy most of the computational complexity . After that, the multiply-accumulate (MAC) operation multiplies the intermediate matrices to output a mitigation probability vector as the mitigated probability result: .

[0076] S3, based on the output sparsity of the mitigation probability vector, pre-obtain the non-zero elements therein to capture the elements that need to be multiplied and accumulated and trace back to the tensor product operation and the multiply-accumulate operation, and convert the matrix-vector multiplication into directly extracting the columns in the mitigation matrix based on the input sparsity of the basis vector to achieve sparsity acceleration.

[0077] After the computational scheme based on the iterative tensor product method, further utilize sparsity to accelerate the computational process. The sparsity scheme is divided into output sparsity and input sparsity.

[0078] (1) Output sparsity. The output mitigation probability vector is inherently sparse. The superposition states generated by quantum algorithms only occupy a part of the Hilbert space. For example, some simple quantum algorithms only generate a few non-zero probability ground states, showing high sparsity: the 16-qubit BV algorithm (Bernstein-Vazirani algorithm) is 98.6%, and the 16-qubit GHZ algorithm (Greenberger-Horne-Zeilinger algorithm) is 96.7%, as Figure 3As shown in (b), "iter" in the abscissa represents the number of iterations, and "q" represents the number of qubits. In addition, as the number of iterations increases, the variational quantum algorithm usually converges to a specific quantum state, resulting in a highly sparse output probability distribution. In Figure 3 In (b), when the number of iterations reaches 40, the sparsities of the 10-qubit VQE (Variational Quantum Eigensolver) algorithm, the 10-qubit QAOA (Quantum Approximate Optimization Algorithm), and the 10-qubit FALQON (Feedback-based Adaptive Learning Quantum Optimizer) algorithm are 65.5%, 80.4%, and 75.3% respectively. Since the ideal probability distribution is usually within the range of the noise probability distribution, in the embodiments, the output sparsity is identified by only calculating the non-zero probability elements in the noise probability vector. Specifically, by pre-obtaining the non-zero elements in the mitigation probability vector, backtracking from each non-zero element to the multiply-accumulate operation to determine the elements that need to perform multiplication and accumulation, and further backtracking to the tensor product operation step to determine the elements that need to perform the tensor product operation. Pruning the mitigation vector can lead to higher accuracy because the entire probability distribution is more concentrated in the range of the ideal state. Through the pruning mitigation vector test, the fidelity of the 12-qubit BV algorithm is increased from 0.59 to 0.63. From an arithmetic perspective, identifying the output sparsity can reduce the complexity of tensor product (TP) operations and multiply-accumulate (MAC) operations as shown in Figure 2 .

[0079] (2) Input sparsity. The basis vectors (ground states) in a quantum system are usually represented as unit vectors, where each vector contains only one element of 1 and the remaining elements are equal to 0. Therefore, input sparsity is essentially bit-level sparsity, for example Figure 2 in . Utilizing this sparsity, the matrix-vector multiplication (MVM) operation is transformed into selecting a column from the mitigation matrix, which is determined by the non-zero indices in the basis vectors. As shown in (a) of Figure 3 , the result of utilizing this sparsity significantly reduces the computational load of the MVM operation by 85.7%. In (a) of Figure 3 , the total redundant calculation involving input and output sparsity is analyzed, and the result is 91.2%, which highlights the great potential for optimization, indicating that further improvement can significantly improve performance and reduce unnecessary computational overhead.

[0080] S4. Perform row-parallel and column-parallel computations on the intermediate vectors in the tensor product operation to generate the rows and columns of the intermediate matrix respectively, achieving parallel acceleration based on sparsity.

[0081] In the embodiment, parallelism is further identified in the sparse data stream with error mitigation. Orthogonal probabilistic-level parallelism and state-level parallelism are identified in the tensor product (TP) operation. Probabilistic-level parallelism is to perform parallel computations on multiple intermediate vectors. State-level parallelism is to perform parallel computations on multiple elements in each intermediate vector. By flattening the intermediate vector obtained from the TP operation into an intermediate matrix , it is equivalent to partitioning the output matrix and performing parallel computations on the elements. As Figure 4 shown in (a) of , probabilistic-level parallelism computes two elements and in the intermediate matrix in parallel. According to the tensor product principle, the computations of these two elements are independent of each other, and six different elements need to be obtained from the input vectors. , . This method outputs the rows of the intermediate matrix, allowing the MAC operation with the noisy probability vector to separately generate the mitigated probability values. As Figure 4 shown in (b) of , state-level parallelism computes the elements of the same intermediate vector, that is, different quantum states of the same intermediate vector, , and . This parallelization is responsible for generating the columns of the intermediate matrix and can reuse the input elements, taking advantage of the inherent data reuse characteristics in the tensor product operation. For example, and are reused to compute and . Therefore, interconnection resources between PEs (processing elements) and on-chip memory are required so that the input elements can be broadcast. Considering that these two parallelization strategies are used together, a block in the intermediate matrix will be generated in one cycle, using the horizontal parallelism

[0082] and the vertical parallelism Figure 5 to represent the parallelism in the two dimensions of probabilistic-level and state-level respectively.

[0083] First, utilize the output sparsity. The specific steps are as follows:

[0084] (1) Trace back the non-zero values in the intermediate matrix to control the redundant computations caused by the sparsity of the mitigated probability vector. Take Figure 5Taking the example of , compress the 0th row (0000), 2nd row (0010), 13th row (1101), and 14th row (1110) of the intermediate matrix

[0085] (2)Identify the operands of the tensor product according to the row index and the qubit grouping scheme. For example, the row index of is . According to the grouping scheme, the indices of the operands in the intermediate value (matrix , , ) can be traced.

[0086] (3)Compress the intermediate matrix into a dense matrix of row indices, which can still implement parallel computing of the tensor product and calculate non-zero outputs.

[0087] (4)Perform MAC operations between the columns of the matrix and the non-zero probability ground states of the vector . .

[0088] When tracing non-zero elements, the block-based parallelization strategy allows the reuse of indices and operands, alleviating the problem of frequent updates of output probabilities. For the values in the same row of a block, the row indices required for their operands are the same. For example, in Figure 5 , and both share the index . Therefore, each block can reuse a set of decoded indices times. Similarly, for the values in the same column of a block, the required operands come from the same index. For example, the tensor products of and both use and . Therefore, each block can reuse the results of matrix-vector multiplication times. Block-based parallel computing can effectively utilize output sparsity while achieving index reuse and operand reuse.

[0089] Next, utilize the input sparsity, and the specific steps are as follows:

[0090] (1)Obtain the corresponding sub-ground states from the ground states according to the grouping scheme. For example, in Figure 5 , according to the grouping scheme, is divided into , , .

[0091] (2) According to its normalization property, the number represented by the ground state points to the position where the value is 1 in the basis vector. For example, in , "11" represents the number 3, indicating that the third element of is 1,

[0092] (3) Based on the position of "1", the columns of the mitigation matrix can be directly extracted from the ground state to complete the matrix-vector multiplication (MVM) operation, expressed as:

[0093] ,

[0094] By directly extracting the columns of the mitigation matrix using bit-level sparsity, the matrix-vector multiplication (MVM) operation can be skipped, thus saving computational resources and time.

[0095] S5. For the tensor product operation, perform the calculation of generating the intermediate matrix in row-major order and column-major order alternately in the iteration rounds to reduce the waiting time of the multiply-accumulate operation between adjacent iteration rounds, and achieve the operation fusion acceleration based on sparsity and parallelism.

[0096] Based on the sparsity acceleration and parallelism acceleration, further implement the operation fusion scheme to achieve acceleration. During the iterative error mitigation process, the MAC operation can be pipelined with the MVM and TP operations because it is element-wise calculated. However, this requires multiple rounds of iteration to improve the fidelity, involving complex data dependencies between different iterations. The output of the MAC operator in the previous iteration is used as the input for the next iteration. Since the generation sequences of the outputs are different, the two iterations cannot be directly overlapped, as shown in (a) of Figure 6 . In the embodiment, this problem is solved by swapping the blocks in the intermediate matrix , which helps to reduce the waiting time between iterations. Specifically, during the first iteration ( Iter.1 ), it follows the row-major order to generate the intermediate matrix row by row, i.e., the Z-order. While during the second iteration ( Iter.2 ), it switches to the column-major order, i.e., the N-order, as shown in 6. By doing so, the three operators are successfully fused and data independence is ensured. The blocks of matrices and are calculated simultaneously, and then the MAC stage is performed independently without waiting for the completion of all MAC operators in the first iteration. This reduces the waiting time of the MAC operator between the two iterations, as shown in Figure 6As shown in (b), for an 8×8 intermediate matrix size, the waiting time is reduced from 64 cycles to 8 cycles, and this improvement significantly enhances the efficiency of iterative calculations.

[0097] Based on the same inventive concept, an embodiment of the present invention also provides a quantum error mitigation accelerator (named SiTA) that utilizes sparsity in the tensor product to achieve fast and accurate quantum readout error mitigation. Figure 7 Figure shows the overall accelerator architecture design of the accelerator SiTA, which consists of a memory and three computing modules. These modules organize the computing process in a fully pipelined manner, including: a matrix-vector multiplication computing unit (MVM unit), a tensor product computing unit (TP unit), a multiply-accumulate computing unit (MAC unit), and a memory (including the SRAM for the input matrix and the SRAM for the output vector). The matrix-vector multiplication computing unit is used to perform a column selection operator on the mitigation matrix according to the strings in the basis vectors, obtain the columns of the mitigation matrix, and flatten them into intermediate values. The tensor product computing unit is used to process the tensor product data stream, in a way of probability-level parallelism and state-level parallelism, calculate the intermediate vector and flatten it into an intermediate matrix. The multiply-accumulate computing unit is used to receive the results from the tensor product computing unit in a block-by-block manner. Each block of the intermediate matrix is generated in row-major order or column-major order and multiplied by the noise probability vector, and the results of different blocks are accumulated and aggregated to obtain the mitigation probability vector as the output. The memory is used to store the input mitigation matrix and the output mitigation probability vector respectively.

[0098] The hardware module can accept configurable parameters and support error mitigation with different granularities. For example, different grouping schemes and iteration numbers. Memory hierarchy: The inputs include 1) the mitigation matrix, 2) the non-zero noise probability vector, and 3) the basis vectors of non-zero probabilities. After dividing the qubits into multiple groups, the mitigation matrix becomes relatively small and can be pre-stored in the on-chip memory. On the other hand, the non-zero noise probability vector is dynamic and varies according to the circuit and the qubits being measured. Therefore, in the embodiment, it is selected to stream the non-zero noise probability vector and its basis vectors through two independent channels, as Figure 7 shown. The on-chip memory system includes two main SRAMs (static random access memories): one for the mitigation matrix (64KB), and the other for the mitigation probability vector (16KB). To facilitate random vector access, the mitigation matrix can be stored in different repositories, up to 4. The 16KB output memory accommodates the mitigation probability vector, which allows up to 8K non-zero FP16 (16-bit floating-point) probability data, sufficient to adapt to the results of current quantum algorithms and quantum devices.

[0099] As Figure 7It also gives an example of the memory behavior when processing the 128-qubit BV algorithm. Explanations of several hardware basic terms involved are as follows: The controller interface represents the communication bridge with external controllers (such as CPUs, DMAs, etc.), responsible for protocol conversion and command parsing; the access control logic represents the control logic execution for data read and write operations; broadcasting means copying data into multiple copies and sending them in parallel to multiple PEs for reception; the sub-index buffer represents the memory area for buffering sub-index data; accumulation means adding up multiple intermediate data; the input buffer represents the memory area for caching input data. In Figure 7 it, the noise probability vector contains 2,896 non-zero values. A grouping scheme that divides 128 bits into 64 groups is adopted, thus obtaining 64 4×4 matrices for iterative tensor product. The mitigation matrix requires 16 KB (4×2×2 KB) of on-chip memory, which contains two sets of 64×4×4 mitigation matrices. In one cycle, four basis vectors enter the 512b index buffer (128b×4) for the decoding stage. The decoding result is used as a read signal to obtain column vectors from the SRAM matrix. After the tensor product calculation is completed through the dynamic multiplier chain, 512b (2×16×16b) of memory is required to store the chunks of two intermediate matrices. The final result is transmitted off-chip through a 6 KB output SRAM (2,896×16b).

[0100] In each computing unit of the accelerator SiTA, the MVM computing unit consists of two register arrays and a cascaded multiplexer, which is used to utilize bit-level sparsity, read the binary strings of a series of basis states, and tell which column needs to be sent to the mitigation matrix to implement the column selection operator. The TP computing unit is a key computing engine involving multiple processing elements (PEs). It receives columns from the mitigation matrix and selects elements according to the qubit grouping scheme (such as , , ). Then, this unit parallelizes the sparse tensor product data stream in a probability-level parallel and state-level parallel manner. The MAC computing unit receives the results from the TP computing unit in a block-by-block manner. Each chunk is multiplied by the noise probability vector in the MAC-Z or MAC-N order. Finally, the results of different blocks are aggregated to output the mitigated vector.

[0101] Since the measured qubits and non-zero probabilities are dynamically determined by the target circuit, a cascaded multiplexer is designed and developed in the MVM computing unit in the embodiment to achieve loading any binary string from the ground state. Obviously, the entire quantum state space is , but only a very small number have non-zero probabilities. Therefore, the multiplexer is used to identify their indices, so that the non-zero values in the input sparsity can be accurately located. As Figure 8As shown, a set of multiplexers are respectively used to select binary strings in the basis vectors and work in a cascaded structure, which realizes a trade-off between latency (which can be overlapped with computation) and area. Figure 8 An example of a cascaded multiplexer extracting the 68th bit in the ground state of 128 qubits is shown. This cascaded multiplexer consists of four levels, including three 4-to-1 multiplexers (4-to-1 MUX) and one 2-to-1 multiplexer (2-to-1 MUX), to select appropriate bit strings at each level. In addition, four dividers are used to generate the selection signals for the multiplexers. For the 68th bit, the 128-bit ground state is divided into four 32-bit binary strings, stored in four registers ( Figure 8 denoted as 32-b in the figure), and these registers serve as the inputs to the 4-to-1 multiplexers. The selection signal is obtained by calculating ⌊68 / 32⌋, which points to the binary string from the 64th bit to the 95th bit, and the division sign represents the divider. The remaining 4 will enter the next stage and operate in a similar manner to select the binary string between the 64th qubit and the 71st qubit. Based on this process, the 68th bit (q 68 ) of the binary string can be obtained, providing the index for column extraction. In addition, such a cascaded multiplexer can also be adapted to quantum systems with different numbers of qubits through a cascaded selection method.

[0102] In addition, a dynamic multiplier chain is further designed in the processing unit. The number of qubits and the grouping scheme of different quantum hardware will vary according to the readout frequency of each qubit, resulting in a dynamic change in the number of multiplications in the tensor product operator. To adapt to different quantum devices, a flexible processing element (PE) is designed in the embodiment, which consists of four multipliers (for calculation) and five demultiplexers. Their configuration signals are generated by software-level control, based on the number of qubits and the corresponding grouping strategy, to ensure the generation of appropriate control signals. Specifically, demultiplexer C0 is used to select the data source between the intermediate results of other PEs and the operands of the tensor product. Similarly, demultiplexers C1, C2, and C3 select the operands of the tensor product to achieve continuous multiplication. Demultiplexer C4 decides whether to output the calculated data or transfer the data to the next PE through a register, as Figure 9 shown. Figure 9 Figure (a) in the figure shows the PE configuration details for dividing a 4-qubit system into 3 groups , where there are 2 multiplication operations, and the tensor product calculation only requires one PE, and no data transfer is required between PEs. PE0 and PE1 respectively complete and of the tensor product calculations. For PE0, demultiplexers C0 and C1 select the tensor product and as the input, while C2 and C3 select 1 as the input without changing the calculation result. The demultiplexer C4 selects to directly output the calculation result to complete the tensor product calculation. Similarly, PE1 selects , and as the operands of the tensor product and directly outputs the calculation result . Figure 9 (b) in shows the details of the PE configuration that divides the 32 - qubit system into 16 groups . Each group contains two qubits, and the requirement to perform 16 multiplication operations exceeds the capacity of the multiplier of a single processing element (PE). Therefore, multiple PEs need to cooperate to perform the tensor product operation, and data transmission occurs between PEs. For PE0 and PE1, the demultiplexers C1, C2, and C3 are all configured to select the input operands of the tensor product operation. The demultiplexer C4 of PE0 transfers the intermediate result to the register and sends it as the input for subsequent calculations to PE1. Finally, after the calculations and transmissions of PE1 and PE2, the tensor product result is calculated on PE3.

[0103] A quantum error mitigation accelerator that utilizes sparsity in the tensor product according to the above - mentioned embodiments, its specific hardware implementation method is: The SiTA accelerator is implemented using Xilinx High - Level Synthesis (HLS) to generate a hardware description language (HDLs), such as Verilog, from high - level C++ code. Its performance is evaluated on the Xilinx Alveo U50 platform, operating at a clock frequency of 300 MHz. The SiTA accelerator is also implemented in RTL and synthesized using Synopsys Design Compiler to estimate the chip area and total power consumption under the FreePDK 45nm standard cell library. The SiTA is synthesized based on 45 - nanometer CMOS technology in the embodiments, and the parameters are shown in Table 1.

[0104] Table 1 Area and Power of the SiTA Architecture

[0105]

[0106] As shown in detail in Table 1, SiTA runs seamlessly at a clock frequency of 1 GHz, with a total power consumption of 0.85 W and a compact chip area of 4.36 mm². Further analysis shows that the dynamic multiplier chain becomes the main contributor to power consumption, accounting for 52.9% of the total power consumption. On the other hand, the on - chip buffer responsible for storing the mitigation matrix and mitigation probability vectors plays a crucial role in chip area utilization, occupying 56.9% of the entire chip area.

[0107] The SiTA provided by the embodiments of the present invention is compared with the existing readout error mitigation methods IBU (Realizing topologically ordered states on a quantum processor), Mthree (Scalable mitigation of measurement errors on quantum computers), QuFEM (Fast and accurate quantum readout calibration using the finite element method), and SpREM (Exploiting hamming sparsity for fast quantum readout error mitigation) in combination with specific test data to evaluate their effectiveness. QuFEM runs on an AMD EPYC 9554 CPU (using numpy v1.24.3), while IBU and Mthree run on an NVIDIA A100 PCIe 80 GB GPU (using jax-gpu v0.4.13 and cuBLAS v12.4). SpREM is a hardware-level technology implemented on the Xilinx Alveo U50 platform and operates at a clock frequency of 300 MHz. For all iterative methods, the number of iterations is set to 2. The pruning threshold of QuFEM is set to 10−5. The baseline parameters are fine-tuned to ensure low latency while maintaining high precision.

[0108] (1) Mitigation matrix and benchmark settings: QuFEM is used to characterize the mitigation matrix by running 1380 matrix characterization circuits on the Quafu and Rigetti quantum platforms. The detailed noise characteristics of the two quantum devices are provided in Table 2, where CZ represents the controlled-phase gate error, XY represents the XY gate error that simulates the XY interaction in quantum computing. The benchmarks in the experiment include five well-known quantum algorithms: Bernstein-Vazirani (BV algorithm), Greenberger-Horne-Zeilinger (GHZ algorithm), and three variational quantum algorithms: the feedback-based quantum optimization algorithm (FALQON algorithm), the variational quantum solver (VQE algorithm), and the quantum approximate optimization algorithm (QAOA algorithm). The number of shots for running these quantum circuits is set to 100,000 to ensure a stable probability vector.

[0109] Table 2 Quantum Device Specifications

[0110]

[0111] The Hellinger fidelity is used as an evaluation metric to measure the effectiveness of mitigation. The Hellinger fidelity, as a measure of the similarity between two probability distributions, is used to evaluate the closeness between the noise distribution and the ideal (i.e., noiseless) probability distribution. The latency of SiTA and SpREM is measured using the Xilinx runtime, while their power consumption is analyzed using xbutil. The GPU execution time is measured by cudaEventElapsedTime, and the power consumption is obtained by nvidia-smi.

[0112] (2) Speedup Beneficial Effects: Figure 10 The speedup effect of SiTA relative to the baseline is shown. SiTA achieves geometric mean speedup ratios of 3.9x, 5.5x, 826.1x, and 1340.8x in comparisons with SpREM, Mthree, IBU, and QuFEM, respectively. These speedup ratios are achieved by exploiting sparsity to eliminate redundant computations. Specifically, SiTA has probabilistic-level and state-level parallelism to accelerate the tensor product operator and uses sparsity to eliminate computations involving zero values. For the 96-qubit GHZ algorithm, SiTA achieves a speedup of up to 1217.7x relative to IBU on an NVIDIA A100 GPU. This is because IBU does not optimize the latency between iterations, resulting in low GPU utilization. SiTA shows a speedup of 3.8x to 4.7x compared to SpREM in the BV and GHZ algorithms. In practical applications, although SpREM uses the Hamming sparse format to compress the mitigation matrix, it lacks the technology to characterize this matrix, which prevents it from directly mitigating large-scale quantum circuits with more than 19 qubits. When considering the characterization time, SiTA achieves an end-to-end speedup of more than 3.8x to 4.7x compared to SpREM.

[0113] (3) Speedup Analysis: Figure 11 Figure (a) in [reference] shows the comparison of the operation counts between SiTA and QuFEM. By exploiting the sparsity of the output and input, the number of operations of SiTA is reduced by 1.6x and 1.1x on average compared to QuFEM, respectively. Output sparsity corresponds to the key operation: the tensor product. For example, in the 12-qubit BV algorithm, this operation accounts for 93.7% of the total redundant computations. Therefore, output sparsity significantly reduces the number of operations when eliminating redundant computations. Figure 3In (b), the speedup decomposition of SiTA relative to QuFEM is shown. The overall speedup of 1340.8x mainly comes from three technologies: sparsity acceleration, parallelism acceleration, and operation fusion acceleration provided by the embodiments of the present invention. First, the professional design of the tensor product calculation unit contributes a speedup of 415.1x by combining probabilistic-level and state-level parallelism. Second, the sparse mitigation data stream skips calculations related to zero values, achieving a speedup of 1.7x. Third, the operation fusion technology reduces the latency between iterations, further enhancing the speedup.

[0114] (4) Throughput beneficial effects: As Figure 12 shown, SiTA achieves a geometric mean throughput of 386.2 GOPS, achieving throughput improvements of 3.2x, 4.3x, and 351.1x compared to SpREM, Mthree, and IBU, respectively. SiTA combines probabilistic-level and state-level parallelism to accelerate the key tensor product operations in the readout error mitigation process. By integrating the Z-order and N-order, SiTA fuses the operators of two iterations, promoting overlapping calculations and significantly improving the overall execution efficiency. Therefore, this method achieves a maximum throughput of 420.1 GOPS. It should be noted that SiTA achieves better performance in the GHZ algorithm. This is because the GHZ algorithm outputs more ground states, which increases the load on the SiTA accelerator and enables it to perform computational mitigation more effectively.

[0115] (5) Energy efficiency beneficial effects: In terms of energy efficiency, SiTA achieves an energy efficiency of 14.6 GOPS / W, which is 1.3x, 3.2x, and 2433.3x higher than SpREM, Mthree, and IBU, respectively, as Figure 12 shown. In addition, SiTA demonstrates an impressive improvement in energy efficiency, achieving a three-order-of-magnitude increase (from 1985.9x to 4187.3x) compared to IBU on the NVIDIA A100 GPU. This significant enhancement is largely attributed to the design of SiTA, which stores the mitigation matrix on-chip, effectively minimizing off-chip memory access. By reducing the need for external memory interactions, SiTA not only improves energy efficiency but also optimizes the overall performance.

[0116] (6) Fidelity beneficial effects: For clarity, a simplified representation of a variational quantum algorithm is provided for easy comparison and discussion. For example, QAOA-10-2L represents a 10-qubit QAOA algorithm with two layers of quantum gates L, as shown in Table 3. The comparison results of the Hellinger fidelity between SiTA and the baseline are presented in Table 3.

[0117] Table 3 Comparison of Hellinger Fidelity between SiTA and Baseline Methods

[0118]

[0119] For the BV and GHZ algorithms, the results of SiTA show the highest output fidelity, with an average increase of 2.3 times, 1.5 times, 1.4 times, 1.3 times, and 1.1 times higher than the original results, respectively, compared to the results of Mthree, SpREM, IBU, and QuFEM. For these three variational quantum algorithms, their quantum circuits contain more two-qubit gates, resulting in increased crosstalk noise during the readout process. However, neither IBU nor Mthree considered crosstalk, while SpREM only simulated local crosstalk. Therefore, the error mitigation effects of Mthree, IBU, and SpREM are relatively low, and they are even ineffective for the QAOA, VQE, and FALQON algorithms. For example, Mthree has no mitigation effect on the VQE-12-2L algorithm on the Rigetti quantum platform. In contrast, SiTA mitigates the crosstalk noise between different qubits through an iterative tensor product formula, achieving an average fidelity improvement of 25.0%. In addition, it is also observed that SiTA has a higher fidelity than QuFEM for the BV, GHZ, and QAOA algorithms, while for the VQE and FALQON algorithms, the fidelities of SiTA and QuFEM are similar. This is attributed to the simple circuit structures of the BV, GHZ, and QAOA algorithms, which have less noise and thus output sparse noise probability vectors. However, in the VQE and FALQON algorithms, their circuit noise levels are relatively high, which makes the measured basis states evenly distributed in the Hilbert space.

[0120] The above-described specific embodiments have elaborated on the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A quantum error mitigation acceleration method using sparsity in the tensor product, characterized in that, Including the following steps: Multiply the noise probability vector of quantum readout by the mitigation matrix to establish a quantum error mitigation transformation equation. Divide the mitigation matrix into several groups of sub-mitigation matrices and divide the basis vectors of the noise probability vector into qubits corresponding to the several groups of sub-mitigation matrices; Divide the quantum error mitigation transformation equation into three-step operations of matrix-vector multiplication, tensor product, and multiply-accumulate. The matrix-vector multiplication performs the multiplication calculation between the mitigation matrix and the basis vector and flattens it to obtain an intermediate value. The tensor product performs the tensor product calculation of the intermediate value to obtain an intermediate vector and flattens it into an intermediate matrix. The multiply-accumulate performs the multiplication and accumulation of the intermediate matrix and the noise probability vector to obtain the mitigation probability vector as the output; Based on the output sparsity of the mitigation probability vector, pre-obtain the non-zero elements therein to capture the elements that need to perform multiplication and accumulation and trace back to the tensor product operation and the multiply-accumulate operation. And based on the input sparsity of the basis vector, convert the matrix-vector multiplication into directly extracting the columns in the mitigation matrix to achieve sparsity acceleration; Perform row-parallel calculation and column-parallel calculation on the intermediate vector in the tensor product operation simultaneously to generate the rows and columns of the intermediate matrix respectively to achieve parallelism acceleration based on sparsity; Alternately perform row-major order and column-major order in the iteration rounds of the tensor product operation to generate the calculation of the intermediate matrix so as to reduce the waiting time of the multiply-accumulate operation between adjacent iteration rounds to achieve operation fusion acceleration based on sparsity and parallelism.

2. The quantum error mitigation acceleration method using sparsity in the tensor product according to claim 1, wherein The multiplying the noise probability vector of quantum readout by the mitigation matrix to establish a quantum error mitigation transformation equation includes: Through the matrix-based quantum error mitigation method, multiply the noise probability vector of quantum readout by the mitigation matrix to obtain the mitigation probability vector, so that the noise probability vector approaches the ideal probability vector. The formula of the quantum error mitigation transformation equation is expressed as: , Among them, represents the remission probability vector, represents the remission matrix, represents the noise probability vector.

3. The quantum error mitigation acceleration method using sparsity in the tensor product according to claim 1 or 2, characterized in that The dividing the mitigation matrix into several groups of sub-mitigation matrices and dividing the basis vectors of the noise probability vector into qubits corresponding to the several groups of sub-mitigation matrices includes: Through a quantum error mitigation method based on the tensor product, the mitigation matrix in the quantum error mitigation transformation equation is partitioned into groups of sub-mitigation matrices , and the formula of the quantum error mitigation transformation equation is transformed into: , Among them, represents the mitigation probability vector, represents the noise probability vector, represents the tensor product; The noise probability vector of the basis vectors is further segmented into qubits corresponding to a group of sub mitigation matrices , and then the quantum error mitigation transformation equation formula is transformed into: , Among them, represents the noise probability vector in the state of the basis vector, represents the set of non-zero probability basis vectors in after qubit measurement.

4. The quantum error mitigation acceleration method using sparsity in the tensor product according to claim 1, characterized in that The pre-obtaining the non-zero elements therein based on the output sparsity of the mitigation probability vector to capture the elements that need to perform multiplication and accumulation and trace back to the tensor product operation and the multiply-accumulate operation includes: Based on the output sparsity of the mitigation probability vector, pre-obtain the non-zero elements in the mitigation probability vector, trace back to the multiply-accumulate operation according to each non-zero element to determine the elements that need to perform multiplication and accumulation, and further trace back to the tensor product operation step to determine the elements that need to perform the tensor product operation.

5. The quantum error mitigation acceleration method using sparsity in the tensor product according to claim 1, characterized in that The converting the matrix-vector multiplication into directly extracting the columns in the mitigation matrix based on the input sparsity of the basis vector includes: The basis vector in the quantum system is represented as a unit vector. By identifying the non-zero indices in the basis vector, convert the matrix-vector multiplication operation into directly selecting the corresponding columns from the mitigation matrix according to the non-zero indices.

6. The quantum error mitigation acceleration method using sparsity in the tensor product according to claim 1, characterized in that The performing row-parallel calculation and column-parallel calculation on the intermediate vector in the tensor product operation simultaneously to generate the rows and columns of the intermediate matrix respectively includes: Identify orthogonal probabilistic parallelism and state-level parallelism in the tensor product operation. Probabilistic parallelism means parallelly computing multiple intermediate vectors, and state-level parallelism means parallelly computing multiple elements in each intermediate vector. Perform row-wise parallel computing of each row element of the intermediate vector in the tensor product operation according to probabilistic parallelism, and at the same time perform column-wise parallel computing of each column element of the intermediate vector in the tensor product operation according to state-level parallelism, so as to generate a chunk in the intermediate matrix in each computing cycle, and finally output the entire intermediate matrix.

7. The quantum error mitigation acceleration method using sparsity in the tensor product according to claim 6, wherein For the values in the same row of each chunk, their operands and row indices are the same. Accordingly, operand reuse and index reuse are performed when calculating the values in the same row of each chunk.

8. The quantum error mitigation acceleration method using sparsity in the tensor product according to claim 1, wherein The calculation of the tensor product operation alternately performs row-major order and column-major order in the iteration rounds to reduce the waiting time of the multiply-accumulate operation between adjacent iteration rounds, including: In the current iteration round, the tensor product operation follows the row-major order to generate the intermediate matrix row by row, and in the next iteration round, the tensor product operation follows the column-major order to generate the intermediate matrix column by column. Then, the multiply-accumulate operation can be independently performed after the tensor product operation in the next iteration round without waiting for all the multiply-accumulate operations in the current iteration round to be completed.

9. A quantum error mitigation accelerator that utilizes sparsity in the tensor product, implemented based on the quantum error mitigation acceleration method using sparsity in the tensor product according to any one of claims 1-8, characterized in that Including: A matrix-vector multiplication calculation unit, a tensor product calculation unit, a multiply-accumulate calculation unit, and a memory; The matrix-vector multiplication calculation unit is used to perform a column selection operator for the mitigation matrix according to the string in the basis vector, obtain the columns of the mitigation matrix, and flatten them into intermediate values; The tensor product calculation unit is used to process the tensor product data stream, in a way of probabilistic parallelism and state-level parallelism, calculate to obtain intermediate vectors and flatten them into an intermediate matrix; The multiply-accumulate calculation unit is used to receive the results from the tensor product calculation unit in a block-by-block manner. Each chunk of the intermediate matrix is generated in row-major order or column-major order and multiplied by the noise probability vector. The results of different chunks are accumulated and aggregated to obtain the mitigation probability vector as the output; The memory is used to store the input mitigation matrix and the output mitigation probability vector respectively.

10. The quantum error mitigation accelerator using sparsity in the tensor product according to claim 9, characterized in that, Design a cascaded multiplexer in the matrix-vector multiplication calculation unit to respectively select the strings in the basis vector and work in a cascaded structure; design a processing unit including a multiplier and a demultiplexer in the tensor product calculation unit. The multiplier is used to perform the tensor product calculation, and the demultiplexer is used to select the input operands or intermediate results of the tensor product operation and configure the output path to select whether to directly output the calculation result or transfer it to the next processing unit through a register.

Citation Information

Patent Citations

  • Methods and systems for multi-type probabilistic quantum error mitigation

    EP4485294A1

  • Quantum readout error mitigation by stochastic matrix inversion

    US20210256410A1