Pipeline Hardware Decoder for Quantum Computing Devices

A hardware-based microarchitecture for quantum error correction addresses the inefficiencies in existing methods by reducing hardware complexity and achieving near-linear time complexity for error correction in quantum computers, enabling scalable fault-tolerant quantum computing.

CN114207632BActive Publication Date: 2025-07-15MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080055614.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-15
Filing Date
2020-06-09
Publication Date
2025-07-15
Estimated Expiration
2040-06-09

AI Technical Summary

Technical Problem

Quantum bits are prone to high error rates in quantum computing, and the existing decoding methods fail to effectively consider underlying equipment technology, resulting in low error correction efficiency and unscalable.

Method used

Design a hardware to realize the microarchitecture of the decoder. Through three-stage pipelined microarchitecture and compression technology, the hardware complexity is reduced, and the error correction of a large number of logical qubits is supported, which is suitable for low-temperature environment work.

Benefits of technology

It realizes nearly linear time complexity and high-throughput error correction capabilities, supports scaling to thousands of logical qubits, reduces hardware cost and power consumption, and improves the reliability of quantum computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114207632B_ABST
    Figure CN114207632B_ABST
Patent Text Reader

Abstract

A quantum computing device includes: at least one quantum register including a plurality of qubits; and a hardware decoder. The hardware decoder is configured to: receive syndrome data from one or more of the plurality of qubits; and decode the received syndrome data by implementing a union-find decoding algorithm via a hardware microarchitecture including two or more pipeline stages.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] Quantum bits are prone to high error rates and benefit from active error correction. Quantum error correction codes can be used to encode logical quantum bits into a collection of physical quantum bits. Then, measurements can be used to detect errors and an error decoder can be used to correct the errors. Quantum bits typically operate at very low temperatures, and data is transferred to the error decoder at higher operating temperatures. SUMMARY OF THE INVENTION

[0002] The present Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. The present Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Additionally, the claimed subject matter is not limited to implementations that solve any or all disadvantages noted in any part of this disclosure.

[0003] A quantum computing device includes: at least one quantum register including a plurality of quantum bits; and a hardware decoder. The hardware decoder is configured to: receive syndrome data from one or more of the plurality of quantum bits; and decode the received syndrome data by implementing a parity-check decoding algorithm via a hardware microarchitecture including two or more pipeline stages. This enables high-throughput error correction with a nearly linear time complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Figure 1 An example of a quantum computing organization is schematically illustrated.

[0005] Figure 2 Aspects of an example quantum computer are schematically illustrated.

[0006] Figure 3 A Bloch sphere is illustrated, which graphically represents the quantum state of a single quantum bit in a quantum computer.

[0007] Figure 4 A logical quantum bit in a lattice having alternating data qubits and parity qubits is schematically illustrated.

[0008] Figure 5 Two consecutive rounds of syndrome measurements are schematically illustrated.

[0009] Figure 6 is a graph indicating the memory capacity required to store syndrome measurement data under varying conditions.

[0010] Figure 7 An example compression scheme is schematically illustrated.

[0011] Figure 8Schematically shows multiple regions on a surface code lattice for geometry-based compression.

[0012] Figure 9 Shows an example method for compressing adjoint data within a quantum computing device.

[0013] Figure 10 Shows an example method for compressing adjoint data within a quantum computing device using a geometry-based compressor.

[0014] Figure 11 Schematically shows an example decoder design.

[0015] Figure 12 Schematically shows an example union-find decoder.

[0016] Figure 13 Schematically shows an example graph generator module.

[0017] Figure 14 Schematically shows the states of the main components in the graph generator module during graph generation.

[0018] Figure 15 Schematically shows an example cluster and root table entry.

[0019] Figure 16 Schematically shows a depth-first search engine and an example graph for error correction.

[0020] Figure 17 Shows an example of peeling for an example error graph executed in a correction engine.

[0021] Figure 18 Shows an example method for implementing a pipelined version of a hardware union-find decoder.

[0022] Figure 19 Schematically shows the baseline organization of L logical qubits.

[0023] Figure 20 Shows an example design of a decoder module.

[0024] Figure 21 Schematically shows an example microarchitecture of a scalable fault-tolerant quantum computer.

[0025] Figure 22 Shows an example method for routing adjoint data within a quantum computing device.

[0026] Figure 23 Schematically shows a Monte Carlo emulator framework.

[0027] Figure 24It is a curve graph indicating the average compression ratio of different error rates across code distances.

[0028] Figure 25 It is a curve graph indicating the average cluster diameter of different error rates and code distances of logical qubits.

[0029] Figure 26 It is a curve graph indicating the total memory capacity required to implement a spanning tree memory.

[0030] Figure 27 It is a curve graph indicating the distribution of the number of edges in a cluster of fixed code distance and physical error rate.

[0031] Figure 28 It is a curve graph showing the average number of edges in a cluster of different code distances and error rates.

[0032] Figure 29 It is a curve graph indicating the correlation between the graph generator module and the execution time in a depth - first search engine.

[0033] Figure 30 It is a curve graph indicating the distribution of the execution time for decoding a 3D graph.

[0034] Figure 31 It shows a schematic diagram of an example classical computing device. Detailed implementation

[0035] Qubits (the basic information unit in a quantum computer) are prone to high error rates. To support fault - tolerant quantum computing, active error correction can be applied to these qubits. Quantum error - correcting codes (QECCs) use redundant data and parity qubits to encode logical qubits. Error correction is performed through a process called error decoding, which diagnoses errors on data qubits by analyzing the measurements of parity qubits. Currently, most decoding methods target qubit errors at the algorithm level and do not take into account the underlying device technology used to design them.

[0036] In this article, aiming at the architectural challenges involved in designing these decoders, a 3 - stage pipelined microarchitecture for the hardware implementation of the union - find decoder is described. The error - correction algorithm is designed to be suitable for hardware implementation. Regarding the storage amount and bandwidth required for implementation, the feasibility of data compression for different noise mechanisms is evaluated. An architecture is disclosed that scales the proposed decoder design for a large number of logical qubits and supports practical fault - tolerant quantum computing. Such a design can reduce the total cost of each of the three pipeline stages by 2x, 4x, and 4x respectively through resource sharing across multiple logical qubits without affecting decoder correctness and error thresholds. As an example, for a code distance of 11 and 10 -3The physical error rate is on the order of magnitude, and the logical error rate is 10 -8 .

[0037] Quantum computing uses quantum mechanical properties to support computations for specific applications that are infeasible to perform on conventional (i.e., non - quantum) state - of - the - art computers within a reasonable amount of time. Example applications include prime factorization, database search, and physical and chemical simulations. The basic computational unit on a quantum computer is the qubit. Qubits inevitably interact with the environment and lose their quantum states. Imperfect quantum gate operations exacerbate this problem, since quantum gates are unitary transformations selected from a continuous set of possible values and thus cannot be implemented with perfect accuracy. To protect quantum states from noise, QECC has been developed. In any QECC, logical qubits are encoded with several physical qubits to support fault - tolerant quantum computing. Fault tolerance incurs a resource overhead (typically 20 - 100x per logical qubit) and is effectively infeasible on currently available small prototypes. However, quantum error correction is considered highly valuable, if not essential, for running useful applications on fault - tolerant quantum computers.

[0038] QECC is different from classical error - correction techniques like triple - modular redundancy (TMR) or single - error - correcting double - error - detecting (SECDED). These differences stem from the fundamental properties of qubits and the high error rate (typically on the order of 10 -2 ). For example, qubits cannot be copied (no - cloning theorem) and lose their quantum states when measured. QECC uses redundant qubits to create an encoded space by using ancilla qubits that interact with the data qubits. By measuring the ancilla qubits, errors on the data qubits can be detected and corrected using a decoder.

[0039] The error - decoding algorithm dictates how the syndrome (the result of the ancilla measurements) will be processed to detect errors in the encoded blocks of data qubits. The design and performance of the decoder depend on the decoding algorithm, QECC, physical error rate, noise model, and implementation technology. For practical purposes, the rate at which the decoder processes syndrome measurements must be faster than the rate at which errors occur. They must also account for specific technical constraints for operation in cryogenic environments and scale to a large number of qubits.

[0040] Perfect error decoding is NP-hard (non-deterministic polynomial-time) with exponential time complexity. Therefore, the best decoding algorithms trade off error correction capabilities to reduce the time complexity. Although the decoder is groundbreaking for fault-tolerant quantum computing, most decoding techniques have been studied only at the algorithmic level and do not consider the underlying implementation technology. Other methods such as lookup-table-based decoders or deep neural decoders are not scalable to a large number of qubits. The Union-Find decoder algorithm is simple and has near-linear time complexity, making it a suitable candidate for scalable fault-tolerant quantum computing. Here, a microarchitecture for the hardware implementation of the Union-Find decoder is disclosed, where the algorithm is redesigned to reduce the hardware complexity and allow for scaling to a large number of logical qubits.

[0041] To support faster processing and reduce the communication latency, the decoder is designed to operate at a temperature very close to that of the physical qubits (77K or 4K), rather than at room temperature (300K). Figure 1 An example quantum computing system 100 indicating the temperature gradient 105 is shown. The quantum computing system 100 includes one or more qubit registers 110 operating at 20mK, a decoder 115 and a controller 120 typically operating at 4K or 77K, and a host computing device 125 and an end-user computing device 130 operating at 300K (room temperature). Depending on whether the decoder 115 operates at 77K or 4K, the underlying implementation technology and design constraints provide different trade-offs. Superconducting logic design at 4K is very close to the physical qubits and has significant energy efficiency, but is limited by device density and memory capacity. Conventional CMOS operating at 77K can drive complex designs for applications with a larger memory footprint, but its energy efficiency is lower than that of superconducting logic, and moving data back and forth from the physical qubits residing at 15 - 20mK incurs data transmission overhead.

[0042] In this paper, a microarchitecture for the hardware implementation of the Union-Find decoder is disclosed. The implementation challenges associated with memory capacity and bandwidth for operation in a cryogenic environment are discussed. The surface code is used as an example for the underlying QECC and various noise models, although other implementations have been considered. The surface code is a promising QECC candidate that arranges a set of qubits in a 2D layout with alternating data qubits and ancilla qubits. Any error in the data qubits can be detected by its neighboring ancilla qubits, thus only requiring nearest-neighbor connections. The feasibility and scalability of such a design for large-scale fault-tolerant quantum computers are described.

[0043] Systems and methods for solving many problems in the field of quantum computing are disclosed herein. For example, the design of QEC decoders is analyzed, including their placement and design complexity in the thermal domain. A microarchitecture for the hardware implementation of the union-find decoder is proposed, which proves to be more practical to operate the decoder at 77K.

[0044] The memory capacity required to store adjoint measurements is calculated, and it is shown that storing them in superconducting memory at 4K may be infeasible. However, transferring data to 77K requires a large bandwidth. To overcome these two challenges, techniques that can be used to compress adjoint measurement data are proposed. The implementation of dynamic zero compression and sparse representation is described. Additionally, a geometry-based compression scheme is proposed, which takes into account the underlying structure of the surface code lattice. Additionally, compression schemes and their applicability are described for different noise mechanisms.

[0045] The union-find decoder algorithm is refined to reduce hardware costs and for enhanced implementation in the noise model. The original union-find decoding algorithm only considered gate errors on paired data qubits in space. The hardware microarchitecture described herein also considers measurement errors, which are paired in time, and decodes them using several rounds of adjoint measurements.

[0046] Additionally, the hardware system architecture for scaling these decoders for a large number of logical qubits is described. Such an implementation can consider the utilization differences of pipeline stages in individual decoding units and support optimized resource sharing among multiple logical qubits to reduce hardware costs.

[0047] A qubit is the basic information unit on a quantum computer. The foundation of quantum computing depends on two quantum mechanical properties: superposition and entanglement. A qubit can be represented as a linear combination of its two basis states. If the basis states are |0> and |1>, then the qubit |ψ> can be represented as |ψ> = α|0> + β|1>, where and |α| 2 + |β| 2 = 1. When the magnitudes or / and phases of the probability amplitudes α, β change, the state of the qubit changes. For example, an amplitude flip (or, bit flip) changes the state of |ψ> to β|0> + α|1>. Alternatively, a phase flip changes its state to α|0> - β|1>. Quantum instructions use quantum gate operations to modify the probability amplitudes, and quantum gate operations are represented using the identity matrix (I) and Pauli matrices. The Pauli matrices X, Z, and Y represent the effects of bit flip, phase flip, or both, respectively.

[0048] In some embodiments, the methods and processes described herein can be associated with a quantum computing system of one or more quantum computing devices. Figure 2Aspects of an example quantum computer 210 configured to perform quantum logic operations are shown (see below). Conventional computer memory stores digital data in arrays of bits and performs bit-by-bit logical operations, while a quantum computer stores data in an array of qubits and performs quantum mechanical operations on the qubits to achieve the desired logic. Accordingly, Figure 2 the quantum computer 210 includes at least one register 212 that includes an array of qubits 214. The illustrated register has a length of eight qubits; registers including longer and shorter arrays of qubits are also contemplated, as are quantum computers including two or more registers of any length.

[0049] The qubits of register 212 can take various forms depending on the desired architecture of quantum computer 210. As a non-limiting example, each of qubits 214a through 214h can include: a superconducting Josephson junction, a trapped ion, a trapped atom coupled to a high-finesse cavity, an atom or molecule confined within a fullerene, an ion or neutral dopant atom confined within a host lattice, a quantum dot presenting discrete spatial or spin states, an electron-hole in a semiconductor junction entrained via an electrostatic trap, a pair of coupled quantum wires, a nucleus accessible via magnetic resonance, a free electron in helium, a molecular magnet, or a metallike carbon nanosphere. More generally, each of qubits 214a through 214h can include any particle or system of particles that can exist in two or more discrete quantum states that can be experimentally measured and manipulated. For example, qubits can also be implemented in multiple processing states corresponding to different optical propagation modes through linear optical elements (e.g., mirrors, beam splitters, and phase shifters), as well as in states accumulated within a Bose-Einstein condensate.

[0050] Figure 3 is a diagram of a Bloch sphere 216 that provides a graphical description of some quantum mechanical aspects of each of qubits 214a through 214h. In this description, the north and south poles of the Bloch sphere correspond to the standard basis vectors |0> and |1> respectively—for example, the up-spin and down-spin states of an electron or other fermion. The set of points on the surface of the Bloch sphere includes all possible pure states |ψ> of the qubit, while the interior points correspond to all possible mixed states. The mixed state of a given qubit may be caused by decoherence, which may occur due to an undesired coupling to external degrees of freedom.

[0051] Now returning to Figure 2, the quantum computer 210 includes a controller 218. The controller may include conventional electronic components, including at least one processor 220 and an associated storage machine 222. The term "conventional" as used herein is applied to any component that can be modeled as a collection of particles without regard to the quantum state of any individual particle. For example, conventional electronic components include integrated microfabricated transistors, resistors, and capacitors. The storage machine 222 may be configured to store program instructions 224 that cause the processor 220 to perform any of the processes described herein. Additional aspects of the controller 218 are described below.

[0052] The controller 218 of the quantum computer 210 is configured to receive a plurality of inputs 226 and provide a plurality of outputs 228. The inputs and outputs may each include digital and / or analog lines. At least some of the inputs and outputs may be data lines through which data is provided to and extracted from the quantum computer. Other inputs may include control lines via which the operation of the quantum computer can be adjusted or otherwise controlled.

[0053] The controller 218 is operably coupled to the register 212 via an interface 230. The interface is configured to exchange data bi-directionally with the controller. The interface is also configured to exchange signals corresponding to the data bi-directionally with the register. Depending on the architecture of the quantum computer 210, such signals may include electrical, magnetic, and / or optical signals. Via the signals transmitted through the interface, the controller can interrogate or otherwise affect the quantum states stored in the register, as defined by the collective quantum state of the qubit array 214. To this end, the interface includes at least one modulator 232 and at least one demodulator 234, each modulator and each demodulator being operably coupled to one or more qubits of the register 212. Each modulator is configured to output a signal to the register based on modulation data received from the controller. Each demodulator is configured to sense a signal from the register and output data to the controller based on that signal. In some scenarios, the data received from the demodulator may be an estimate of the observable value of a measurement of the quantum state stored in the register.

[0054] More specifically, a suitably configured signal from modulator 232 can physically interact with one or more of qubits 214a through 214h in register 212 to trigger a measurement of the quantum state(s) stored in one or more qubits. Demodulator 234 can then sense the resulting signal released by one or more qubits based on the measurement and can provide data corresponding to the resulting signal to the controller. In other words, the demodulator can be configured to reveal an estimate of an observable value that reflects the quantum state of one or more qubits of the register based on the received signal and provide that estimate to controller 218. In one non-limiting example, the modulator can provide a suitable voltage pulse or sequence of pulses to the electrodes of one or more qubits based on data from the controller to initiate a measurement. In short, the demodulator can sense photon emissions from one or more qubits and can assert a corresponding digital voltage level on the interface line(s) into the controller. Generally, any measurement of a quantum mechanical state is defined by an operator corresponding to the observable being measured; the result R of the measurement is guaranteed to be one of the allowed eigenvalues. In quantum computer 210, R is statistically related to the register state prior to the measurement but is not uniquely determined by that register state.

[0055] According to appropriate inputs from controller 218, interface 230 can also be configured to implement one or more quantum logic gates to operate on the quantum states stored in register 212. The functionality of each logic gate in a conventional computer system is described according to a corresponding truth table, while the functionality of each quantum gate is described by a corresponding operator matrix. The operator matrix operates (i.e., multiplies) on the complex vector representing the register state and implements a specific rotation of that vector in Hilbert space.

[0056] Continuing Figure 2 , a suitably configured signal from modulator 232 of interface 230 can physically interact with one or more of qubits 214a through 214h in register 212 in order to assert any desired quantum gate operation. As described above, a desired quantum gate operation is a specifically defined rotation of the complex vector representing the register state. To achieve the desired rotation one or more modulators of interface 230 can apply a predetermined signal level S i for a predetermined duration T i .

[0057] In some examples, multiple signal levels may be applied to multiple sequences or other associated durations. In more specific examples, multiple signal levels and durations are arranged to form a composite signal waveform that may be applied to one or more qubits of a register. Generally, each signal level S i and each duration T i are control parameters that are adjustable through appropriate programming of the controller 218. In other quantum computing architectures, different sets of adjustable control parameters may control the quantum operations applied to the register states.

[0058] Qubits inevitably lose their quantum states through interactions with different degrees of freedom in their surroundings. Even if qubits can be perfectly isolated from environmental noise, quantum gate operations are imperfect and cannot be applied with exact precision. This poses various limitations to running any application on a quantum computer. Therefore, the quantum states manipulated by a quantum computer must be error-corrected using quantum error-correcting codes (QECCs). QECCs encode logical qubits into a collection of physical qubits such that the error rate of the logical qubits is lower than the physical error rate. As long as the physical error rate is below an acceptable threshold, QECCs can support fault-tolerant quantum computing at the cost of an increased number of physical qubits. In recent years, several error-correction protocols have been proposed. The surface code is applied in this article and is considered to be the most promising QECC for fault-tolerant quantum computing. QEC models arbitrary noise as a superposition of quantum operations. Therefore, QECCs use Pauli matrices to capture the effects of errors as bit flips, phase flips, or a combination of both.

[0059] The surface code is widely considered suitable for scalable fault-tolerant quantum computing. It encodes logical qubits in a lattice using alternating data qubits and parity qubits. Figure 4 A schematic representation of such a lattice is shown at 400, with the lattice 400 having a 2D distance-3 (d = 3) surface code. Each X stabilizer 402 and Y stabilizer 404 is coupled to its adjacent data qubits 406. Each data qubit 406 interacts only with its nearest-neighbor parity qubits 408, such that errors on the data qubit 406 can be diagnosed by measuring the locally supported operators, as shown at 410. In this example, a Z error 412 on data qubit A 414 is captured by parity qubits P0 416 and P1 418. Similarly, an X error 420 on data qubit B 422 is captured by parity qubits P2 424 and P3 426. In the simplest implementation, a surface code of distance d uses (2d - 1) 2 physical qubits to store a single logical qubit, where d is a measure of redundancy and fault tolerance. Larger code distances result in greater redundancy and increased fault tolerance.

[0060] Logical operators consist of a string of single-qubit operators between two opposite edges. The encoding space is the subspace where all stabilizer generators (as shown in Figure 4 ) have an eigenvalue of +1. By construction, logical states are invariant under the application of stabilizer generators. Any closed loop of Pauli operator applications will keep the logical state unchanged. The measurement of stabilizer generators detects the endpoints of a series of errors. Error correction is based on this information, and stabilizer measurements are called syndromes.

[0061] In QEC, the effect of an error is reversed by applying appropriate Pauli gates. For example, if a qubit encounters a bit-flip error, applying a Pauli X gate flips it back to the expected state. It has been shown previously that as long as Clifford gates are applied to qubits, no active error correction needs to be performed. Instead, it is sufficient to keep track of the Pauli frame in software. Thus, the main focus of quantum error correction is error decoding rather than error correction. Optimal error decoding is a computationally difficult problem. A quantum error decoder takes the syndrome measurement as input and returns an estimate of the error in the data qubits. In addition to the ability to detect errors, the decoder also relies on high operating speed to prevent the accumulation of errors. In other words, errors must be detected faster than they occur.

[0062] Since error decoding must be fast, the decoder must provide high performance and operate near the physical qubits. As described regarding Figure 1 , qubits typically operate at 15 - 20 mK. Depending on whether the decoder is designed to operate at 4K or at 77K, the implementation technologies offer different trade-offs, as shown in Table 1. Due to proximity to the physical qubits, hardware designed to operate at 4K must meet strict power requirements. This is to ensure that thermal noise is controlled. Additionally, these designs must be cooled using complex and expensive liquid helium coolers. The decoder can be designed using CMOS or superconducting logic at 4K. CMOS has power consumption and thus cannot be used for large-scale quantum computers. Superconducting logic offers low power consumption, but has major drawbacks such as limited device density and low storage capacity, which make it extremely difficult to fabricate complex and large designs. Conventional CMOS operating at 77K offers the ability to design complex systems with larger memory and power budgets. The cooling overhead associated with 77K is an order of magnitude lower than that of 4K. However, a decoder designed to operate at 77K must account for transmission delays and meet the bandwidth required to transfer data back and forth between 4K and 77K. The trade-offs between superconducting technology at 4K and CMOS at 77K are listed in Table 1.

[0063] Parameter Superconducting technology Traditional CMOS Operating temperature 4K 77K Operating frequency 10 GHz 4 GHz Memory capacity 123 - 512 bytes 4 Gb Power budget 1W N / A Feature size 248 nm 7 - 16 nm Cooling overhead 1000x / 400x 10x

[0064] Table 1

[0065] In this paper, the challenges in designing the microarchitecture of decoders for quantum error correction under realistic noise models are examined. Qubit errors can be broadly classified into three types: decoherence errors, gate errors, and measurement errors. Qubits maintain their quantum states only for a short duration (referred to as the decoherence time), leading to decoherence errors. Non-ideal gate operations result in gate errors on qubits. Imperfections in qubit measurements lead to measurement errors. The decoder may misinterpret an accompanying measurement error as a data qubit error and correct the non-error data qubit, thereby introducing an error. The decoder must take such accompanying measurement errors into account while decoding errors. This directly affects the microarchitecture and design choices of the decoder.

[0066] Figure 5 Diagram 500 is shown, which indicates two consecutive rounds of syndrome measurements 502 and 504, and shows how measurement errors 506 pair up in time and data qubit errors 508 pair up in space. Diagram 500 shows that if the decoder only examines the measurement results of the 0th round 502, it will misinterpret the error on the parity qubit P0 510 and force a correction of the error-free D0 512. Current decoders address accompanying measurement errors by examining d rounds of measurements, where d is the code distance. The data generated by d rounds of syndrome measurements and the error log per data qubit must be stored for the decoder to work correctly. This requires storage space of up to several megabytes (depending on the code distance and the number of logical qubits).

[0067] Figure 6 An example graph 600 indicating the memory capacity (in KB) required to store the syndrome measurement data for d (code distance) rounds and the error log for N logical qubits is shown. The required capacity is much higher than the available memory in superconducting logic at 4K. To perform error decoding at 77K, the measurement data must be transferred from 4K to 77K. For a given qubit plane with L logical qubits and each qubit encoded using a surface code with a code distance of d, 2d(d - 1)L bits must be sent at the end of each syndrome measurement cycle. Assuming a reasonable number of logical qubits and code distance, the 4K - 77K link requires a bandwidth range on the order of several Gb / s. Data transfer at lower bandwidths reduces the remaining effective time for error decoding, as it must provide an estimate of the error within d syndrome measurement cycles (e.g., the surface code cycle can be broken down into d syndrome measurement cycles). Thus, the main challenge in designing any decoder at 77K is the very large required bandwidth.

[0068] One way to effectively handle the capacity and bandwidth requirements in cache and main memory is data compression. The sparsity of the measured data can be analyzed and estimated as described herein. For example, let p be the probability of a Z error on a data qubit, and let u be the error indicator vector for n data qubits (note that the same analysis also applies to the X adjoint). If there are 4 data qubits and the first two have Z errors, then u = 1100. Assuming an identical and independent error distribution (iid), the upper bound of the adjoint Hamming weight is given by Equation (1), where |u| is the Hamming weight of the error indicator vector u (e.g., the number of 1s).

[0069]

[0070] Thus, the probability of having m or more errors is given by Equation 2:

[0071]

[0072] Using the union - bound, the upper bound of the total number of adjoint bits s(Z u ) is given by Equation 3.

[0073] s(Z u ) ≤ 2|u| (Equation 3)

[0074] Assume the code distance is 11 and the error rate is 10 -3 , the probability of having 10 or more errors (for a given code distance, the number of errors is quite large) is 6.2x10 -14 . Thus, the probability of observing an adjoint with a large Hamming weight is extremely low. This analysis shows that it is possible to compress adjoint data to reduce the storage overhead for storage and / or meet the bandwidth requirements. Different compression techniques for adjoint data are described herein, as the usefulness of the compression technique depends on the entropy of the data. In this article, three compression techniques are described, although other techniques have been considered. Different noise mechanisms to which these techniques can be applied are analyzed. The examples described have compression schemes that use simple coding and do not require high hardware complexity, respectively.

[0075] Dynamic zero compression (DZC) was initially introduced to reduce the energy required for cache access of zero - valued bytes. A similar technique can be adopted to compress adjoint data. Figure 7An example is shown at 700. An adjoint 705 of length L is grouped into K blocks each of W bits, where W is the compression width 710. If the number of bits in the last block is less than W bits, additional padding zeros can be added. A K-bit wide zero indicator bit (ZIB) vector 715 includes 1 bit per block. If all bits of the i-th block are 0, the corresponding bit in the ZIB (ZIB[i]) can be set to 1. Otherwise, the bit can be set to 0. The data 720 to be transmitted can be obtained by appending non-zero blocks 725 at the end of the ZIB vector.

[0076] As shown at 750, the sparse representation can be considered similar to traditional techniques for storing sparse matrices, in which non-zero elements of a sparse matrix 760 are stored by storing only the row and column indices 765. A sparse representation bit (SRB) 755 is used to indicate whether all adjoint bits are zero. If there is one or more non-zero bits in the adjoint, the SRB can be left unset, and the indices of the non-zero elements 755 can be sent together with the SRB in the transmitted data 775.

[0077] Geometry-based compression (Geo-Comp) can be considered an adaptation of DZC that also takes into account the geometry of the surface code lattice. The geometry-based compression scheme can compress regions of X and Z adjoints together, rather than compressing X and Z adjoints separately. The entire surface code lattice can be divided into multiple regions, where each region roughly contains an equal number of adjoint bits (similar to the compression width of DZC). Figure 8 A surface code lattice 800 for geometry-based compression, including multiple regions (801, 802, 803, 804; represented by dashed lines), is schematically shown. Figure 8 An example of how a surface code lattice with a code distance of 5 is divided into 4 regions is shown. Using ZIB for each region and transmitting only adjoint data from non-zero regions can compress the adjoints. When a Y error occurs on a data qubit, the X and Z adjoint bits flip to indicate the error. When the two types of adjoints are compressed independently, for a given compression width, the total number of non-zero blocks is higher. For example, if Figure 8 the data qubit D0 810 shown in encounters a Y error, the X adjoint bits X0 811 and X1 812 and the Z adjoint bits Z0 813 and Z1 814 flip. In a compression scheme such as DZC, (X0 811, X1 812) and (Z0 813, Z1 814) are located on different data blocks and are compressed separately. However, if the geometry of the lattice is considered, non-zero adjoint bits are typically within the same region, unless the data qubit is on a region boundary (e.g., D1 815 in the lattice 800).

[0078] In general, for a given noise model, the number and size of regions can be adjusted by calculating the expected number of blocks that contain the trivial (all-zero) syndrome. However, larger-sized regions lead to complex hardware by increasing the logical depth. Therefore, even for very low error rates, smaller region sizes (depending on the code distance) can be analyzed. The regions do not need to be of equal size, and the size and / or number of regions can be determined based on the expected number of data blocks that contain the trivial syndrome.

[0079] Figure 9 An example method 900 for compressing syndrome data within a quantum computing device is shown. In some examples, method 900 can be implemented by a quantum computing device that includes a union-find decoder (e.g., Figure 12 the decoder schematically depicted in).

[0080] At 910, method 900 includes generating syndrome data from at least one quantum register that includes l logical qubits, where l is a positive integer. The generated syndrome data can include at least X syndrome data and Z syndrome data.

[0081] Continuing at 920, method 900 includes, for each logical qubit: routing the generated syndrome data to a compression engine, which is configured to compress the syndrome data. The quantum computing device can include multiple compression engines. In some examples, at least one compression engine is configured to compress the syndrome data using dynamic zero compression. In some examples, at least one compression engine is configured to compress the syndrome data using sparse representation. In some examples, at least one compression engine is configured to compress the syndrome data using geometry-based compression. The quantum computing device can include two or more logical qubit sectors coupled to two or more types of compression engines. In some examples, method 900 can include operating the compression engine at 4K. However, higher (e.g., 8K) or lower (e.g., 2K) temperatures can be used.

[0082] Continuing at 930, method 900 includes routing the compressed syndrome data to a decompression engine, which is configured to: receive the compressed syndrome data; and decompress the received compressed syndrome data. At 940, method 900 includes routing the decompressed syndrome data to a decoder block. In some examples, the decompressed syndrome data can be routed to a graph generator module of the decoder block. In some examples, method 900 can include operating the decompression engine and / or the decoder block at 77K. However, higher (e.g., 85K) or lower (e.g., 70K) temperatures can be used. In some examples, the quantum computing device includes a set of d decoder blocks, where d < 2*l.

[0083] Figure 10 Illustrates an example method 1000 for compressing syndrome data using geometry-based compression within a quantum computing device. In some examples, method 1000 may be implemented by a quantum computing device including a union-find decoder (e.g., Figure 12 the decoder schematically depicted in).

[0084] At 1010, method 1000 includes generating syndrome data from at least one surface code lattice including l logical qubits, where l is a positive integer. For example, the surface code lattice is divided into two or more regions based on the lattice geometry, as Figure 8 shown in. In some examples, the number of regions may be determined based on the expected number of data blocks containing trivial syndromes.

[0085] At 1020, method 1000 includes, for each logical qubit: routing the generated syndrome data to a compression engine configured to compress the syndrome data using geometry-based compression. At 1030, method 1000 includes compressing the syndrome data using zero indicator bits for each of two or more regions of the surface code lattice. At 1040, method 1000 includes transmitting the syndrome data only from non-zero regions. In other words, it can be assumed that if no data is received from a region, that region includes only trivial (e.g., all-zero) data.

[0086] At 1050, method 1000 includes routing the compressed syndrome data to a decompression engine configured to: receive the compressed syndrome data; and decompress the received compressed syndrome data. The decompression engine may be programmed based on the geometry-based compression scheme used by the compression engine.

[0087] The decoder for QEC is used to process the syndrome measurement data and identify errors in the corrupted data qubits. Here, the microarchitecture of the hardware implementation of the union-find decoder for the surface code is improved. In the surface code, local operators on the lattice of qubits are measured and the decoder is used to process the syndrome to generate an estimate of the most likely error on the data qubits. The decoder microarchitecture is designed to prevent the accumulation of errors while maintaining low hardware complexity to meet the strict power budget for operation in a cryogenic environment. The architecture proposed here is designed to support scaling up to thousands of logical qubits to support fault-tolerant quantum computing.

[0088] Quantum error decoding is an NP-hard problem. Thus, most decoding algorithms trade off between error thresholds and lower time and algorithmic complexities. A promising error decoding technique is the graph-based minimum weight perfect matching (MWPM) decoder. Although the MWPM decoder offers a high error threshold, it suffers from a high time complexity (O(n2)). Alternatively, a simple way to design a decoder is based on the use of a lookup table. The table is indexed by syndrome bits, and the corresponding entries store the error information of data qubits. However, the lookup table decoder is not scalable and requires terabytes of memory even for small code distances. Deep neural decoders are popular and learn the probability density function of possible errors corresponding to the measured syndrome sequence during the training phase. Using inference, the error pattern for a given syndrome is evaluated. However, neural decoders require more hardware for computation and are not scalable as the code distance increases. The recently proposed union-find decoder presents an algorithm that forms clusters around non-trivial syndromes (non-zero syndromes) and performs error correction in nearly linear time using graph traversal. The union-find decoder thus offers simplicity, time complexity, and a high error threshold.

[0089] The operation of the union-find decoder is as Figure 11 shown. At 1100, each edge on graph 1102 represents a data qubit, and each vertex represents a parity qubit (e.g., 1 1 0 4, 11 0 6). Decoding begins by growing a spanning forest 1108 to cover all error syndrome bits, thus forming one or more even clusters, as shown at 1110. Data qubits A 1112 and B 1114 can be assigned unknown Pauli errors 1116 and 1118, respectively. By traversing the forest, errors can be detected, as shown at 1120. The cluster traversal steps (shown at 1122, 1124) can be used to detect, classify (e.g., Z error), and correct errors.

[0090] As Figure 12 shown in the block diagram 1200 of, an adaptation of the algorithm can be implemented. The compressed syndrome data 1210 is routed to the decompression engine 1215. The decompressed syndrome data 1220 is then routed to the graph generator (Gr-Gen) module 1225. The Gr-Gen module 1225 can be configured to generate spanning tree memory (STM) data. The depth-first search engine (DFS) 1230 can be configured to access the STM data and generate an edge stack based on the STM data. The correction (Corr) engine 1235 can be configured to access the edge stack, generate memory requests based on the accessed edge stack, and update the error log 1240.

[0091] If adjoint measurement errors are ignored, decoding is performed using the 2D graph generated from the adjoint measurements of a single round. To account for erroneous measurements, d consecutive adjoint measurements, where d is the code distance, must be decoded together, resulting in a 3D graph. The union-find decoder can be used for both cases. The main difference is that the required memory amount grows quadratically (for 2D) or cubically (for 3D) with the code distance of the surface code. For simplicity, the microarchitecture of the union-find decoder is described in 2D and extended to 3D. All relevant results described are obtained for 3D graphs. The decoding design includes 3 pipeline stages, enabling improved design scalability.

[0092] The Gr-Gen module takes the adjoint as input after decompression and generates a spanning forest by growing clusters around non-trivial adjoint bits (non-zero adjoint bits). The spanning forest can be constructed using two basic graph operations: Union() and Find(). Figure 13 An example Gr-Gen module 1300 is schematically shown. Module 1300 includes a spanning tree memory (STM) 1310, a zero data register (ZDR) 1315, a root table 1320, a size table 1325, a parity register 1330, and a fused edge stack 1335. This design is slightly different from the union-find algorithm described earlier for reducing hardware resource costs. The size of each component is a function of the code distance d. STM 1310 stores 1 bit for each vertex and 2 bits for each edge. Each edge uses 2 bits because, according to the original algorithm, the cluster grows half the edge width around the vertex or the existing cluster boundary. ZDR 1315 stores 1 bit per STM row. If the content of the row is 0, the bit stores 0, and if at least one bit in the row is 1, the corresponding ZDR bit for the row stores 1. Since the adjoint data is sparse and the total number of edges in the spanning forest will be low, ZDR 1315 accelerates the traversal of STM 1310. FES 1335 stores the newly grown edges so that they can be added to the existing clusters. The root table 1320 and the size table 1325 store the root and size of the clusters, respectively. The tree traversal register 1340 stores the vertices of each cluster accessed in the Find() operation. The interface 1345 between the Gr-Gen module and the DFS engine can allow the DFS engine to access the data stored at STM 1310.

[0093] As Figure 14As shown, the root table entry (root table[i]) is initialized to the index (i). As shown at 1400, the size table entry for non-trivial adjoint bits is initialized to 1. These tables help the Union() and Find() operations to merge clusters into the final state shown at 1420 after the growth phase shown at 1410. They are indexed by the cluster index. The size of the table is set for the maximum possible number of clusters, which is equal to the total number of vertices in the surface code lattice. A boundary list for each cluster can be stored. However, in the noise mechanisms relevant to practical applications, the average cluster diameter is very small. The cluster diameter can be defined as the maximum distance between two vertices on the cluster boundary. Therefore, instead of storing the boundary list, the boundary index can be calculated during the cluster growth phase. The original algorithm grows all odd clusters until the parity is even. Therefore, odd clusters must be detected quickly. To do this, a parity register can be used, as Figure 11 shown. The parity register can store 1 bit of parity per cluster, depending on whether it is odd or even. For a reasonable code distance of 11, seven 32-bit registers may be sufficient. For larger code distances, additional parity information can be stored in memory and prefetched to hide the memory latency.

[0094] The control logic can read the parity register and grow the clusters with odd parity (referred to as the growth phase) by writing to the STM, ZDR, and adding the newly added edges that touch the boundaries of other clusters to the FES. For the edges connecting to other clusters, the STM may not be updated to prevent double growth. It can be updated when the clusters are merged by reading from the FES. This logic can check whether the newly added edges connect two clusters by reading the root table entries of the vertices connected by the edges (call these the primary vertices). This is equivalent to the Find() operation. As Figure 15 shown at 1500 in, the vertices visited on the path to find the root of each primary vertex are stored in the tree traversal register. As shown at 1510, the root table entries of these vertices can be updated to point directly to the root of the cluster to minimize the depth of the tree for future traversals. This operation (referred to as path compression) is included in the union-find algorithm and allows the depth of the tree to be kept short, thus amortizing the cost of the Find() operation. For example, at 1500, Figure 15 shows the state of two clusters and the root table at a certain moment. Suppose that after a growth step, vertices 0 and 6 are connected and the two clusters must be merged. The tree traversal register can be used to update the root of vertex 0, as shown at 1500. Since the depth of the tree is continuously compressed, only a few registers are sufficient. In one example, 5 registers are used per primary vertex, although more or fewer registers can also be used. If the primary vertices belong to different clusters, the root of the smaller cluster can be updated to point to the root of the larger cluster.

[0095] The DFS engine can process STM data generated by Gr-Gen that stores a set of even clusters growing. It can use the DFS algorithm to generate a list of edges that form a spanning tree for each cluster in the STM. In other examples, breadth-first search exploration can be used, although DFS is generally more memory-efficient. At Figure 16 An example DFS engine is shown at 1600 of. This logic can be implemented using a finite state machine 1610 and two stacks 1620 and 1622. Stacks can be used because the order of visiting edges in the spanning tree can be reversed to perform correction by peeling. The edge stack 1620 can store a list of visited edges, while the pending edge stack 1622 can store edges that will be visited later in the ongoing DFS. For example, as Figure 16 shown at 1630 of, when the FSM visits vertex 1 of the spanning tree, edge a is pushed onto the edge stack and edge c is pushed onto the pending edge stack. When the end of the current path is reached, the pending edge can be popped and traversed. To support pipelined operation and improve performance, the microarchitecture can be designed to include an alternative edge stack 1632. When there is more than one cluster, the correction engine can work on the edge list of another cluster being traversed while the DFS engine traverses one cluster via the Corr engine interface 1640. As shown at 1630, if edges a, b, c, and d belong to cluster C0 and edges e and f belong to cluster C1, the DFS engine 1600 can traverse C1 while the Corr engine processes the correction of C0. This can help set the size of the stack to handle the average cluster size rather than the worst-case cluster size. In the case where the DFS engine 1600 encounters a cluster large enough that it cannot fit in one stack, the alternative stack 1632 can be used, and an overflow bit can be set to indicate that stacks 1620 and 1632 hold edges corresponding to a single cluster. This proposed implementation can include a number of memory reads proportional to the size of the cluster. By checking the STM1310 line by line, the effective cost of generating clusters is reduced. ZDR1315 reduces the cost of traversing the STM 1310 line by line.

[0096] The Corr engine can perform the peeling process of the decoder and can identify the Pauli corrections to be applied. The Corr engine can access the edge list (which is stored on the stack) and the syndrome bits corresponding to the vertices along the edge list. The syndrome bits can be accessed by decompressing the compressed syndrome and / or by accessing the STM. However, the former may increase the logic complexity and latency, while the latter may increase the number of memory requests that the STM needs to handle. To reduce the memory traffic and eliminate the need for additional decompression logic, the syndrome information can be deposited by the DFS engine together with the edge index information. The temporary syndrome changes caused by peeling are stored in local registers. An example of the peeling of the error graph performed in the Corr engine is asFigure 17 as shown in Figure 17 Figure 17 shows an example of step 1 1700, step 2 1710, and step 3 1720 along with examples of save registers, error logs, edge stacks, and error graphs. The Corr engine can also read the last surface code cycle error log and can update the Pauli corrections for the current edges. For example, if the error on edge e0 was Z in the previous logical cycle and it also encounters a Z error in the current cycle, the Pauli error for e0 can be updated to I, as shown at 1720.

[0097] Figure 18 Figure 18 shows an example decoding method 1800 for a quantum computing device. In some examples, the decoding method 1800 can be implemented by a quantum computing device including a union-find decoder (such as Figure 12 the decoder schematically depicted in

[0098] At 1805, method 1800 includes receiving syndrome data from one or more qubits (such as logical qubits residing in a quantum register). The received syndrome data can include X syndrome data and / or Z syndrome data.

[0099] At 1810, method 1800 includes decoding the received syndrome data with a union-find decoder implemented in hardware including two or more pipeline stages. As an example, this can include Figure 12 a hardware-implemented union-find decoder including three pipeline stages as shown in

[0100] Optionally, at 1820, decoding the syndrome data can include generating a spanning forest in the Gr-Gen module by growing clusters around non-trivial syndrome bits. In some examples, Union() and Find() graph operations can be used to generate a spanning tree.

[0101] Optionally, at 1825, decoding the syndrome data can include storing data about the spanning forest in a spanning tree memory (STM) and a zero data register at the Gr-Gen module. In some examples, newly grown edges can be stored at the fused edge stack.

[0102] Optionally, at 1830, decoding the syndrome data can include accessing the data stored in the STM at the DFS engine. Optionally, at 1835, decoding the syndrome data can include generating one or more edge stacks at the DFS engine based on the data stored in the STM. For example, as Figure 16As shown, generating one or more edge stacks based on data stored in the STM may include generating a primary edge stack that includes a list of visited edges. Additionally or alternatively, generating one or more edge stacks based on data stored in the STM may include generating a pending edge stack that includes a list of edges to be visited. Additionally or alternatively, generating one or more edge stacks based on data stored in the STM may include generating an alternative edge stack that is configured to hold the remaining edges of clusters from the spanning forest.

[0103] Optionally, at 1840, decoding the adjoint data may include accessing one or more of the generated edge stacks at the Corr engine. Optionally, at 1845, decoding the adjoint data may include generating a memory request at the Corr engine based on the accessed edge stacks. Optionally, at 1850, decoding the adjoint data may include performing iterative peeling decoding on each of the accessed edge stacks at the Corr engine. Optionally, at 1855, decoding the adjoint data may include updating the decoder's error log at the Corr engine based on the results of the iterative peeling decoding.

[0104] As discussed herein, decoding based on a single round of measurements will not account for adjoint measurement errors. To cope with measurement errors, the decoder examines d (code distance) rounds of measurements. This type of error correction can be accomplished with minimal changes to the design. For example, the decoder can analyze a 3D graph instead of forming a graph on a 2D plane. Each vertex can be connected to up to 4 neighbors. However, for a 3D graph, each vertex can now have up to two additional edges corresponding to the previous and next rounds of measurements. To reduce the storage overhead, the STM for each round of adjoint measurements can be stored. The STM can be optimized such that each row of the STM stores the vertices of a row of the surface code lattice, the edge information of the vertices of the next row, and the edge information connecting the corresponding vertices in the surface code lattice of the next round.

[0105] The compression techniques described herein can reduce the amount of memory required to store the adjoint data and error logs of data qubits. However, the microarchitecture of the union-find decoder also uses memory, and the total capacity required is far from the total capacity provided by superconducting memory. Therefore, the design can be implemented by using conventional CMOS operating at 77K. This can also reduce the thermal noise generated in the cryogenic environment near the quantum substrate, since the design is physically far from the quantum substrate.

[0106] For the baseline design, a naive implementation can allocate a decoder for each X adjoint and each Z adjoint of each logical qubit, as Figure 19 shown at 1900 of Figure 19The system structure of a large number of logical qubits within the quantum register 1910 is schematically shown. The quantum register 1910 is shown to include logical qubit 0 1910a, logical qubit 1 1910b, and logical qubit l 1910l as representative logical qubits operating at 15 - 20 mK. Each logical qubit is configured to receive signals from the control logic 1915 and output adjoint data to the compression engines (e.g., 1920a, 1920b... 1920l). The control logic 1915 and the compression engines 1920a... 1920l are shown to operate at a higher temperature of 4K than the quantum register. However, higher (e.g., 8K) or lower (e.g., 2K) temperatures can be used.

[0107] Each compression engine routes the compressed adjoint data to the decompression engines (1925a, 1925b... 1925l) operating at 77K. The decompression engines decompress the compressed adjoint data and route the decompressed X adjoint data and Z adjoint data to the decoding block 1930. In this example, each decompression engine is coupled to a pair of pipelined union-find decoders (1935a, 1935b, 1935c, 1935d... 1935k, 1935l) operating at 77K. Each union-find decoder analyzes the adjoint data received from the decompression engine and updates the error log 1940. Although shown to operate at 77K, higher (e.g., 85K) or lower (e.g., 70K) temperatures can be used to operate the decompression engines and decoders, although the operating temperatures of the decompression engines and decoders are generally higher than those of the compression engines.

[0108] Thus, for the baseline design, the decoding logic can use 2L union-find decoders per logical qubit. In this implementation, each logical qubit uses its own dedicated decoder. However, the utilization rate of each pipeline stage may be different. Therefore, the architecture shown at 1900 may not provide the optimal allocation of resources. For a large number of qubits, the on-chip components are underutilized and dissipate heat. Since the entire system operates at 77K, the increased power consumption linearly increases the cost of cooling.

[0109] Thus, an architecture including a decoder block with a reduced number of pipeline units can be used. In Figure 20FIG. 2000 shows an example design of such a decoder block. A qubit register 2005 including a plurality of logical qubits transfers adjoint data to a set of Gr-Gen modules. A set of Gr-Gen modules 2010 may share one or more DFS engines 2020, and a set of DFS engines 2020 may share one or more Corr engines 2030. The hardware overhead includes a first set of multiplexers 2035 coupling the set of Gr-Gen modules 2010 to a DFS engine 2020, and a second set of multiplexers 2040 coupling the set of DFS engines 2020 to a Corr engine 2030. Memory requests generated by the Corr engine 2030 may be routed to the correct memory location using a demultiplexer 2045. Selection logic 2050 may prioritize the first ready component and may use round robin arbitration to generate appropriate selection signals for the multiplexers 2035 and 2040. For example, if four Gr-Gen modules 2010 share a DFS engine 2020 and the second Gr-Gen module finishes cluster formation earlier than the others, it may access the corresponding DFS engine 2020 first. Thus, the round robin strategy ensures fairness while sharing resources.

[0110] In Figure 21 FIG. 2100 shows an example system architecture. A qubit register 2105 includes a plurality of logical qubits 2110 coupled to control logic 2115. Each logical qubit is coupled to a compression engine 2120, and each compression engine 2120 is in turn coupled to a decompression engine 2125. A block of N logical qubits 2110 shares a decoder block 2130, and the decoder block 2130 updates an error log 2135 for each coupled logical qubit 2110. As described with respect to Figure 19 the operating temperature may vary with the indicated temperatures of 4K and 77K. If N logical qubits share a decoder block 2130, then for a quantum register 2105 having L logical qubits 2110, the total number of decoder blocks 2130 required is L / N. An example microarchitecture uses L Gr-Gen modules, (a) L DFS engines, and (b) L Corr engines. Resource savings depend on the parameters (a) and (b). The values of (a) and (b) can be calculated to minimize the overall hardware cost. This can be framed as a constrained optimization problem.

[0111] One way to decode a large-scale system is to assign a decoder to each logical qubit. However, this method results in linear growth with respect to hardware, and thus linear growth in power cost. Consequently, this design is not very efficient and is not scalable by itself. The design herein supports the reuse of specific design components in order to reduce the actual cost when decoder blocks are scaled for a large number of logical qubits.

[0112] Resources can be shared within and / or across decoding units. Given the distribution of decoding times, it is unlikely that several very long adjoint vectors will need to be decoded simultaneously, so resources can be shared.

[0113] This sharing is independent of the decoder or decoding algorithm, including cases where the decoding algorithm has a runtime that depends on the adjoint, so some adjoints may be more difficult or take longer to decode than others. For example, some machine learning-based decoders do not depend on adjoints. A machine learning decoder can have a multi-layer neural network. Once decoding is performed on one qubit on the first layer, a second qubit can use the first layer while the first qubit works on the second layer of the network.

[0114] Figure 22 An example method 2200 of a quantum computing device is shown. Method 2200 can be performed by a multiplexed quantum computing device (such as, Figure 20 and Figure 21 the computing device shown). At 2205, method 2200 includes generating an adjoint from at least one quantum register including l logical qubits, where l is a positive integer. The generated adjoint can include X and Z adjoints. At 2210, method 2200 includes routing the generated adjoint to a set of d decoder blocks coupled to the at least one quantum register, where d < 2*l. As described with respect to Figure 20 and Figure 21 this allows for scalability of the quantum computing device, as fewer than two decoders are needed to handle the processing of X and Z adjoints for each logical qubit.

[0115] In some examples, each decoder block is configured to receive a decoding request from a set of n logical qubits, where n > 1. In some examples, each decoder block includes g Gr-Gen modules, where 0 < g ≤ l, and each Gr-Gen module is configured to generate spanning tree memory (STM) data based on the received adjoint. In some examples, each decoder block further includes α*l DFS engines, where 0 < α < 1. In some examples, two or more Gr-Gen modules are coupled to each DFS engine via a multiplexer in a first set of multiplexers.

[0116] Optionally, at 2215, method 2200 includes accessing, at each DFS engine, STM data generated by two or more Gr-Gen modules via one of a first set of multiplexers. Optionally, at 2220, method 2200 includes generating an edge stack at each DFS engine based on the STM data. In some examples, each decoder block further includes β*l Corr engines, where 0 < β < 1. In some examples, two or more DFS engines are coupled to each Corr engine via one of a second set of multiplexers.

[0117] Optionally, at 2225, method 2200 includes accessing, at each Corr engine, the edge stack generated by two or more DFS engines via one of a second set of multiplexers. Optionally, at 2230, method 2200 includes generating a memory request based on the accessed edge stack. Optionally, at 2235, method 2200 includes routing, via one or more demultiplexers, the memory requests generated by each Corr engine to a memory location. Optionally, at 2240, method 2200 includes routing return signals via each multiplexer in the first set of multiplexers and the second set of multiplexers based on round-robin arbitration.

[0118] When decoding d rounds of syndrome measurements within one logical cycle (T), error correction is successful, which limits the maximum latency that the decoder can tolerate. When the decoder fails to decode all syndromes within a logical cycle, errors may not be detected. This type of failure can be referred to as a timeout failure. Since the decoder is imperfect and exhibits threshold behavior, there is also a possibility of a logical error occurring when the correction generated by the decoder changes the logical state of a qubit. Therefore, the failure of the decoder can be attributed to a timeout failure or a logical error. To keep the error threshold the same and prevent the system failure rate from increasing, the probability of a timeout failure (ptof) must be lower than the probability of a logical error occurring (plog), as shown in Equation (4). For an optimized design, resource sharing is possible as long as ptof is small enough.

[0119] p tof ≤p log (Equation 4)

[0120] Assume that N logical qubits with exactly the same error rate share k decoding units. The total execution time for decoding N logical qubits is given by Equation (5):

[0121]

[0122] where τ i represents the execution time for decoding the syndrome of the i-th logical qubit. In this case, the probability of a timeout failure p tof must satisfy Equation (6).

[0123]

[0124] The optimization goal is to minimize the number of decoding units k for a given number of logical qubits N such that the constraints given by equation (4) are satisfied. The ptof can be modeled using the execution time obtained from the simulator.

[0125] The decoder performance can be modeled by studying the number of reads. The write operations performed can be read-modify-write, and the write-back may not be on the critical path. For memory access and a 4 GHz clock frequency, a latency of 4 cycles is assumed. The total number of memory requests in Gr-Gen for a given syndrome is proportional to the cluster diameter (Di). However, it is proportional to the cluster size (Si) in the DFS engine and the Corr engine. Equations (7) and (8) give the execution times spent in Gr-Gen (TGG), DFS engine (TDFS), and Corr engine (TCE) for a syndrome with n clusters.

[0126]

[0127] τ DFS = τ CE = Σ i S i (Equation 8)

[0128] In the optimized design, each Gr-Gen unit grows clusters for X and Z syndromes. Two or more Gr-Gen units use one DFS engine module, and two or more DFS engines use one Corr engine. These numbers of units to be shared can be determined by the fraction of the total execution time spent in each pipeline stage.

[0129] Next, the simulation infrastructure for making design choices in the decoder microarchitecture is discussed. This infrastructure supports the estimation of some key statistics of the union-find decoder and also supports the study of the performance of the compression techniques described in this paper.

[0130] The Monte Carlo simulator is used to analyze the performance of different compression techniques and obtain statistical data on the performance of the union-find decoder. Figure 23The Monte Carlo simulator 2300 is shown schematically. Different configurations span four different physical error rates, ten different code distances, and four noise models, each simulated for one million trials. The selected error rates are 10-6 (most optimistic), 10-4, 10-3, and 10-2 (most pessimistic). The simulator 2300 accepts a code distance 2302, a noise model 2304, and a compression algorithm 2306. Based on the code distance 2302, the simulator 2300 generates a surface code lattice via a lattice generator 2308. Depending on the selected noise model 2304, the simulator injects errors on the data qubits of the surface code lattice via error injection 2310 and generates an output syndrome 2312. The output syndrome 2312 is then compressed via a compressor 2314 according to the input compression algorithm 2306 to generate a compressed syndrome 2316. The simulator 2300 then outputs a compression ratio. As a figure of merit for determining the most suitable compression scheme, the compression ratio (determined by equation (9)) and the percentage of incompressible syndromes are used. The simulation is repeated one million times to calculate the average compression ratio and the percentage of incompressible syndromes.

[0131]

[0132] The simulator also runs a union-find decoding algorithm on the syndrome 2312 via a decoder 2318. A statistics generator 2320 then analyzes the distribution of cluster sizes, the average number of clusters on a given lattice, and the execution time spent in each pipeline stage of the decoder 2318 by modeling the hardware. These statistics and performance numbers provide insights that help in the microarchitecture design of the decoder hardware implementation and drive scalable design.

[0133] The performance of the decoder depends to a large extent on the noise model of the underlying qubits. Therefore, four different error models are explored. Assuming identically and independently distributed (iid) errors, the depolarizing noise model is selected as the most basic noise model. In the depolarizing noise model, if the error rate is p, the probability that each physical qubit encounters an error is p, and the probability of remaining error-free is (1 - p). Additionally, in this error model, X, Y, and Z errors occur with equal probability p / 3. The other three noise models assume different probabilities for X and Z errors, as shown in Table 2.

[0134]

[0135] Table 2

[0136] This paper discusses the results of syndrome compression, baseline union-find decoder design, and scalability analysis. The results of the baseline decoder and scalability analysis are based on the d (code distance) rounds of syndrome measurements described in this paper.

[0137] The performance of each compression scheme depends on the noise model. For the depolarizing noise model, compression schemes such as DZC and Geo-Comp provide better performance at low code distances that depend on the error rate, compared to sparse representation. For noise models with relative biases for specific types of errors (such as Px = 10Pz and Px = 100Pz), DZC performs better than Geo-Comp. For lower code distances, even though sparse representation provides a higher compression ratio, for larger error rates, the percentage of incompressible syndromes is higher (up to 6%). For noise models where one type of error probability is much larger than the other, better compression ratios are obtained by separately compressing the X and Z syndromes at the cost of greater hardware complexity. If only one type of compression can be used due to hardware constraints, for lower code distances, DZC has better performance. Table 3 specifies different noise mechanisms and the appropriate compression schemes that are most effective in each mechanism. Overall, for most cases in mechanisms with low error rates, sparse representation has better performance.

[0138]

[0139] Table 3

[0140] Figure 24 A graph 2400 is shown, and graph 2400 shows the mean compression ratio of the X syndrome of a depolarizing noise channel using the selected compression scheme for different physical error rates and noise mechanisms. The depolarizing noise channel is considered a representative candidate. Similar results are observed for the Z syndrome.

[0141] The distribution of the cluster diameter was determined from the simulation. As defined herein, the cluster diameter is the maximum distance between any two boundary vertices of the cluster. Figure 25 A graph 2500 is shown, and graph 2500 indicates the average cluster diameter for different error rates and code distances of logical qubits. The average cluster diameter is low. This result is used to eliminate the storage cost incurred when maintaining the boundary list of each cluster in the hardware (i.e., the function used in the original union-find algorithm). This reduces the hardware cost of the Gr-Gen module. The probability that the cluster diameter will become smaller increases as the error rate decreases.

[0142] The spanning tree memory (STM) used by the Gr-Gen module and the DFS engine occupies most of the storage cost. Figure 26Figure 2600 is shown, which indicates the total memory capacity required for a spanning tree memory (STM) for a given code distance (d) and number of logical qubits (N). This shown result is a 3D plot constructed using d-round measurements. Figure 2600 shows that even for a large number of logical qubits (such as 1000), for very large code distances (d) and d-round measurements, the total memory required to decode both the X adjoint and Z adjoint is less than 10 MB. If the decoder does not need to consider d-round measurements (assuming perfect measurements may be possible in the future), the required total memory capacity will be reduced by a factor of d.

[0143] The maximum possible number of entries in the root table and the size table is the total number of adjoint bits for d (code distance) round adjoint measurements (equal to 2d(d - 1)). Each root table entry includes a root that can be uniquely identified using log22d2(d - 1) bits. Similarly, the maximum feasible cluster size includes all adjoint bits. Thus, for each logical qubit, the total size of the root table and the size table is 2d2(d - 1)log22d2(d - 1) bits.

[0144] The size of the stack can be determined by analyzing the maximum number of edges within a cluster from Monte Carlo simulations. The number of edges in a cluster follows a Poisson distribution. Figure 27 An example Figure 2700 is shown indicating such a distribution for code distance d = 11 and physical error rate p = 10-3. Thus, the stack size can be designed as half of the maximum number of edges. Figure 28 Figure 2800 is a plot showing the average number of edges in a cluster for different code distances and error rates. Each stack stores two vertices (log24d2(d - 1) bits), a growth direction (2 bits), and 1 bit of adjoint. It is noted that each DFS engine includes 2 stacks for pipelining. If the size of a cluster is greater than what each stack can hold, an overflow bit can be set and an alternative stack used when available.

[0145] Figure 29 Figure 2900 is shown indicating the correlation between Gr-Gen and the execution time in the DFS engine. This means that more time is spent in the Gr-Gen unit during decoding. As described herein, this data is used to select the number of resources to be shared within the decoder block. Figure 30 Figure [not clear what is to be filled here] is shown indicating for a single decoder block with code distance (d) of 11 and error rate (p) of 0.5x10-3 (e.g., as Figure 20A graph 3000 of the execution time distribution as shown. The shaded area indicates events that increase the probability of a timeout failure. In the case of implementing resource sharing, the probability ptof of a timeout failure is lower than the probability of a logical error rate of 10-8. For L logical qubits, the numbers of Gr-Gen modules, DFS engines, and Corr engines used in this architecture are L, L / 2, and L / 2 respectively. Therefore, the total numbers of Gr-Gen modules, DFS engines, and Corr engines are reduced by 2x, 4x, and 4x respectively.

[0146] Error correction is an integral part of classical computing related to quantum computers. Error decoding algorithms are designed to achieve higher error correction capabilities (thresholds). In this article, a microarchitecture for the hardware implementation of a union-find decoder is disclosed, which uses CMOS operating at 77K. Accompanying compression is feasible to meet the bandwidth requirements of 4K–77K links. Different compression schemes perform differently under different noise mechanisms, where sparse data representation generally performs better for lower error rates and larger code distances. The disclosed microarchitecture is designed to scale the decoder to thousands of logical qubits. The architecture includes three pipeline stages and is tuned for high performance, high throughput, and low hardware complexity. The design can be scaled for a larger number of logical qubits for practical fault-tolerant quantum computing. The time spent in each pipeline stage is different, so the utilization rate of each stage is also different. By considering this, an architecture that relies on resource sharing across multiple logical qubits is disclosed. Supporting such resource sharing enables the logical error rate to remain unaffected and minimizes the system failure rate caused by the inability to decode errors due to a lack of decoding resources.

[0147] In some embodiments, the methods and processes described herein can be associated with the computing systems of one or more computing devices. Specifically, such methods and processes can be implemented as computer applications or services, application programming interfaces (APIs), libraries, and / or other computer program products.

[0148] Figure 31 A non-limiting embodiment of a computing system 3100 is schematically shown, and the computing system 3100 can implement one or more of the above methods and processes. The computing system 3100 is shown in a simplified form. The computing system 3100 can embody the host computer device described above and shown in Figure 1 . The computing system 3100 can take the form of one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, gaming devices, mobile computing devices, mobile communication devices (e.g., smart phones), and / or other computing devices, as well as wearable computing devices (such as smart watches and head-mounted augmented reality devices).

[0149] The computing system 3100 includes a logical processor 3102, volatile memory 3104, and a non-volatile storage device 3106. The computing system 3100 may optionally include a display subsystem 3108, an input subsystem 3110, a communication subsystem 3112, and / or Figure 31 other components not shown.

[0150] The logical processor 3102 includes one or more physical devices configured to execute instructions. For example, the logical processor may be configured to execute instructions as part of one or more applications, programs, routines, libraries, objects, components, data structures, or other logical constructs. Such instructions may be implemented to perform tasks, implement data types, transform the state of one or more components, achieve a technical effect, or otherwise achieve a desired result.

[0151] The logical processor may include one or more physical processors (hardware) configured to execute software instructions. Additionally or alternatively, the logical processor may include one or more hardware logic circuits or firmware devices configured to execute hardware-implemented logic or firmware instructions. The processors of the logical processor 3102 may be single-core or multi-core, and the instructions executed thereon may be configured for sequential, parallel, and / or distributed processing. Optionally, the various components of the logical processor may be distributed among two or more separate devices, which may be located remotely and / or configured for cooperative processing. Aspects of the logical processor may be virtualized and executed by remotely accessible networked computing devices configured in a cloud computing configuration. In such a case, it can be understood that these virtualized aspects run on different physical logical processors of various different machines.

[0152] The non-volatile storage device 3106 includes one or more physical devices configured to store instructions executable by the logical processor to implement the methods and processes described herein. When implementing such methods and processes, the state of the non-volatile storage device 3106 may be transformed—for example, to store different data.

[0153] The non-volatile storage device 3106 may include removable and / or built-in physical devices. The non-volatile storage device 3106 may include optical memories (e.g., CDs, DVDs, HD-DVDs, Blu-ray discs, etc.), semiconductor memories (e.g., ROMs, EPROMs, EEPROMs, flash memories, etc.), and / or magnetic memories (e.g., hard disk drives, floppy disk drives, tape drives, MRAMs, etc.) or other mass storage device technologies. The non-volatile storage device 3106 may include non-volatile, dynamic, static, read / write, read-only, sequential access, location-addressable, file-addressable, and / or content-addressable devices. It will be understood that the non-volatile storage device 3106 is configured to save instructions even when the non-volatile storage device 3106 is powered off.

[0154] The volatile memory 3104 may include a physical device containing random access memory. The logical processor 3102 typically utilizes the volatile memory 3104 to temporarily store information during software instruction processing. It will be understood that when the volatile memory 3104 is powered off, the volatile memory 3104 generally does not continue to store instructions.

[0155] Aspects of the logical processor 3102, the volatile memory 3104, and the non-volatile storage device 3106 may be integrated together into one or more hardware logic components. For example, such hardware logic components may include field programmable gate arrays (FPGAs), program and application specific integrated circuits (PASIC / ASICs), program and application specific standard products (PSSP / ASSPs), systems on a chip (SOCs), and complex programmable logic devices (CPLDs).

[0156] When included, the display subsystem 3108 may be used to present a visual representation of data saved by the non-volatile storage device 3106. The visual representation may take the form of a graphical user interface (GUI). Since the methods and processes described herein change the data saved by the non-volatile storage device, thereby transforming the state of the non-volatile storage device, the state of the display subsystem 3108 may likewise be transformed to visually represent the changes in the underlying data. The display subsystem 3108 may include one or more display devices utilizing almost any type of technology. Such display devices may be combined with the logical processor 3102, the volatile memory 3104, and / or the non-volatile storage device 3106 in a shared enclosure, or such display devices may be peripheral display devices.

[0157] When included, the input subsystem 3110 can include one or more user input devices (such as a keyboard, mouse, touch screen, or game controller), or interface with user input devices. In some embodiments, the input subsystem can include selected natural user input (NUI) components, or interface with selected NUI components. Such components can be integrated or peripheral, and the conversion and / or processing of input actions can be handled on or off the system. Example NUI components can include a microphone for voice and / or speech recognition; infrared, color, stereo, and / or depth cameras for machine vision and / or gesture recognition; a head tracker, eye tracker, accelerometer, and / or gyroscope for motion detection and / or intent recognition; and an electric field sensing component for evaluating brain activity; and / or any other suitable sensors.

[0158] When included, the communication subsystem 3112 can be configured to communicatively couple the various computing devices described herein to each other, and to other devices. The communication subsystem 3112 can include wired and / or wireless communication devices compatible with one or more different communication protocols. As a non-limiting example, the communication subsystem can be configured to communicate via a wireless telephone network, or a wired or wireless local or wide area network (such as HDMI over a Wi-Fi connection). In some embodiments, the communication subsystem can allow the computing system 3100 to send messages to and / or receive messages from other devices via a network such as the Internet.

[0159] As an example, a quantum computing device includes: at least one quantum register including a plurality of qubits; and a hardware decoder configured to: receive adjoint data from one or more of the plurality of qubits; and decode the received adjoint data by implementing a union-find decoding algorithm via a hardware microarchitecture including two or more pipeline stages. In such an example or any other example, the two or more pipeline stages additionally or alternatively include: a graph generator (Gr-Gen) module configured to: generate a spanning forest by growing clusters around non-trivial adjoint bits; and store data about the spanning forest in a spanning tree memory (STM) and a zero data register; a depth first search (DFS) engine configured to: access data stored in the STM; and generate one or more edge stacks based on the data stored in the STM; and a correction (Corr) engine configured to: access one or more edge stacks; generate memory requests based on the accessed edge stacks; and update an error log based on the decoded adjoint data. In any of the foregoing examples or any other example, the DFS engine is additionally or alternatively configured to generate: a primary edge stack including a list of visited edges; a pending edge stack including a list of edges to be visited; and an alternative edge stack configured to hold remaining edges of clusters from the spanning forest. In any of the foregoing examples or any other example, at least one of the two or more pipeline stages is additionally or alternatively coupled to two or more upstream pipeline stages via a multiplexer.

[0160] In another example, a quantum computing device includes: at least one quantum register including a plurality of logical qubits; and a hardware decoder configured to receive syndrome data from one or more of the plurality of logical qubits and decode the received syndrome data. The hardware decoder includes: a graph generator (Gr-Gen) module configured to: generate a spanning forest by growing clusters around non-trivial syndrome bits; and store data regarding the spanning forest in a spanning tree memory (STM) and a zero data register; a depth-first search (DFS) engine configured to: access data stored in the STM; and generate one or more edge stacks based on the data stored in the STM; and a correction (Corr) engine configured to: access one or more edge stacks; and generate memory requests based on the accessed edge stacks. In such an example or any other example, the Corr engine is additionally or alternatively configured to: perform iterative peeling decoding on each accessed edge stack; and update an error log of the hardware decoder based on the results of the iterative peeling decoding. In any of the foregoing examples or any other example, the hardware decoder additionally or alternatively analyzes d consecutive rounds of received syndrome data together in a 3D graph, where d is the code distance. In any of the foregoing examples or any other example, the Gr-Gen module is additionally or alternatively configured to generate a spanning forest using Union() and Find() graph operations. In any of the foregoing examples or any other example, the Gr-Gen module additionally or alternatively includes a fused edge stack configured to store newly grown edges. In any of the foregoing examples or any other example, additionally or alternatively, the STM is updated based on an indication received from the fused edge stack that two or more clusters have been merged. In any of the foregoing examples or any other example, the DFS engine is additionally or alternatively configured to generate: a primary edge stack including a list of visited edges; a pending edge stack including a list of edges to be visited; and an alternative edge stack configured to hold remaining edges from one or more clusters from the spanning forest. In any of the foregoing examples or any other example, the decompressed syndrome data is additionally or alternatively routed from a decompression engine to the Gr-Gen module. In any of the foregoing examples or any other example, two or more decompression engines are additionally or alternatively coupled to the hardware decoder.

[0161] In yet another example, a decoding method for a quantum computing device includes: receiving syndrome data from one or more qubits among a plurality of qubits; and decoding the syndrome data with a union-find decoder implemented in hardware including two or more pipeline stages. In such an example or any other example, decoding the syndrome data with a union-find decoder implemented in hardware including two or more pipeline stages additionally or alternatively includes: at a graph generator (Gr-Gen) module: generating a spanning forest by growing clusters around non-trivial syndrome bits; and storing data regarding the spanning forest in a spanning tree memory (STM) and a zero data register; at a depth-first search (DFS) engine: accessing the data stored in the STM; and generating one or more edge stacks based on the data stored in the STM; and at a correction (Corr) engine: accessing one or more generated edge stacks; and generating memory requests based on the accessed edge stacks. In any of the foregoing examples or any other example, the decoding method additionally or alternatively includes, at the Corr engine: performing iterative peeling decoding on each accessed edge stack; and updating an error log of the hardware-implemented union-find decoder based on the results of the iterative peeling decoding. In any of the foregoing examples or any other example, the decoding method additionally or alternatively includes analyzing d consecutive rounds of syndrome data together at the hardware-implemented union-find decoder in a 3D graph, where d is the code distance. In any of the foregoing examples or any other example, the decoding method additionally or alternatively includes using Union() and Find() graph operations at the Gr-Gen module to generate the spanning forest. In any of the foregoing examples or any other example, the decoding method additionally or alternatively includes storing newly grown edges in a merged edge stack at the Gr-Gen module. In any of the foregoing examples or any other example, the decoding method additionally or alternatively includes, at the DFS engine: generating a primary edge stack that includes a list of visited edges; generating a pending edge stack that includes a list of edges to be visited; and generating an alternative edge stack that is configured to hold the remaining edges of clusters from the spanning forest.

[0162] It will be understood that the configurations and / or methods described herein are exemplary in nature and these specific embodiments or examples should not be considered limiting as many variations are possible. The specific routines or methods described herein may represent one or more of any number of processing strategies. As such, the various acts shown and / or described may be performed in the order shown and / or described, in other orders, in parallel, or omitted. Likewise, the order of the above processes may be changed.

[0163] The subject matter of the present disclosure includes all novel and non-obvious combinations and sub-combinations of the various processes, systems, and configurations, as well as other features, functions, acts, and / or properties disclosed herein, and any and all equivalents thereof.

Claims

1. A quantum computing device, comprising: At least one quantum register, including a plurality of qubits; And A hardware decoder configured to: Receive syndrome data from one or more of the plurality of qubits; And Decode the received syndrome data by implementing a union-find decoding algorithm via a hardware microarchitecture including two or more pipeline stages, wherein the two or more pipeline stages include a graph generator Gr-Gen module, a depth-first search DFS engine, and a correction Corr engine.

2. The quantum computing device according to claim 1, wherein The graph generator Gr-Gen module is configured to: Generate a spanning forest by growing clusters around non-trivial adjoint bits; And Store data about the spanning forest in a spanning tree memory STM and a zero data register; The depth-first search DFS engine is configured to: Access the data stored in the STM; And Generate one or more edge stacks based on the data stored in the STM; and The correction Corr engine is configured to: Access the one or more edge stacks; Generate a memory request based on the accessed edge stack; and Update an error log based on the decoded syndrome data.

3. The quantum computing device according to claim 2, wherein the DFS engine is further configured to generate: A primary edge stack, including a list of visited edges, A pending edge stack, including a list of edges to be visited; and An alternative edge stack configured to hold the remaining edges of a cluster from the spanning forest.

4. The quantum computing device according to claim 1, wherein at least one of the two or more pipeline stages is coupled to two or more upstream pipeline stages via a multiplexer.

5. A quantum computing device, comprising: At least one quantum register, including a plurality of logical qubits; And A hardware decoder configured to receive syndrome data from one or more of the plurality of logical qubits and decode the received syndrome data, the hardware decoder including: A graph generator Gr-Gen module configured to: Generate a spanning forest by growing clusters around non-trivial syndrome bits; and Store data about the spanning forest in a spanning tree memory STM and a zero data register; A depth-first search DFS engine configured to: Access the data stored in the STM; and Generate one or more edge stacks based on the data stored in the STM; and A correction Corr engine configured to: Access one or more edge stacks; and Generate a memory request based on the accessed edge stack.

6. The quantum computing device according to claim 5, wherein the Corr engine is further configured to: Perform iterative peeling decoding on each accessed edge stack; and Update the error log of the hardware decoder based on the result of the iterative peeling decoding.

7. The quantum computing device according to claim 5, wherein the hardware decoder analyzes d consecutive received syndrome data together in a 3D graph, where d is the code distance.

8. The quantum computing device according to claim 5, wherein the Gr-Gen module is configured to generate the spanning forest using Union() and Find() graph operations.

9. The quantum computing device according to claim 5, wherein the Gr-Gen module further includes a fused edge stack configured to store newly grown edges.

10. The quantum computing device according to claim 9, wherein the STM is updated based on an indication received from the fused edge stack that two or more clusters have been merged.

11. The quantum computing device according to claim 5, wherein the DFS engine is configured to generate: a primary edge stack including a list of visited edges, a pending edge stack including a list of edges to be visited; and an alternative edge stack configured to hold the remaining edges from one or more clusters of the spanning forest.

12. The quantum computing device according to claim 5, wherein compressed adjoint data is routed from the compression engine to the Gr-Gen module.

13. The quantum computing device according to claim 12, wherein two or more compression engines are coupled to the hardware decoder.

14. A decoding method for a quantum computing device, comprising: receiving adjoint data from one or more qubits among a plurality of qubits; and decoding the adjoint data with a union-find decoder implemented in hardware including two or more pipeline stages, wherein the two or more pipeline stages include a graph generator Gr-Gen module, a depth-first search DFS engine, and a correction Corr engine.

15. The decoding method according to claim 14, wherein decoding the adjoint data with a union-find decoder implemented in hardware including two or more pipeline stages includes: at the graph generator Gr-Gen module: generating a spanning forest by growing clusters around non-trivial adjoint bits; and storing data about the spanning forest in a spanning tree memory STM and a zero data register; at the depth-first search DFS engine: accessing the data stored in the STM; and generating one or more edge stacks based on the data stored in the STM; and at the correction Corr engine: accessing one or more of the generated edge stacks; and generating a memory request based on the accessed edge stack.

16. The decoding method according to claim 15 further includes: At the Corr engine: performing iterative peeling decoding on each accessed edge stack; and updating an error log of the hardware-implemented union-find decoder based on the result of the iterative peeling decoding.

17. The decoding method according to claim 15 further comprises: Analyzing d consecutive rounds of adjoint data together in a 3D graph at the hardware-implemented union-find decoder, where d is the code distance.

18. The decoding method according to claim 15 further includes: At the Gr-Gen module, generating the spanning forest using Union() and Find() graph operations.

19. The decoding method according to claim 15 further includes: At the Gr-Gen module, storing newly grown edges at the fused edge stack.

20. The decoding method according to claim 15 further includes: At the DFS engine: generating a primary edge stack, the primary edge stack including a list of visited edges; generating a pending edge stack, the pending edge stack including a list of edges to be visited; and generate an alternative edge stack configured to store remaining edges from clusters of the spanning forest.

Citation Information

Patent Citations

  • Methods and devices for symbols detection in multi antenna systems

    CN107276716A

  • A multicast routing method based on fidelity measurement in a quantum communication network

    CN109714261A