Companion data compression for quantum computing devices
By designing a mechanism for compression and decompression of accompanying data in quantum computing devices, the problem of high error rate of quantum bits is solved, efficient data transmission and error correction are achieved, and the efficiency and reliability of the device are improved.
Patent Information
- Application Number
- CN202510151544.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-18
- Filing Date
- 2020-06-09
- Publication Date
- 2025-05-30
AI Technical Summary
In quantum computing devices, quantum bits are prone to high error rates, and the prior art is difficult to effectively solve this problem, especially in the process of data transmission and error correction.
A quantum computing device is designed, including at least one quantum register and a plurality of logic qubits, each logic qubit coupled with a compression engine and a decompression engine for compressing and decompressing accompanying data, thereby reducing bandwidth overhead and supporting high-throughput transmission.
By compressing the accompanying data, bandwidth overhead is reduced and high-throughput transmission of accompanying data from quantum registers is supported, improving the efficiency and reliability of quantum computing devices.
Smart Images

Figure CN120069102A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application is a divisional application of the patent application with application number 202080055436.4 and invention title "Adjoint Data Compression for Quantum Computing Devices". Background Art
[0003] Qubits are prone to high error rates and thus benefit from active error correction. Quantum error correction codes can be used to encode logical qubits into a set of physical qubits. Then, measurements can be used to detect errors and an error decoder can be used to correct the errors. Qubits typically operate at very low temperatures, and data is transferred to the error decoder at a higher operating temperature. Summary of the Invention
[0004] The present invention content is provided to introduce a selection of concepts that are further described below in the detailed description in a simplified form. The present invention content is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Additionally, the claimed subject matter is not limited to implementations that solve any or all of the disadvantages noted in any part of this disclosure.
[0005] A quantum computing device includes at least one quantum register that includes a plurality of logical qubits. A compression engine is coupled to each of the plurality of logical qubits. Each compression engine is configured to compress adjoint data. A decompression engine is coupled to each compression engine. Each decompression engine is configured to receive the compressed adjoint data, decompress the received compressed adjoint data, and route the decompressed adjoint data to a decoder block. This reduces the bandwidth overhead and supports high-throughput transmission of adjoint data from the quantum register. Brief Description of the Drawings
[0006] Figure 1 An example of a quantum computing organization is schematically shown.
[0007] Figure 2 Aspects of an example quantum computer are schematically shown.
[0008] Figure 3 A Bloch sphere is illustrated, which graphically represents the quantum state of a single qubit in a quantum computer.
[0009] Figure 4 Logical qubits in a lattice with alternating data qubits and parity qubits are schematically shown.
[0010] Figure 5 Two consecutive rounds of adjoint measurements are schematically shown.
[0011] Figure 6 It is a graph indicating the memory capacity required to store adjoint measurement data under varying conditions.
[0012] Figure 7 Schematically shows an example compression scheme.
[0013] Figure 8 Schematically shows multiple regions on a surface code lattice for geometry-based compression.
[0014] Figure 9 Shows an example method for compressing adjoint data within a quantum computing device.
[0015] Figure 10 Shows an example method for compressing adjoint data within a quantum computing device using a geometry-based compressor.
[0016] Figure 11 Schematically shows an example decoder design.
[0017] Figure 12 Schematically shows an example union-find decoder.
[0018] Figure 13 Schematically shows an example graph generator module.
[0019] Figure 14 Schematically shows the states of the main components in the graph generator module during graph generation.
[0020] Figure 15 Schematically shows an example cluster and root table entry.
[0021] Figure 16 Schematically shows a depth-first search engine and an example graph for error correction.
[0022] Figure 17 Shows an example of peeling for an example error graph performed in a correction engine.
[0023] Figure 18 Shows an example method for implementing a pipelined version of a hardware union-find decoder.
[0024] Figure 19 Schematically shows the baseline organization of L logical qubits.
[0025] Figure 20 Shows an example design of a decoder module.
[0026] Figure 21 Schematically shows an example microarchitecture of a scalable fault-tolerant quantum computer.
[0027] Figure 22 Shows an example method for routing adjoint data within a quantum computing device.
[0028] Figure 23 Schematically shows a Monte Carlo emulator framework.
[0029] Figure 24 Is a graph indicating the mean compression ratio of different error rates across code distances.
[0030] Figure 25 Is a graph indicating the average cluster diameter of different error rates and code distances of logical qubits.
[0031] Figure 26 Is a graph indicating the total memory capacity required to implement a spanning tree memory.
[0032] Figure 27 Is a graph indicating the distribution of the number of edges in a cluster of fixed code distance and physical error rate.
[0033] Figure 28 Is a graph showing the average number of edges in clusters of different code distances and error rates.
[0034] Figure 29 Is a graph indicating the correlation between the graph generator module and the execution time in a depth - first search engine.
[0035] Figure 30 Is a graph indicating the distribution of the execution time for decoding a 3D graph.
[0036] Figure 31 Shows a schematic diagram of an example classical computing device. Detailed Description
[0037] Qubits (the basic information units in a quantum computer) are prone to high error rates. To support fault - tolerant quantum computing, active error correction can be applied to these qubits. Quantum error - correcting codes (QECCs) use redundant data and parity qubits to encode logical qubits. Error correction is performed through a process called error decoding, which diagnoses errors on data qubits by analyzing the measurements of parity qubits. Currently, most decoding methods target qubit errors at the algorithmic level and do not take into account the underlying device technology used to design them.
[0038] In this paper, aiming at the architectural challenges involved in designing these decoders, a three-stage pipelined microarchitecture for the hardware implementation of the union-find decoder is described. The error correction algorithm is designed to be suitable for hardware implementation. Regarding the storage amount and bandwidth required for implementation, the feasibility of data compression for different noise mechanisms is evaluated. An architecture is disclosed that scales the proposed decoder design for a large number of logical qubits and supports practical fault-tolerant quantum computing. Such a design can reduce the total cost of each of the three pipeline stages by 2x, 4x, and 4x respectively through resource sharing across multiple logical qubits without affecting decoder correctness and error thresholds. As an example, for a code distance of 11 and a physical error rate on the order of 10 -3 the logical error rate is 10 -8 .
[0039] Quantum computing uses the properties of quantum mechanics to enable computations for specific applications that are infeasible to perform on conventional (i.e., non-quantum) state-of-the-art computers within a reasonable amount of time. Example applications include prime factorization, database search, and physical and chemical simulations. The basic computational unit on a quantum computer is the qubit. Qubits inevitably interact with the environment and lose their quantum states. Imperfect quantum gate operations exacerbate this problem since quantum gates are unitary transformations chosen from a continuous set of possible values and thus cannot be implemented with perfect accuracy. To protect quantum states from noise, QECCs have been developed. In any QECC, logical qubits are encoded with several physical qubits to support fault-tolerant quantum computing. Fault tolerance incurs a resource overhead (typically 20-100x per logical qubit) and is not practically feasible on currently available small prototypes. However, quantum error correction is considered highly valuable, if not entirely necessary, in order to run useful applications on fault-tolerant quantum computers.
[0040] QECCs differ from classical error correction techniques such as triple modular redundancy (TMR) or single error correction double error detection (SECDED). These differences stem from the fundamental properties of qubits and their high error rates (typically on the order of 10 -2 ). For example, qubits cannot be copied (no-cloning theorem) and lose their quantum states when measured. QECCs use redundant qubits to create an encoded space by using ancilla qubits that interact with the data qubits. By measuring the ancilla qubits, a decoder can be used to detect and correct errors on the data qubits.
[0041] The error decoding algorithm specifies how the syndrome (the result of the ancillary measurement) will be processed to detect errors in the encoded blocks of data qubits. The design and performance of the decoder depend on the decoding algorithm, QECC, physical error rate, noise model, and implementation technology. For practical purposes, the rate at which the decoder processes the syndrome measurements must be faster than the rate at which errors occur. They must also take into account the specific technical constraints for operation in cryogenic environments and scale to a large number of qubits.
[0042] Perfect error decoding is NP-hard (non-deterministic polynomial-time) with exponential time complexity. Thus, the best decoding algorithms trade off error correction capabilities to reduce the time complexity. Although decoders are groundbreaking for fault-tolerant quantum computing, most decoding techniques have been studied only at the algorithmic level and do not consider the underlying implementation technology. Other methods such as lookup table-based decoders or deep neural decoders do not scale to a large number of qubits. The Union-Find decoder algorithm is simple and has near-linear time complexity, making it a suitable candidate for scalable fault-tolerant quantum computing. Here, a microarchitecture for the hardware implementation of the Union-Find decoder is disclosed, where the algorithm is redesigned to reduce the hardware complexity and allow scaling to a large number of logical qubits.
[0043] To support faster processing and reduce the communication latency, the decoder is designed to operate at a temperature very close to that of the physical qubits (77K or 4K) rather than at room temperature (300K). Figure 1 An example quantum computing system 100 indicating the temperature gradient 105 is shown. The quantum computing system 100 includes one or more qubit registers 110 operating at 20mK, a decoder 115 and a controller 120 typically operating at 4K or 77K, and a host computing device 125 and an end-user computing device 130 operating at 300K (room temperature). Depending on whether the decoder 115 operates at 77K or 4K, the underlying implementation technology and design constraints provide different trade-offs. Superconducting logic design at 4K is very close to the physical qubits and has significant energy efficiency, but is limited by device density and memory capacity. Conventional CMOS operating at 77K can drive complex designs for applications with larger memory footprints, but its energy efficiency is lower than that of superconducting logic, and moving data back and forth from the physical qubits residing at 15 - 20mK incurs data transmission overhead.
[0044] In this paper, a microarchitecture for the hardware implementation of a union-find decoder is disclosed. Implementation challenges associated with memory capacity and bandwidth for operation in cryogenic environments are discussed. Surface codes are used as an example for underlying QECC and various noise models, although other implementations have been considered. Surface codes are a promising QECC candidate that arranges a set of qubits in a two-dimensional layout with alternating data qubits and ancilla qubits. Any error in a data qubit can be detected by its neighboring ancilla qubits, thus only requiring nearest-neighbor connections. The feasibility and scalability of such a design for large-scale fault-tolerant quantum computers are described.
[0045] Here, systems and methods for solving many problems in the field of quantum computing are disclosed. For example, the design of QEC decoders is analyzed, including their placement and design complexity in the thermal domain. A microarchitecture for the hardware implementation of a union-find decoder is proposed, which proves to be more practical to operate the decoder at 77K.
[0046] The memory capacity required to store syndrome measurements is calculated, and it is shown that storing them in superconducting memory at 4K may be infeasible. However, transferring data to 77K requires a large bandwidth. To overcome these two challenges, techniques that can be used to compress syndrome measurement data are proposed. The implementation of dynamic zero compression and sparse representation is described. Additionally, a geometry-based compression scheme is proposed, which takes into account the underlying structure of the surface code lattice. Additionally, compression schemes and their applicability are described for different noise mechanisms.
[0047] The union-find decoder algorithm is refined in order to reduce hardware costs and for enhanced implementation in noise models. The original union-find decoding algorithm only considered gate errors on paired data qubits in space. The hardware microarchitecture described in this paper also considers measurement errors, which are paired in time, and decodes them using several rounds of syndrome measurements.
[0048] Additionally, the hardware system architecture for scaling these decoders for a large number of logical qubits is described. Such an implementation can consider the utilization differences of pipeline stages in individual decoding units and support optimized resource sharing between multiple logical qubits to reduce hardware costs.
[0049] A qubit is the basic unit of information on a quantum computer. The foundation of quantum computing relies on two quantum mechanical properties: superposition and entanglement. A qubit can be represented as a linear combination of its two basis states. If the basis states are |0> and |1>, then the qubit |ψ> can be represented as |ψ> = α|0> + β|1>, where and |α| 2 + |β| 2= 1. When the magnitudes and / or phases of the probability amplitudes α and β change, the state of the qubit changes. For example, an amplitude flip (or, bit flip) changes the state of |ψ> to β|0> + α|1>. Alternatively, a phase flip changes its state to α|0> - β|1>. Quantum instructions use quantum gate operations to modify the probability amplitudes, and quantum gate operations are represented using the identity matrix (I) and the Pauli matrices. The Pauli matrices X, Z, and Y represent the effects of bit flips, phase flips, or both, respectively.
[0050] In some embodiments, the methods and processes described herein may be associated with a quantum computing system of one or more quantum computing devices. Figure 2 Aspects of an example quantum computer 210 configured to perform quantum logic operations are shown (see below). Conventional computer memory stores digital data in arrays of bits and performs bit-by-bit logical operations, while a quantum computer stores data in an array of qubits and performs quantum mechanical operations on the qubits to achieve the desired logic. Accordingly, Figure 2 The quantum computer 210 includes at least one register 212 that includes an array of qubits 214. The illustrated register has a length of eight qubits; registers including longer and shorter arrays of qubits are also contemplated, as are quantum computers including two or more registers of any length.
[0051] The qubits of register 212 may take various forms depending on the desired architecture of quantum computer 210. As a non-limiting example, each qubit 214 may include: a superconducting Josephson junction, a trapped ion, a trapped atom coupled to a high-finesse cavity, an atom or molecule confined within a fullerene, an ion or neutral dopant atom confined within a host lattice, a quantum dot presenting discrete spatial or spin states, an electron-hole in a semiconductor junction trapped by an electrostatic well, a pair of coupled quantum wires, a nucleus accessible by magnetic resonance, a free electron in helium, a molecular magnet, or a metallike carbon nanosphere. More generally, each qubit 214 may include any particle or system of particles that can exist in two or more discrete quantum states that can be experimentally measured and manipulated. For example, qubits may also be implemented in multiple processing states corresponding to different light propagation modes through linear optical elements (e.g., mirrors, beam splitters, and phase shifters), as well as in states accumulated within a Bose-Einstein condensate.
[0052] Figure 3is a diagram of the Bloch sphere 216, which provides a graphical description of some of the quantum mechanical aspects of the individual qubits 214. In this description, the north and south poles of the Bloch sphere correspond to the standard basis vectors |0> and |1> respectively - for example, the up-spin and down-spin states of an electron or other fermion. The set of points on the surface of the Bloch sphere includes all possible pure states |ψ> of the qubit, while the points inside correspond to all possible mixed states. The mixed state of a given qubit may be caused by decoherence, which may occur due to an unwanted coupling to external degrees of freedom.
[0053] Now returning to Figure 2 , the quantum computer 210 includes a controller 218. The controller may include conventional electronic components, including at least one processor 220 and an associated storage machine 222. The term "conventional" as used herein applies to any component that can be modeled as a collection of particles, regardless of the quantum state of any individual particle. For example, conventional electronic components include integrated microfabricated transistors, resistors, and capacitors. The storage machine 222 may be configured to hold program instructions 224 that cause the processor 220 to perform any of the processes described herein. Additional aspects of the controller 218 are described below.
[0054] The controller 218 of the quantum computer 210 is configured to receive a plurality of inputs 226 and provide a plurality of outputs 228. The inputs and outputs may each include digital and / or analog lines. At least some of the inputs and outputs may be data lines through which data is provided to and extracted from the quantum computer. Other inputs may include control lines via which the operation of the quantum computer can be adjusted or otherwise controlled.
[0055] The controller 218 is operably coupled to the register 212 via an interface 230. The interface is configured to exchange data bi-directionally with the controller. The interface is also configured to exchange signals corresponding to the data bi-directionally with the register. Depending on the architecture of the quantum computer 210, such signals may include electrical, magnetic, and / or optical signals. Via the signals transmitted through the interface, the controller can interrogate or otherwise affect the quantum states stored in the register, as defined by the collective quantum state of the qubit array 214. To this end, the interface includes at least one modulator 232 and at least one demodulator 234, each modulator and each demodulator being operably coupled to one or more qubits of the register 212. Each modulator is configured to output a signal to the register based on modulation data received from the controller. Each demodulator is configured to sense a signal from the register and output data to the controller based on that signal. In some scenarios, the data received from the demodulator may be an estimate of the observable value of a measurement of the quantum state stored in the register.
[0056] More specifically, a suitably configured signal from the modulator 232 can physically interact with one or more qubits 214 of the register 212 to trigger a measurement of the quantum state stored in the one or more qubits. The demodulator 234 can then sense the resulting signal released by the one or more qubits based on the measurement and can provide data corresponding to the resulting signal to the controller. In other words, the demodulator can be configured to reveal an estimate of an observable value that reflects the quantum state of one or more qubits of the register based on the received signal and provide the estimate to the controller 218. In one non-limiting example, the modulator can provide a suitable voltage pulse or sequence of pulses to the electrodes of one or more qubits based on data from the controller to initiate the measurement. In short, the demodulator can sense photon emissions from one or more qubits and can assert a corresponding digital voltage level on the interface line entering the controller. Generally, any measurement of a quantum mechanical state is defined by an operator corresponding to the observable to be measured; the result R of the measurement is guaranteed to be one of the allowed eigenvalues. In the quantum computer 210, R is statistically related to the register state prior to the measurement but is not uniquely determined by that register state.
[0057] According to appropriate inputs from the controller 218, the interface 230 can also be configured to implement one or more quantum logic gates to operate on the quantum states stored in the register 212. The function of each logic gate in a conventional computer system is described according to a corresponding truth table, while the function of each quantum gate is described by a corresponding operator matrix. The operator matrix operates (i.e., multiplies) on the complex vector representing the register state and implements a specific rotation of that vector in Hilbert space.
[0058] Continuing Figure 2 , a suitably configured signal from the modulator 232 of the interface 230 can physically interact with one or more qubits 214 of the register 212 in order to assert any desired quantum gate operation. As described above, the desired quantum gate operation is a specifically defined rotation of the complex vector representing the register state. To achieve the desired rotation one or more modulators of the interface 230 can apply a predetermined signal level S i for a predetermined duration T i .
[0059] In some examples, multiple signal levels can be applied to multiple sequences or other associated durations. In more specific examples, multiple signal levels and durations are arranged to form a composite signal waveform that can be applied to one or more qubits of a register. Typically, each signal level S i and each duration T i are control parameters that can be adjusted through appropriate programming of the controller 218. In other quantum computing architectures, different sets of adjustable control parameters can control the quantum operations applied to the register states.
[0060] Qubits inevitably lose their quantum states through interactions with different degrees of freedom in their surroundings. Even if qubits can be perfectly isolated from environmental noise, quantum gate operations are imperfect and cannot be applied with exact accuracy. This poses various limitations to running any application on a quantum computer. Therefore, the quantum states manipulated by a quantum computer must be error-corrected using quantum error-correcting codes (QECCs). QECCs encode logical qubits into a collection of physical qubits such that the error rate of the logical qubits is lower than the physical error rate. As long as the physical error rate is below an acceptable threshold, QECCs can support fault-tolerant quantum computing at the cost of an increased number of physical qubits. In recent years, several error-correction protocols have been proposed. The surface code is applied in this article and is considered to be the most promising QECC for fault-tolerant quantum computing. QEC models arbitrary noise as a superposition of quantum operations. Therefore, QECCs use Pauli matrices to capture the effects of errors as bit flips, phase flips, or a combination of both.
[0061] The surface code is widely considered suitable for scalable fault-tolerant quantum computing. It encodes logical qubits in a lattice using alternating data qubits and parity qubits. Figure 4 A schematic representation of such a lattice is shown at 400, where the lattice 400 has a 2D distance-3 (d = 3) surface code. Each X stabilizer 402 and Y stabilizer 404 is coupled to its adjacent data qubits 406. Each data qubit 406 interacts only with its nearest-neighbor parity qubits 408, such that errors on the data qubit 406 can be diagnosed by measuring the locally supported operators, as shown at 410. In this example, a Z error 412 on data qubit A 414 is captured by parity qubits P0 416 and P1 418. Similarly, an X error 420 on data qubit B 422 is captured by parity qubits P2 424 and P3 426. In the simplest implementation, a surface code of distance d uses (2d - 1) 2 physical qubits to store a single logical qubit, where d is a measure of redundancy and fault tolerance. Larger code distances result in greater redundancy and increased fault tolerance.
[0062] Logical operators consist of a string of single - qubit operators between two opposite edges. The encoded space is the subspace where all stabilizer generators (as shown in Figure 4 ) have an eigenvalue of +1. By construction, logical states are invariant under the application of stabilizer generators. Any closed loop of Pauli operator applications will leave the logical state unchanged. Measurement of stabilizer generators detects the endpoints of a series of errors. Error correction is based on this information, and stabilizer measurement is called syndromic.
[0063] In QEC, the effect of an error is reversed by applying appropriate Pauli gates. For example, if a qubit encounters a bit - flip error, applying a Pauli X gate flips it back to the expected state. It has been shown previously that as long as Clifford gates are applied to qubits, no active error correction needs to be performed. Instead, it is sufficient to keep track of the Pauli frame in software. Thus, the main focus of quantum error correction is error decoding rather than error correction. Optimal error decoding is a computationally difficult problem. A quantum error decoder takes syndromic measurements as input and returns an estimate of the error in the data qubits. In addition to the ability to detect errors, the decoder also relies on high operating speed to prevent the accumulation of errors. In other words, errors must be detected faster than they occur.
[0064] Since error decoding must be fast, the decoder must provide high performance and operate near the physical qubits. As described regarding Figure 1 , qubits typically operate at 15 - 20 mK. Depending on whether the decoder is designed to work at 4K or 77K, the implementation technologies offer different trade - offs, as shown in Table 1. Due to proximity to the physical qubits, hardware designed to work at 4K must meet strict power requirements. This is to ensure that thermal noise is controlled. Additionally, these designs must be cooled using complex and expensive liquid - helium coolers. Decoders can be designed using CMOS or superconducting logic at 4K. CMOS has power consumption and thus cannot be used for large - scale quantum computers. Superconducting logic offers low power consumption, but has major drawbacks such as limited device density and low storage capacity, which make it extremely difficult to fabricate complex and large designs. Conventional CMOS operating at 77K offers the ability to design complex systems with larger memory and power budgets. The cooling overhead associated with 77K is an order of magnitude lower than that of 4K. However, a decoder designed to work at 77K must account for transmission delays and meet the bandwidth requirements for data transfer back and forth between 4K and 77K. The trade - offs between superconducting technology at 4K and CMOS at 77K are listed in Table 1.
[0065] Parameters Superconducting Technology Traditional CMOS Operating Temperature 4K 77K Operating Frequency 10 GHz 4 GHz Memory Capacity 123 - 512 bytes 4 Gb Power Budget 1W N / A Feature Size 248 nm 7 - 16 nm Cooling Overhead 1000x / 400x 10x
[0066] Table 1
[0067] In this paper, the challenges in designing the microarchitecture of decoders for quantum error correction under realistic noise models are examined. Qubit errors can be broadly classified into three types: decoherence errors, gate errors, and measurement errors. Qubits maintain their quantum states only for a short duration (referred to as the decoherence time), leading to decoherence errors. Non-ideal gate operations result in gate errors on qubits. Imperfections in qubit measurements lead to measurement errors. The decoder may misinterpret the accompanying measurement errors as data qubit errors and correct the non-error data qubits, thus introducing errors. The decoder must consider such accompanying measurement errors while decoding errors. This directly affects the microarchitecture and design choices of the decoder.
[0068] Figure 5 Diagram 500 is shown, which indicates two consecutive rounds of syndrome measurements 502 and 504, and shows how measurement errors 506 pair up in time and data qubit errors 508 pair up in space. Diagram 500 shows that if the decoder only examines the measurement results of the 0th round 502, it will misinterpret the error on the parity qubit P0 510 and force a correction on the error-free D0 512. Current decoders address accompanying measurement errors by examining d rounds of measurements, where d is the code distance. The data generated by d rounds of syndrome measurements and the error log per data qubit must be stored for the decoder to work correctly. This requires storage space of up to several megabytes (depending on the code distance and the number of logical qubits).
[0069] Figure 6 Example graph 600 is shown indicating the memory capacity (in KB) required to store the syndrome measurement data for d (code distance) rounds and the error logs for N logical qubits. The required capacity is much higher than the available memory in superconducting logic at 4K. To perform error decoding at 77K, the measurement data must be transferred from 4K to 77K. For a given qubit plane with L logical qubits and each qubit encoded using a surface code with code distance d, 2d(d - 1)L bits must be sent at the end of each syndrome measurement cycle. Assuming a reasonable number of logical qubits and code distance, the 4K - 77K link requires a bandwidth in the order of several Gb / s. Data transfer at lower bandwidths reduces the remaining effective time for error decoding, as it must provide an estimate of the errors within d syndrome measurement cycles (e.g., a surface code cycle can be broken down into d syndrome measurement cycles). Thus, the main challenge in designing any decoder at 77K is the very large bandwidth required.
[0070] One way to effectively handle the capacity and bandwidth requirements in the cache and main memory is data compression. The sparsity of the measured data can be analyzed and estimated as described herein. For example, let p be the probability of a Z error on a data qubit, and let u be the error indicator vector for n data qubits (note that the same analysis also applies to the X adjoint). If there are 4 data qubits and the first two have Z errors, then u = 1100. Assuming an identical and independent error distribution (iid), the upper bound of the adjoint Hamming weight is given by Equation (1), where |u| is the Hamming weight of the error indicator vector u (e.g., the number of 1s).
[0071]
[0072] Therefore, the probability of having m or more errors is given by Equation 2:
[0073]
[0074] Using the union - bound, the upper bound of the total number of adjoint bits s(Z u ) is given by Equation 3.
[0075] s(Z u ) ≤ 2|u| (Equation 3)
[0076] Assume the code distance is 11 and the error rate is 10 -3 , the probability of having 10 or more errors (for a given code distance, the number of errors is quite large) is 6.2x10 -14 . Therefore, the probability of observing an adjoint with a large Hamming weight is extremely low. This analysis indicates that it is possible to compress adjoint data to reduce the storage overhead for storage and / or meet the bandwidth requirements. Different compression techniques for adjoint data are described herein, as the usefulness of the compression technique depends on the entropy of the data. In this article, three compression techniques are described, although other techniques have been considered. Different noise mechanisms to which these techniques can be applied are analyzed. The examples described have compression schemes that use simple coding and do not require high hardware complexity respectively.
[0077] Dynamic zero compression (DZC) was originally introduced to reduce the energy required for cache access of zero - valued bytes. A similar technique can be adopted to compress adjoint data. Figure 7An example is shown at 700. The syndrome of length L is grouped into K blocks each of W bits, where W is the compression width 710. If the number of bits in the last block is less than W bits, additional padding zeros can be added. The K-bit wide zero indicator bit (ZIB) vector 715 includes 1 bit per block. If all bits of the i-th block are 0, the corresponding bit in the ZIB (ZIB[i]) can be set to 1. Otherwise, the bit can be set to 0. The data to be transmitted 720 can be obtained by appending the non-zero blocks 725 at the end of the ZIB vector.
[0078] As shown at 750, the sparse representation can be considered similar to traditional techniques for storing sparse matrices, where the non-zero elements of the sparse matrix 760 are stored by storing only the row and column indices 765. The sparse representation bit (SRB) 755 is used to indicate whether all syndrome bits are zero. If there is one or more non-zero bits in the syndrome, the SRB can be left unset, and the indices of the non-zero elements 755 can be sent together with the SRB in the transmitted data 775.
[0079] Geometry-based compression (Geo-Comp) can be considered an adaptation of DZC that also takes into account the geometry of the surface code lattice. The geometry-based compression scheme can compress regions of the X and Z syndromes together, rather than compressing the X and Z syndromes separately. The entire surface code lattice can be divided into multiple regions, where each region roughly contains an equal number of syndrome bits (similar to the compression width of DZC). Figure 8 Schematically shown is a surface code lattice 800 for geometry-based compression that includes multiple regions (801, 802, 803, 804; represented by dashed lines). Figure 8 An example is shown of how a surface code lattice with a code distance of 5 is divided into 4 regions. Using ZIB for each region and transmitting only the syndrome data from non-zero regions can compress the syndrome. When a Y error occurs on a data qubit, the X and Z syndrome bits flip to indicate the error. When the two types of syndromes are compressed independently, for a given compression width, the total number of non-zero blocks is higher. For example, if Figure 8 the data qubit D0 810 shown in encounters a Y error, the X syndrome bits X0 811 and X1 812 and the Z syndrome bits Z0 813 and Z1 814 flip. In a compression scheme such as DZC, (X0 811,X1 812) and (Z0813,Z1 814) are on different data blocks and are compressed separately. However, if the geometry of the lattice is considered, the non-zero syndrome bits are typically within the same region, unless the data qubit is on a region boundary (e.g., D1815 in the lattice 800).
[0080] In general, for a given noise model, the number and size of regions can be adjusted by calculating the expected number of blocks that contain the trivial (all-zero) syndrome. However, larger-sized regions lead to complex hardware by increasing the logical depth. Thus, even for very low error rates, smaller region sizes (depending on the code distance) can be analyzed. The regions do not need to be of equal size, and the size and / or number of regions can be determined based on the expected number of data blocks that contain the trivial syndrome.
[0081] Figure 9 An example method 900 for compressing syndrome data within a quantum computing device is shown. In some examples, method 900 can be implemented by a quantum computing device that includes a union-find decoder (e.g., Figure 12 the decoder schematically depicted in).
[0082] At 910, method 900 includes generating syndrome data from at least one quantum register that includes l logical qubits, where l is a positive integer. The generated syndrome data can include at least X syndrome data and Z syndrome data.
[0083] Continuing at 920, method 900 includes, for each logical qubit: routing the generated syndrome data to a compression engine that is configured to compress the syndrome data. The quantum computing device can include multiple compression engines. In some examples, at least one compression engine is configured to compress the syndrome data using dynamic zero compression. In some examples, at least one compression engine is configured to compress the syndrome data using a sparse representation. In some examples, at least one compression engine is configured to compress the syndrome data using geometry-based compression. The quantum computing device can include two or more logical qubit sectors that are coupled to two or more types of compression engines. In some examples, method 900 can include operating the compression engine at 4K. However, higher (e.g., 8K) or lower (e.g., 2K) temperatures can be used.
[0084] Continuing at 930, method 900 includes routing the compressed syndrome data to a decompression engine that is configured to: receive the compressed syndrome data; and decompress the received compressed syndrome data. At 940, method 900 includes routing the decompressed syndrome data to a decoder block. In some examples, the decompressed syndrome data can be routed to a graph generator module of the decoder block. In some examples, method 900 can include operating the decompression engine and / or the decoder block at 77K. However, higher (e.g., 85K) or lower (e.g., 70K) temperatures can be used. In some examples, the quantum computing device includes a set of d decoder blocks, where d < 2*l.
[0085] Figure 10 Illustrates an example method 1000 for compressing syndrome data using geometry-based compression within a quantum computing device. In some examples, method 1000 may be implemented by a hardware including a union-find decoder (e.g., the decoder schematically depicted in Figure 12 ) of a quantum computing device.
[0086] At 1010, method 1000 includes generating syndrome data from at least one surface code lattice including l logical qubits, where l is a positive integer. For example, the surface code lattice is divided into two or more regions based on the lattice geometry, as shown in Figure 8 . In some examples, the number of regions may be determined based on the expected number of data blocks containing trivial syndromes.
[0087] At 1020, method 1000 includes, for each logical qubit: routing the generated syndrome data to a compression engine configured to compress the syndrome data using geometry-based compression. At 1030, method 1000 includes compressing the syndrome data using zero indicator bits for each of two or more regions of the surface code lattice. At 1040, method 1000 includes transmitting the syndrome data only from non-zero regions. In other words, it may be assumed that if no data is received from a region, that region includes only trivial (e.g., all-zero) data.
[0088] At 1050, method 1000 includes routing the compressed syndrome data to a decompression engine configured to: receive the compressed syndrome data; and decompress the received compressed syndrome data. The decompression engine may be programmed based on the geometry-based compression scheme used by the compression engine.
[0089] The decoder of QEC is used to process syndrome measurement data and identify errors in the corrupted data qubits. Here, the microarchitecture of the hardware implementation of the union-find decoder for the surface code is improved. In the surface code, local operators on the lattice of qubits are measured and the syndrome is processed using the decoder to generate an estimate of the most likely error on the data qubits. The decoder microarchitecture is designed to prevent the accumulation of errors while maintaining a low hardware complexity to meet the strict power budget for operation in a cryogenic environment. The architecture proposed here is designed to support scaling up to thousands of logical qubits to support fault-tolerant quantum computing.
[0090] Quantum error decoding is an NP-hard problem. Thus, most decoding algorithms trade off between error thresholds and lower time and algorithmic complexity. A promising error decoding technique is the graph-based minimum weight perfect matching (MWPM) decoder. Although the MWPM decoder offers a high error threshold, it suffers from a high time complexity (O(n 2 ))). Alternatively, a simple way to design a decoder is based on using a lookup table. The table is indexed by syndrome bits, and the corresponding entries store the error information of the data qubits. However, the lookup table decoder is not scalable and requires terabytes of memory even for small code distances. Deep neural decoders are popular and learn the probability density function of the possible errors corresponding to the measured syndrome sequences during the training phase. Using inference, the error pattern for a given syndrome is evaluated. However, neural decoders require more hardware for computation and are not scalable as the code distance increases. The recently proposed union-find decoder presents an algorithm that forms clusters around non-trivial syndromes (non-zero syndromes) and performs error correction in almost linear time using graph traversal. The union-find decoder thus offers simplicity, time complexity, and a high error threshold.
[0091] The operation of the union-find decoder is as Figure 11 shown. At 1100, each edge on graph 1102 represents a data qubit, and each vertex represents a parity qubit (e.g., 1104, 1106). Decoding begins by growing a spanning forest 1108 to cover all error syndrome bits, thereby forming one or more even clusters, as shown at 1110. Data qubits A 1112 and B 1114 can be assigned unknown Pauli errors 1116 and 1118, respectively. By traversing the forest, errors can be detected, as shown at 1120. The cluster traversal steps (shown at 1122, 1124) can be used to detect, classify (e.g., Z errors), and correct the errors.
[0092] As Figure 12 shown in the block diagram 1200 of, an adaptation of the algorithm can be implemented. The compressed syndrome data 1210 is routed to a decompression engine 1215. The decompressed syndrome data 1220 is then routed to a graph generator (Gr-Gen) module 1225. The Gr-Gen module 1225 can be configured to generate spanning tree memory (STM) data. A depth-first search engine (DFS) 1230 can be configured to access the STM data and generate an edge stack based on the STM data. A correction (Corr) engine 1235 can be configured to access the edge stack, generate memory requests based on the accessed edge stack, and update an error log 1240.
[0093] If adjoint measurement errors are ignored, decoding is performed using 2D graphs generated from single-round adjoint measurements. To account for erroneous measurements, d consecutive adjoint measurements, where d is the code distance, must be decoded together, resulting in 3D graphs. The union-find decoder can be used for both cases. The main difference is that the required memory grows quadratically (for 2D) or cubically (for 3D) with the code distance of the surface code. For simplicity, the microarchitecture of the union-find decoder is described in 2D and generalized to 3D. All relevant results described are obtained for 3D graphs. The decoding design includes 3 pipeline stages, enabling improved design scalability.
[0094] The Gr-Gen module takes the adjoint as input after decompression and generates a spanning forest by growing clusters around non-trivial adjoint bits (non-zero adjoint bits). The spanning forest can be constructed using two basic graph operations: Union() and Find(). Figure 13 An example Gr-Gen module 1300 is schematically shown. Module 1300 includes a spanning tree memory (STM) 1310, a zero data register (ZDR) 1315, a root table 1320, a size table 1325, a parity register 1330, and a fused edge stack 1335. This design is slightly different from the union-find algorithm described previously for reducing hardware resource costs. The size of each component is a function of the code distance d. STM 1310 stores 1 bit for each vertex and 2 bits for each edge. Each edge uses 2 bits because, according to the original algorithm, clusters grow half the edge width around vertices or existing cluster boundaries. ZDR 1315 stores 1 bit per STM row. If the content of the row is 0, the bit stores 0, and if at least one bit in the row is 1, the corresponding ZDR bit for the row stores 1. Since adjoint data is sparse and the total number of edges in the spanning forest will be low, ZDR 1315 accelerates the traversal of STM 1310. FES1335 stores newly grown edges so that they can be added to existing clusters. The root table 1320 and size table 1325 store the roots and sizes of the clusters, respectively. The tree traversal register 1340 stores the vertices of each cluster accessed during the Find() operation. The interface 1345 between the Gr-Gen module and the DFS engine can allow the DFS engine to access the data stored at STM 1310.
[0095] As Figure 14As shown, the root table entry (root table[i]) is initialized to the index (i). As shown at 1400, the size table entry for non-trivial adjoint bits is initialized to 1. These tables assist the Union() and Find() operations in merging clusters into the final state shown at 1420 after the growth phase shown at 1410. They are indexed by cluster index. The size of the tables is set for the maximum possible number of clusters, which is equal to the total number of vertices in the surface code lattice. A boundary list for each cluster can be stored. However, in the noise mechanisms relevant to practical applications, the average cluster diameter is very small. The cluster diameter can be defined as the maximum distance between two vertices on the cluster boundary. Thus, instead of storing the boundary list, the boundary index can be calculated during the cluster growth phase. The original algorithm grows all odd clusters until the parity is even. Thus, odd clusters must be detected quickly. To do this, a parity register can be used, as Figure 11 shown. The parity register can store 1 bit of parity per cluster, depending on whether it is odd or even. For a reasonable code distance of 11, seven 32-bit registers can be sufficient. For larger code distances, additional parity information can be stored in memory and prefetched to hide the memory latency.
[0096] The control logic can read the parity register and grow the clusters with odd parity (referred to as the growth phase) by writing to the STM, ZDR, and adding newly added edges that touch the boundaries of other clusters to the FES. For edges connecting to other clusters, the STM may not be updated to prevent double growth. It can be updated when clusters are merged by reading from the FES. This logic can check whether the newly added edge connects two clusters by reading the root table entries of the vertices connected by the edge (call these the primary vertices). This is equivalent to the Find() operation. As Figure 15 shown at 1500 in, the vertices visited on the path to find the root of each primary vertex are stored in the tree traversal register. As shown at 1510, the root table entries of these vertices can be updated to point directly to the root of the cluster to minimize the depth of the tree for future traversals. This operation (referred to as path compression) is included in the union-find algorithm and allows the depth of the tree to be kept short, thus amortizing the cost of the Find() operation. For example, at 1500, Figure 15 shows the state of two clusters and the root table at a certain moment. Assume that after a growth step, vertices 0 and 6 are connected and the two clusters must be merged. The tree traversal register can be used to update the root of vertex 0, as shown at 1500. Since the depth of the tree is continuously compressed, only a few registers are sufficient. In one example, 5 registers are used per primary vertex, although more or fewer registers can also be used. If the primary vertices belong to different clusters, the root of the smaller cluster can be updated to point to the root of the larger cluster.
[0097] The DFS engine can process STM data generated by Gr-Gen that stores a set of even clusters growing. It can use the DFS algorithm to generate a list of edges forming a spanning tree for each cluster in the STM. In other examples, breadth-first search exploration can be used, although DFS is generally more memory-efficient. At Figure 16 An example DFS engine is shown at 1600 of. This logic can be implemented using a finite state machine 1610 and two stacks 1620 and 1622. Stacks can be used because the order of visiting edges in the spanning tree can be reversed to perform correction by peeling. The edge stack 1620 can store a list of visited edges, while the pending edge stack 1622 can store edges that will be visited later in the ongoing DFS. For example, as Figure 16 shown at 1630 of, when the FSM visits vertex 1 of the spanning tree, edge a is pushed onto the edge stack, and edge c is pushed onto the pending edge stack. When the end of the current path is reached, the pending edge can be popped and traversed. To support pipelined operation and improve performance, the microarchitecture can be designed to include an alternative edge stack 1632. When there is more than one cluster, the correction engine can work on the edge list of another cluster being traversed while the DFS engine traverses one cluster via the Corr engine interface 1640. As shown at 1630, if edges a, b, c, and d belong to cluster C0, and edges e and f belong to cluster C1, then when the Corr engine processes the correction of C0, the DFS engine 1600 can traverse C1. This can help set the size of the stack to handle the average cluster size rather than the worst-case cluster size. In the case where the DFS engine 1600 encounters a large enough cluster that cannot fit in one stack, the alternative stack 1632 can be used, and an overflow bit can be set to indicate that stacks 1620 and 1632 hold edges corresponding to a single cluster. This proposed implementation can include a number of memory reads proportional to the size of the cluster. By checking the STM1310 line by line, the effective cost of generating clusters is reduced. ZDR 1315 reduces the cost of traversing the STM 1310 line by line.
[0098] The Corr engine can perform the peeling process of the decoder and can identify the Pauli corrections to be applied. The Corr engine can access the edge list (stored on the stack) and the syndrome bits corresponding to the vertices along the edge list. The syndrome bits can be accessed by decompressing the compressed syndrome and / or by accessing the STM. However, the former may increase the logic complexity and latency, while the latter may increase the number of memory requests that the STM needs to handle. To reduce the memory traffic and eliminate the need for additional decompression logic, the syndrome information can be stored by the DFS engine together with the edge index information. The temporary syndrome changes caused by peeling are stored in local registers. An example of the peeling of an error graph performed in the Corr engine is asFigure 17 as shown in Figure 17 Figure 17 shows examples of step 1 1700, step 2 1710, and step 3 1720 along with examples of save registers, error logs, edge stacks, and error graphs. The Corr engine can also read the last surface code cycle error log and can update the Pauli corrections for the current edges. For example, if the error on edge e0 was Z in the previous logical cycle and it also encounters a Z error in the current cycle, the Pauli error for e0 can be updated to I, as shown at 1720.
[0099] Figure 18 Figure 18 shows an example decoding method 1800 for a quantum computing device. In some examples, the decoding method 1800 can be implemented by a quantum computing device including a union-find decoder (such as, Figure 12 the decoder schematically depicted in
[0100] At 1805, method 1800 includes receiving syndrome data from one or more qubits (such as logical qubits residing in a quantum register). The received syndrome data can include X syndrome data and / or Z syndrome data.
[0101] At 1810, method 1800 includes decoding the received syndrome data with a union-find decoder implemented in hardware including two or more pipeline stages. As an example, this can include Figure 12 a union-find decoder implemented in hardware including three pipeline stages as shown in
[0102] Optionally, at 1820, decoding the syndrome data can include generating a spanning forest in the Gr-Gen module by growing clusters around non-trivial syndrome bits. In some examples, the Union() and Find() graph operations can be used to generate a spanning tree.
[0103] Optionally, at 1825, decoding the syndrome data can include storing data about the spanning forest in a spanning tree memory (STM) and a zero data register at the Gr-Gen module. In some examples, newly grown edges can be stored at the fused edge stack.
[0104] Optionally, at 1830, decoding the syndrome data can include accessing the data stored in the STM at the DFS engine. Optionally, at 1835, decoding the syndrome data can include generating one or more edge stacks at the DFS engine based on the data stored in the STM. For example, as Figure 16As shown, generating one or more edge stacks based on data stored in the STM may include generating a primary edge stack that includes a list of edges visited. Additionally or alternatively, generating one or more edge stacks based on data stored in the STM may include generating a pending edge stack that includes a list of edges to be visited. Additionally or alternatively, generating one or more edge stacks based on data stored in the STM may include generating an alternative edge stack that is configured to hold the remaining edges of clusters from the spanning forest.
[0105] Optionally, at 1840, decoding the adjoint data may include accessing one or more of the generated edge stacks at the Corr engine. Optionally, at 1845, decoding the adjoint data may include generating a memory request at the Corr engine based on the accessed edge stacks. Optionally, at 1850, decoding the adjoint data may include performing iterative peeling decoding on each of the accessed edge stacks at the Corr engine. Optionally, at 1855, decoding the adjoint data may include updating the decoder's error log at the Corr engine based on the results of the iterative peeling decoding.
[0106] As discussed herein, decoding based on a single round of measurements will not account for adjoint measurement errors. To cope with measurement errors, the decoder examines d (code distance) rounds of measurements. This type of error correction can be achieved with minimal changes to the design. For example, the decoder can analyze a 3D graph instead of forming a graph on a 2D plane. Each vertex can be connected to at most 4 neighbors. However, for a 3D graph, each vertex can now have up to two additional edges corresponding to the previous round of measurements and the next round of measurements. To reduce the storage overhead, the STM for each round of adjoint measurements can be stored. The STM can be optimized such that each row of the STM stores the vertices of a row of the surface code lattice, the edge information of the vertices of the next row, and the edge information connecting the corresponding vertices in the surface code lattice of the next round.
[0107] The compression techniques described herein can reduce the amount of memory required to store the adjoint data and error logs of data qubits. However, the microarchitecture of the union-find decoder also uses memory, and the total capacity required is far from the total capacity provided by superconducting memory. Therefore, the design can be implemented by using conventional CMOS operating at 77K. This can also reduce the thermal noise generated in the cryogenic environment close to the quantum substrate, as the design is physically far from the quantum substrate.
[0108] For the baseline design, a naive implementation can allocate a decoder for each X adjoint and each Z adjoint of each logical qubit, as Figure 19 shown at 1900. Figure 19The system architecture of a large number of logical qubits within the quantum register 1910 is schematically shown. The quantum register 1910 is shown to include logical qubit 0 1910a, logical qubit 1 1910b, and logical qubit l 1910l as representative logical qubits operating at 15 - 20 mK. Each logical qubit is configured to receive signals from the control logic 1915 and output adjoint data to a compression engine (e.g., 1920a, 1920b... 1920l). The control logic 1915 and the compression engines 1920a... 1920l are shown to operate at a higher temperature of 4K than the quantum register. However, higher (e.g., 8K) or lower (e.g., 2K) temperatures can be used.
[0109] Each compression engine routes the compressed adjoint data to a decompression engine (1925a, 1925b... 1925l) operating at 77K. The decompression engine decompresses the compressed adjoint data and routes the decompressed X adjoint data and Z adjoint data to the decoding block 1930. In this example, each decompression engine is coupled to a pair of pipelined union - find decoders (1935a, 1935b, 1935c, 1935d... 1935k, 1935l) operating at 77K. Each union - find decoder analyzes the adjoint data received from the decompression engine and updates the error log 1940. Although shown to operate at 77K, higher (e.g., 85K) or lower (e.g., 70K) temperatures can be used to operate the decompression engine and the decoder, although the operating temperature of the decompression engine and the decoder can generally be higher than that of the compression engine.
[0110] Thus, for the baseline design, the decoding logic can use 2L union - find decoders per logical qubit. In this implementation, each logical qubit uses its own dedicated decoder. However, the utilization rate of each pipeline stage may be different. Therefore, the architecture shown at 1900 may not provide the optimal allocation of resources. For a large number of qubits, the on - chip components are under - utilized and dissipate heat. Since the entire system operates at 77K, the increased power consumption linearly increases the cost of cooling.
[0111] Thus, an architecture including a decoder block with a reduced number of pipeline units can be used. In Figure 20FIG. 2000 shows an example design of such a decoder block. A qubit register 2005 including a plurality of logical qubits transfers accompanying data to a set of Gr-Gen modules. A set of Gr-Gen modules 2010 may share one or more DFS engines 2020, and a set of DFS engines 2020 may share one or more Corr engines 2030. Hardware overhead includes a first set of multiplexers 2035 coupling the set of Gr-Gen modules 2010 to a DFS engine 2020, and a second set of multiplexers 2040 coupling the set of DFS engines 2020 to a Corr engine 2030. Memory requests generated by the Corr engine 2030 may be routed to the correct memory location using a demultiplexer 2045. Selection logic 2050 may prioritize the first ready component and may use round robin arbitration to generate appropriate selection signals for the multiplexers 2035 and 2040. For example, if four Gr-Gen modules 2010 share a DFS engine 2020 and the second Gr-Gen module finishes cluster formation earlier than the others, it may access the corresponding DFS engine 2020 first. Thus, the round robin strategy ensures fairness while sharing resources.
[0112] In Figure 21 FIG. 2100 shows an example system architecture. A qubit register 2105 includes a plurality of logical qubits 2110 coupled to control logic 2115. Each logical qubit is coupled to a compression engine 2120, and each compression engine 2120 is in turn coupled to a decompression engine 2125. A block of N logical qubits 2110 shares a decoder block 2130, and the decoder block 2130 updates an error log 2135 for each coupled logical qubit 2110. As described with respect to Figure 19 the operating temperature may vary with the indicated temperatures of 4K and 77K. If N logical qubits share a decoder block 2130, then for a quantum register 2105 having L logical qubits 2110, the total number of decoder blocks 2130 required is L / N. An example microarchitecture uses L Gr-Gen modules, (a) L DFS engines, and (b) L Corr engines. Resource savings depend on parameters (a) and (b). The values of (a) and (b) can be calculated to minimize the overall hardware cost. This can be framed as a constrained optimization problem.
[0113] One way to decode a large-scale system is to assign a decoder to each logical qubit. However, this method results in linear growth with respect to hardware, thus leading to linear growth in power cost. Consequently, this design is not very efficient and is not scalable by itself. The design herein supports the reuse of specific design components in order to reduce the actual cost when decoder blocks are scaled up for a large number of logical qubits.
[0114] Resources can be shared within and / or across decoding units. Given the distribution of decoding times, it is unlikely that several very long adjoint vectors will need to be decoded simultaneously, so resources can be shared.
[0115] This sharing is independent of the decoder or decoding algorithm, including cases where the decoding algorithm has a run time that depends on the adjoint, so some adjoints may be more difficult or take longer to decode than others. For example, some machine learning-based decoders do not depend on the adjoint. A machine learning decoder can have a multi-layer neural network. Once decoding is performed on one qubit on the first layer, the second qubit can use the first layer while the first qubit works on the second layer of the network.
[0116] Figure 22 An example method 2200 of a quantum computing device is shown. Method 2200 can be performed by a multiplexed quantum computing device (such as, Figure 20 and Figure 21 the computing device shown). At 2205, method 2200 includes generating an adjoint from at least one quantum register including l logical qubits, where l is a positive integer. The generated adjoint can include X and Z adjoints. At 2210, method 2200 includes routing the generated adjoint to a set of d decoder blocks coupled to the at least one quantum register, where d < 2*l. As described with respect to Figure 20 and Figure 21 this allows for scalability of the quantum computing device, as fewer than two decoders are needed to handle the processing of X and Z adjoints for each logical qubit.
[0117] In some examples, each decoder block is configured to receive a decoding request from a set of n logical qubits, where n > 1. In some examples, each decoder block includes g Gr-Gen modules, where 0 < g ≤ l, and each Gr-Gen module is configured to generate spanning tree memory (STM) data based on the received adjoint. In some examples, each decoder block further includes α*l DFS engines, where 0 < α < 1. In some examples, two or more Gr-Gen modules are coupled to each DFS engine via a multiplexer in a first set of multiplexers.
[0118] Optionally, at 2215, method 2200 includes accessing, at each DFS engine, STM data generated by two or more Gr-Gen modules via one of a first set of multiplexers. Optionally, at 2220, method 2200 includes generating, at each DFS engine, an edge stack based on the STM data. In some examples, each decoder block further includes β*l Corr engines, where 0 < β < 1. In some examples, two or more DFS engines are coupled to each Corr engine via one of a second set of multiplexers.
[0119] Optionally, at 2225, method 2200 includes accessing, at each Corr engine, an edge stack generated by two or more DFS engines via one of a second set of multiplexers. Optionally, at 2230, method 2200 includes generating a memory request based on the accessed edge stack. Optionally, at 2235, method 2200 includes routing, via one or more demultiplexers, the memory requests generated by each Corr engine to a memory location. Optionally, at 2240, method 2200 includes routing return signals via each multiplexer in the first set of multiplexers and the second set of multiplexers based on round-robin arbitration.
[0120] When decoding d rounds of syndromes within one logical cycle (T), error correction is successful, which limits the maximum latency that the decoder can tolerate. When the decoder cannot decode all syndromes within the logical cycle, errors may not be detected. This type of failure can be referred to as a timeout failure. Since the decoder is imperfect and exhibits threshold behavior, there is also a possibility of a logical error occurring when the correction generated by the decoder changes the logical state of the qubit. Therefore, the failures of the decoder can be attributed to timeout failures or logical errors. To keep the error threshold the same and prevent the system failure rate from increasing, the probability of a timeout failure (p tof ) must be lower than the occurrence probability of a logical error (p log ), as shown in Equation (4). For an optimized design, resource sharing is possible as long as p tof is small enough.
[0121] p tof ≤p log (Equation 4)
[0122] Assume that N logical qubits with exactly the same error rate share k decoding units. The total execution time for decoding N logical qubits is given by Equation (5):
[0123]
[0124] where τ iDenotes the execution time for decoding the adjoint of the \(i\)-th logical qubit. In this case, the probability \(p\) of a timeout failure tof Must satisfy Equation (6).
[0125]
[0126] The optimization goal is to minimize the number \(k\) of decoding units for a given number \(N\) of logical qubits such that the constraints given by Equation (4) are satisfied. \(p\) tof The execution time obtained from the emulator can be used for modeling.
[0127] The decoder performance can be modeled by studying the number of reads. The write operations performed can be read-modify-write, and the write-back may not be on the critical path. For memory access and a 4 GHz clock frequency, a 4-cycle delay is assumed. The total number of memory requests in Gr-Gen for a given adjoint is proportional to the cluster diameter (\(D_i\)). However, it is proportional to the cluster size (\(S_i\)) in the DFS engine and the Corr engine. Equations (7) and (8) give the execution times spent in Gr-Gen (\(T_{GG}\)), DFS engine (\(T_{DFS}\)), and Corr engine (\(T_{CE}\)) for an adjoint with \(n\) clusters.
[0128]
[0129] \(\tau\) DFS \(=\tau\) CE \(=\sum\) i \(S\) i (Equation 8)
[0130] In the optimized design, each Gr-Gen unit grows clusters for \(X\) and \(Z\) adjoints. Two or more Gr-Gen units use one DFS engine module, and two or more DFS engines use one Corr engine. These numbers of units to be shared can be determined by the fraction of the total execution time spent in each pipeline stage.
[0131] Below, the simulation infrastructure for making design choices in the decoder microarchitecture is discussed. This infrastructure supports the estimation of some key statistics of the union-find decoder and also supports the study of the performance of the compression techniques described in this article.
[0132] A Monte Carlo emulator is used to analyze the performance of different compression techniques and obtain statistical data on the performance of the union-find decoder. Figure 23 The Monte Carlo emulator 2300 is schematically shown. Different configurations span four different physical error rates, ten different code distances, and four noise models, each simulated one million times. The error rates selected are \(10\) -6 (Most optimistic), \(10\) -4, 10 -3 and 10 -2 (Most pessimistic). The simulator 2300 receives the code distance 2302, the noise model 2304, and the compression algorithm 2306. Based on the code distance 2302, the simulator 2300 generates a surface code lattice via the lattice generator 2308. Depending on the selected noise model 2304, the simulator injects errors on the data qubits of the surface code lattice via error injection 2310 and generates an output syndrome 2312. The output syndrome 2312 is then compressed via the compressor 2314 according to the input compression algorithm 2306 to generate a compressed syndrome 2316. The simulator 2300 then outputs the compression ratio. As a figure of merit for determining the most suitable compression scheme, the compression ratio (determined by Equation (9)) and the percentage of incompressible syndromes are used. The simulation is repeated one million times to calculate the average compression ratio and the percentage of incompressible syndromes.
[0133]
[0134] The simulator also runs the union-find decoding algorithm on the syndrome 2312 via the decoder 2318. The statistics generator 2320 then analyzes the distribution of the cluster sizes, the average number of clusters on a given lattice, and the execution time spent in each pipeline stage of the decoder 2318 by modeling the hardware. These statistics and performance numbers provide insights that help in the microarchitecture design of the decoder hardware implementation and drive scalable design.
[0135] The performance of the decoder depends to a large extent on the noise model of the underlying qubits. Therefore, four different error models were explored. Assuming identically and independently distributed (iid) errors, the depolarizing noise model was chosen as the most basic noise model. In the depolarizing noise model, if the error rate is p, the probability that each physical qubit encounters an error is p, and the probability of remaining error-free is (1 - p). Additionally, in this error model, X, Y, and Z errors occur with equal probability p / 3. The other three noise models assume different probabilities for X and Z errors, as shown in Table 2.
[0136]
[0137] Table 2
[0138] This paper discusses the results of syndrome compression, baseline union-find decoder design, and scalability analysis. The results of the baseline decoder and scalability analysis are based on the d (code distance) round syndrome measurements described in this paper.
[0139] The performance of each compression scheme depends on the noise model. For the depolarizing noise model, compression schemes such as DZC and Geo-Comp provide better performance at low code distances that depend on the error rate compared to sparse representation. For noise models with relative biases for specific types of errors (such as P x = 10P z and P x = 100P z ), DZC performs better than Geo-Comp. For lower code distances, even though sparse representation provides a higher compression ratio, for larger error rates, the percentage of uncompressible syndrome is also higher (up to 6%). For noise models where one type of error probability is much larger than the other, better compression ratios are obtained by separately compressing the X syndrome and Z syndrome at the cost of greater hardware complexity. If only one type of compression can be used due to hardware constraints, for lower code distances, DZC has better performance. Table 3 specifies different noise mechanisms and the appropriate compression schemes that are most effective in each mechanism. Generally speaking, for most cases in mechanisms with low error rates, the performance of sparse representation is better.
[0140]
[0141] Table 3
[0142] Figure 24 shows graph 2400, which shows the mean compression ratio of the X syndrome of the depolarizing noise channel using the selected compression scheme for different physical error rates and noise mechanisms. The depolarizing noise channel is considered a representative candidate. Similar results are observed for the Z syndrome.
[0143] The distribution of the cluster diameter was determined from the simulation. As defined herein, the cluster diameter is the maximum distance between any two boundary vertices of the cluster. Figure 25 shows graph 2500, which indicates the average cluster diameter for different error rates and code distances of logical qubits. The average cluster diameter is low. This result is used to eliminate the storage cost incurred when maintaining the boundary list of each cluster in the hardware (i.e., the function used in the original union-find algorithm). This reduces the hardware cost of the Gr-Gen module. The probability that the cluster diameter will become smaller increases as the error rate decreases.
[0144] The spanning tree memory (STM) used by the Gr-Gen module and the DFS engine occupies most of the storage cost. Figure 26Figure 2600 is shown, which indicates the total memory capacity required for a spanning tree memory (STM) for a given code distance (d) and number of logical qubits (N). This shown result is a 3D graph constructed using d-round measurements. Figure 2600 shows that even for a large number of logical qubits (such as 1000), for very large code distances (d) and d-round measurements, the total memory required to decode both the X adjoint and Z adjoint is less than 10 MB. If the decoder does not need to consider d-round measurements (assuming perfect measurements may be possible in the future), the required total memory capacity will be reduced by a factor of d.
[0145] The maximum possible number of entries in the root table and the size table is the total number of adjoint bits for d (code distance) round adjoint measurements (equal to 2d(d - 1)). Each root table entry includes a root that can be uniquely identified using log 2 2d 2 (d - 1) bits. Similarly, the maximum feasible cluster size includes all adjoint bits. Thus, for each logical qubit, the total size of the root table and the size table is 2d 2 (d - 1)log 2 2d 2 (d - 1) bits.
[0146] The size of the stack can be determined by analyzing the maximum number of edges within a cluster from Monte Carlo simulations. The number of edges in a cluster follows a Poisson distribution. Figure 27 Figure 2700 is shown, which indicates such a distribution for code distance d = 11 and physical error rate p = 10 -3 . Thus, the stack size can be designed to be half of the maximum number of edges. Figure 28 Figure 2800 is a graph showing the average number of edges in a cluster for different code distances and error rates. Each stack stores two vertices (log 2 4d 2 (d - 1) bits), a growth direction (2 bits), and 1 bit adjoint. It is worth noting that each DFS engine includes 2 stacks for pipelining. If the size of the cluster is greater than what each stack can hold, an overflow bit can be set and an alternative stack can be used when available.
[0147] Figure 29 Figure 2900 is shown, which indicates the correlation between Gr-Gen and the execution time in the DFS engine. This means that more time is spent in the Gr-Gen unit during decoding. As described herein, this data is used to select the number of resources to be shared within the decoder block. Figure 30 Figure [X] is shown, which indicates for a single decoder block with code distance (d) of 11 and error rate (p) of 0.5x10 -3 (e.g., as Figure 20The graph 3000 of the execution time distribution as shown). The shaded area indicates events that may increase the probability of timeout failures. In the case of implementing resource sharing, the probability p of timeout failures tof is lower than the probability of a logical error rate of 10 -8 . For L logical qubits, the numbers of Gr-Gen modules, DFS engines, and Corr engines used in this architecture are L, L / 2, and L / 2, respectively. Therefore, the total numbers of Gr-Gen modules, DFS engines, and Corr engines are reduced by 2x, 4x, and 4x, respectively.
[0148] Error correction is an integral part of classical computing related to quantum computers. Error decoding algorithms are designed to achieve higher error correction capabilities (thresholds). In this paper, a microarchitecture for the hardware implementation of the union-find decoder is disclosed, which uses CMOS operating at 77K. Accompanying compression is feasible to meet the bandwidth requirements of 4K–77K links. Different compression schemes perform differently under different noise mechanisms, where sparse data representation generally performs better for lower error rates and larger code distances. The disclosed microarchitecture is designed to scale the decoder to thousands of logical qubits. The architecture includes three pipeline stages and is tuned for high performance, high throughput, and low hardware complexity. The design can be scaled for a larger number of logical qubits for practical fault-tolerant quantum computing. The time spent in each pipeline stage is different, so the utilization rate of each stage is also different. By considering this, an architecture that relies on resource sharing across multiple logical qubits is disclosed. Supporting such resource sharing enables the logical error rate to remain unaffected and minimizes the system failure rate due to errors that cannot be decoded due to a lack of decoding resources.
[0149] In some embodiments, the methods and processes described herein may be associated with the computing systems of one or more computing devices. Specifically, such methods and processes may be implemented as computer applications or services, application programming interfaces (APIs), libraries, and / or other computer program products.
[0150] Figure 31 A non-limiting embodiment of the computing system 3100 is schematically shown, and the computing system 3100 may implement one or more of the above methods and processes. The computing system 3100 is shown in a simplified form. The computing system 3100 may embody the host computer device described above and shown in Figure 1 . The computing system 3100 may take the form of one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, gaming devices, mobile computing devices, mobile communication devices (e.g., smart phones), and / or other computing devices, as well as wearable computing devices (such as smart watches and head-mounted augmented reality devices).
[0151] The computing system 3100 includes a logical processor 3102, a volatile memory 3104, and a non-volatile storage device 3106. The computing system 3100 may optionally include a display subsystem 3108, an input subsystem 3110, a communication subsystem 3112, and / or Figure 31 other components not shown.
[0152] The logical processor 3102 includes one or more physical devices configured to execute instructions. For example, the logical processor may be configured to execute instructions as part of one or more applications, programs, routines, libraries, objects, components, data structures, or other logical constructs. Such instructions may be implemented to perform tasks, implement data types, transform the state of one or more components, achieve a technical effect, or otherwise achieve a desired result.
[0153] The logical processor may include one or more physical processors (hardware) configured to execute software instructions. Additionally or alternatively, the logical processor may include one or more hardware logic circuits or firmware devices configured to execute hardware-implemented logic or firmware instructions. The processors of the logical processor 3102 may be single-core or multi-core, and the instructions executed thereon may be configured for sequential, parallel, and / or distributed processing. Optionally, the various components of the logical processor may be distributed among two or more separate devices, which may be located remotely and / or configured for cooperative processing. Aspects of the logical processor may be virtualized and executed by remotely accessible networked computing devices configured in a cloud computing configuration. In such a case, it can be understood that these virtualized aspects run on different physical logical processors of various different machines.
[0154] The non-volatile storage device 3106 includes one or more physical devices configured to store instructions executable by the logical processor to implement the methods and processes described herein. When implementing such methods and processes, the state of the non-volatile storage device 3106 may be transformed—for example, to store different data.
[0155] The non-volatile storage device 3106 may include removable and / or built-in physical devices. The non-volatile storage device 3106 may include optical memories (e.g., CD, DVD, HD-DVD, Blu-ray Disc, etc.), semiconductor memories (e.g., ROM, EPROM, EEPROM, flash memory, etc.), and / or magnetic memories (e.g., hard disk drive, floppy disk drive, tape drive, MRAM, etc.) or other mass storage device technologies. The non-volatile storage device 3106 may include non-volatile, dynamic, static, read / write, read-only, sequential access, location-addressable, file-addressable, and / or content-addressable devices. It will be understood that the non-volatile storage device 3106 is configured to save instructions even when the non-volatile storage device 3106 is powered off.
[0156] The volatile memory 3104 may include a physical device containing random access memory. The logical processor 3102 typically utilizes the volatile memory 3104 to temporarily store information during software instruction processing. It will be understood that when the volatile memory 3104 is powered off, the volatile memory 3104 generally does not continue to store instructions.
[0157] Aspects of the logical processor 3102, the volatile memory 3104, and the non-volatile storage device 3106 may be integrated together into one or more hardware logic components. For example, such hardware logic components may include field programmable gate arrays (FPGA), program and application specific integrated circuits (PASIC / ASIC), program and application specific standard products (PSSP / ASSP), system on a chip (SOC), and complex programmable logic devices (CPLD).
[0158] When included, the display subsystem 3108 may be used to present a visual representation of data saved by the non-volatile storage device 3106. The visual representation may take the form of a graphical user interface (GUI). Since the methods and processes described herein change the data saved by the non-volatile storage device, thereby transforming the state of the non-volatile storage device, the state of the display subsystem 3108 may likewise be transformed to visually represent the changes in the underlying data. The display subsystem 3108 may include one or more display devices utilizing almost any type of technology. Such display devices may be combined with the logical processor 3102, the volatile memory 3104, and / or the non-volatile storage device 3106 in a shared enclosure, or such display devices may be peripheral display devices.
[0159] When included, input subsystem 3110 can include one or more user input devices (such as a keyboard, mouse, touch screen, or game controller), or interface with user input devices. In some embodiments, the input subsystem can include selected natural user input (NUI) components, or interface with selected NUI components. Such components can be integrated or peripheral, and the translation and / or processing of input actions can be handled on or off the system. Example NUI components can include a microphone for voice and / or speech recognition; infrared, color, stereo, and / or depth cameras for machine vision and / or gesture recognition; a head tracker, eye tracker, accelerometer, and / or gyroscope for motion detection and / or intent recognition; and an electric field sensing component for evaluating brain activity; and / or any other suitable sensors.
[0160] When included, communication subsystem 3112 can be configured to communicatively couple the various computing devices described herein to each other and to other devices. Communication subsystem 3112 can include wired and / or wireless communication devices compatible with one or more different communication protocols. As a non-limiting example, the communication subsystem can be configured to communicate via a wireless telephone network, or a wired or wireless local or wide area network (such as HDMI over a Wi-Fi connection). In some embodiments, the communication subsystem can allow computing system 3100 to send messages to and / or receive messages from other devices via a network such as the Internet.
[0161] In one example, a quantum computing device includes: at least one quantum register including a plurality of logical qubits; compression engines, each coupled to each of the plurality of logical qubits, each compression engine configured to compress adjoint data; and decompression engines, each coupled to each compression engine, each decompression engine configured to: receive the compressed adjoint data; decompress the received compressed adjoint data; and route the decompressed adjoint data to a decoder block. In such an example or any other example, at least one compression engine is additionally or alternatively configured to compress adjoint data using dynamic zero compression. In any of the foregoing examples or any other example, at least one compression engine is additionally or alternatively configured to compress adjoint data using sparse representation. In any of the foregoing examples or any other example, at least one compression engine is additionally or alternatively configured to compress adjoint data using geometry-based compression. In any of the foregoing examples or any other example, the plurality of logical qubits are additionally or alternatively divided into two or more sectors, where a first sector of one or more sectors is coupled to a first type of compression engine configured to compress adjoint data using a first type of compression, and a second sector of one or more sectors is coupled to a second type of compression engine configured to compress adjoint data using a second type of compression. In any of the foregoing examples or any other example, the compression engines are additionally or alternatively configured to operate at a higher temperature than the quantum register. In any of the foregoing examples or any other example, the decompression engines and the decoder block are additionally or alternatively configured to operate at a higher temperature than the compression engines. In any of the foregoing examples or any other example, each decompression engine additionally or alternatively routes the decompressed adjoint data to a graph generator module of the decoder block. In any of the foregoing examples or any other example, the decompressed adjoint data additionally or alternatively includes at least X adjoint data and Z adjoint data. In any of the foregoing examples or any other example, the plurality of logical qubits additionally or alternatively includes l logical qubits, where the quantum computing device includes a set of d decoder blocks, where d < l.
[0162] In another example, a method for a quantum computing device includes: generating adjoint data from at least one quantum register including l logical qubits, where l is a positive integer; and for each logical qubit: routing the generated adjoint data to a compression engine configured to compress the adjoint data; routing the compressed adjoint data to a decompression engine configured to: receive the compressed adjoint data; and decompress the received compressed adjoint data; and routing the decompressed adjoint data to a decoder block. In such an example or any other example, at least one compression engine is additionally or alternatively configured to compress the adjoint data using dynamic zero compression. In any of the foregoing examples or any other example, at least one compression engine is additionally or alternatively configured to compress the adjoint data using a sparse representation. In any of the foregoing examples or any other example, the method additionally or alternatively includes operating the compression engine at a higher temperature than the quantum register. In any of the foregoing examples or any other example, the method additionally or alternatively includes operating the decompression engine and the decoder block at a higher temperature than the compression engine. In any of the foregoing examples or any other example, each decompression engine additionally or alternatively routes the decompressed adjoint data to a graph generator module of the decoder block. In any of the foregoing examples or any other example, the quantum computing device additionally or alternatively includes a set of d decoder blocks, where d < 2*l.
[0163] In yet another example, a method for a quantum computing device includes: generating adjoint data from at least one surface code lattice including l logical qubits, where l is a positive integer and the surface code lattice is divided into two or more regions based on a lattice geometry; and for each logical qubit: routing the generated adjoint data to a compression engine configured to compress the adjoint data using geometry-based compression; routing the compressed adjoint data to a decompression engine configured to: receive the compressed adjoint data; and decompress the received compressed adjoint data; and routing the decompressed adjoint data to a decoder block. In such an example or any other example, compressing the adjoint data using geometry-based compression additionally or alternatively includes: using zero indicator bits to compress the adjoint data for each of the two or more regions of the surface code lattice; and transmitting only the adjoint data from non-zero regions. In any of the foregoing examples or any other example, the number of regions is additionally or alternatively determined based on an expected number of data blocks containing trivial adjoints.
[0164] It will be understood that the configurations and / or methods described herein are exemplary in nature, and these specific embodiments or examples should not be considered limiting as many variations are possible. The specific routines or methods described herein may represent one or more of any number of processing strategies. As such, the various acts shown and / or described may be performed in the order shown and / or described, in other orders, in parallel, or omitted. Likewise, the order of the above processes may be changed.
[0165] The subject matter of the present disclosure includes all novel and non-obvious combinations and sub-combinations of the various processes, systems, and configurations, as well as other features, functions, acts, and / or properties disclosed herein, and any and all equivalents thereof.
Claims
1. A quantum computing device, comprising: A surface code lattice, the surface code lattice comprising l logical qubits, where l is a positive integer, and the surface code lattice being divided into two or more regions based on the lattice geometry; Compression engines, each compression engine coupled to each of the l logical qubits, each compression engine being configured to compress syndrome data generated by the surface code lattice using a geometry-based compression scheme; and Decompression engines, each decompression engine coupled to each compression engine, each decompression engine being configured to: Receive the compressed syndrome data; Decompress the received compressed syndrome data; and Route the decompressed syndrome data to a decoder block.
2. The quantum computing device according to claim 1, wherein the total number of decoder blocks < l.
3. The quantum computing device according to claim 1, wherein at least one decoder block is included in a hardware implementation of a joint-finding decoder.
4. The quantum computing device according to claim 1, wherein the two or more regions include regions of unequal size.
5. The quantum computing device according to claim 4, wherein the number of regions is determined based on the expected number of data blocks containing trivial syndromes, the trivial syndromes including all-zero values.
6. The quantum computing device according to claim 1, wherein the number of qubits in a region does not increase with an increasing error rate.
7. The quantum computing device according to claim 1, wherein the geometry-based compression scheme compresses a region of X syndromes and a region of Z syndromes together.
8. The quantum computing device according to claim 7, wherein the compression engine uses zero indicator bits for each of the two or more regions of the surface code lattice to compress the syndrome data such that the compressed syndrome data includes only syndrome data from non-zero regions.
9. The quantum computing device according to claim 8, wherein one or more of the data qubits are positioned at a region boundary such that the associated syndrome bits are in adjacent regions.
10. The quantum computing device according to claim 7, wherein the decompression engine is programmed based on the geometry-based compression scheme used by the compression engine.
11. A method for a quantum computing device, comprising: Generating syndrome data from at least one surface code lattice comprising l logical qubits, where l is a positive integer, each logical qubit coupled to one or more X syndromes and one or more Z syndromes, the surface code lattice being divided into two or more regions based on the lattice geometry; Routing the generated syndrome data to a compression engine, the compression engine using a geometry-based compression scheme to compress the syndrome data, the geometry-based compression scheme compressing a region of X syndromes and a region of Z syndromes together; Send the compressed syndrome data to a decompression engine that uses the geometry-based compression scheme to decompress the compressed syndrome data; and Route the decompressed syndrome data to a decoder block.
12. The method according to claim 11, wherein compressing the syndrome data using geometry-based compression includes compressing the syndrome data using zero indicator bits for each of the two or more regions of the surface code lattice.
13. The method according to claim 12, wherein the compressed syndrome data includes only syndrome data from non-zero regions.
14. The method according to claim 12, wherein one or more of the data qubits are located on a region boundary such that the associated syndrome bits are in adjacent regions.
15. The method according to claim 11, wherein two or more decompression engines send the compressed syndrome data to a common decoder engine.
16. The method according to claim 11, wherein routing the decompressed syndrome data includes routing to a decoder block included in a hardware implementation of a union-find decoder.
17. A method for a quantum computing device, comprising: Partition a surface code lattice into two or more regions based on lattice geometry, the surface code lattice including l logical qubits, where l is a positive integer; Generate syndrome data from the surface code lattice; Route the generated syndrome data to a compression engine that is trained to compress the syndrome data using a geometry-based compression scheme; the compression engine operates at a higher temperature than the surface code lattice; Send the compressed syndrome data to a decompression engine that operates at a higher temperature than the compression engine; At the decompression engine, decompress the sent compressed syndrome data; and Route the decompressed syndrome data to a decoder block.
18. The method according to claim 17, wherein each logical qubit of the surface code lattice is coupled to one or more X syndromes and one or more Z syndromes, and wherein the compression engine compresses the regions of the X syndromes and the regions of the Z syndromes together into the compressed syndrome data.
19. The method according to claim 17, wherein partitioning the surface code lattice into two or more regions is further based on an expected number of data blocks containing trivial syndromes, the trivial syndromes including all-zero values.
20. The method according to claim 17, wherein partitioning the surface code lattice into two or more regions generates two or more regions of unequal size.