Special accelerator for hierarchical greedy decoding algorithm
By designing a dedicated accelerator for the hierarchical greedy decoding algorithm and utilizing the sparsity and parallelism of qLDPC codes, the latency and resource usage issues of existing decoding algorithms on superconducting quantum platforms are resolved, achieving low-latency and efficient quantum error correction and supporting real-time quantum computing.
Patent Information
- Application Number
- CN202511171413.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Existing quantum low-density parity-check code (qLDPC) decoding algorithms are difficult to achieve real-time and accurate error correction on superconducting quantum platforms. Traditional decoding algorithms such as BP and BP+OSD have problems with decoding delay and high resource usage, and cannot meet the needs of fault-tolerant quantum computing.
A dedicated accelerator for the hierarchical greedy decoding algorithm is designed. By utilizing the sparsity of qLDPC codes and the parallelism of the hierarchical greedy decoding algorithm, the data flow is optimized and dedicated computing units, including transform units, decoding cores, and permutation units, are designed to achieve parallel computing and data reuse, thereby reducing decoding latency.
It achieves a decoding delay of less than 1μs, can process larger-scale qLDPC codes, has low resource utilization, and is suitable for integration into the FPGA control processor of the superconducting quantum platform, supporting real-time and accurate quantum error correction.
Smart Images

Figure CN120671862A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of quantum computing technology, and in particular relates to a special accelerator for a hierarchical greedy decoding algorithm. Background Art
[0002] Quantum error correction (QEC) is one of the key technologies for achieving large-scale fault-tolerant quantum computing. Within the QEC field, quantum low-density parity-check (qLDPC) codes have attracted considerable attention due to their constant coding rate and high error threshold, making them a promising quantum error-correcting code capable of realizing large-scale fault-tolerant quantum systems. However, due to the large, sparse, and irregular parity check matrix of qLDPC codes, real-time decoding remains a significant challenge.
[0003] Traditional decoding algorithms, such as belief propagation (BP) and its improved version, belief propagation plus ordered statistical decoding (BP+OSD), cannot meet the requirements for accurate and real-time decoding on superconducting quantum platforms. Specifically, while BP offers advantages in computational complexity and parallelism, making it potentially suitable for decoding acceleration on hardware such as FPGAs and ASICs, it is limited by quantum degeneracy (i.e., multiple possible error modes correspond to the same error symptom). BP often converges to incorrect error modes, resulting in low decoding accuracy and ineffective decoding. For example, in decoding the bivariate cyclic (BB) code [[784,24,24]], the logical error rate of BP is 1649.5 times higher than that of BP+OSD. To increase the probability of BP finding the correct solution, the number of BP iterations must be increased. However, BP iterations can only be performed serially, and increasing the number of BP iterations significantly increases decoding time. Furthermore, as the size of the error-correcting code increases, the time required for each BP iteration also increases. For example, for BB code [[90,8,10]], the decoding time of BP exceeds the real-time decoding requirement (1μs) on the superconducting platform.
[0004] In order to improve the decoding accuracy of BP, BP+OSD introduced OSD as a post-processing technology to correct the decoding results of BP by solving linear systems. However, OSD involves computationally intensive sequential operations such as sorting and matrix inversion, which makes it difficult to deploy on hardware acceleration platforms such as FPGA and ASIC, resulting in significant delays. This high delay makes it difficult for BP+OSD to meet the requirements of real-time decoding for fault-tolerant quantum computing. Quantum systems usually require the decoder to complete the operation within each QEC cycle (usually less than 1μs). If the decoding speed cannot keep up with the speed of error symptom generation, error symptoms will occur, which will delay subsequent quantum operations, and this delay overhead increases exponentially with the depth of the circuit. Even for small BB codes [[72,12,6]], the decoding time of BP+OSD is as high as about This significantly exceeds the 1μs requirement for real-time decoding, highlighting the limitations of existing software decoding algorithms in achieving high-performance hardware acceleration.
[0005] To meet the demands of real-time decoding, the industry has devoted significant attention to dedicated decoding hardware accelerators. However, general-purpose processors (such as CPUs or GPUs) are inefficient when handling the fine-grained, bit-level operations common in QEC. This is because these operations are incompatible with the wide vector execution units of CPUs and GPUs, resulting in low utilization and inefficient memory access. In contrast, FPGAs (field-programmable gate arrays) are more suitable for such tasks due to their reconfigurability and high parallelism. Furthermore, FPGA-based decoding accelerators can be tightly integrated with superconducting quantum systems, as error symptoms are typically generated and transmitted by FPGA-implemented read modules. This tightly coupled decoding accelerator allows direct access to error symptom data, avoiding the latency associated with transmitting data to an external processor and further reducing decoding time.
[0006] Therefore, there is an urgent need for a dedicated hardware accelerator that can fully utilize the sparsity of qLDPC codes and the parallelism of decoding algorithms to overcome the bottlenecks of existing decoders in accuracy and latency and realize real-time qLDPC decoding on superconducting quantum platforms. Summary of the Invention
[0007] In light of the above, the purpose of this invention is to provide a dedicated accelerator for the hierarchical greedy decoding algorithm. By designing a hardware accelerator that fully exploits the sparsity of qLDPC codes and the parallelism of the hierarchical greedy decoding algorithm, this accelerator enables real-time qLDPC decoding. Furthermore, by optimizing data flows and designing dedicated computational units, this accelerator addresses the inability of traditional processors like CPUs and GPUs to support fine-grained bit-level operations, significantly reducing decoding latency.
[0008] To achieve the above-mentioned purpose, the present invention provides the following technical solutions: An embodiment of the present invention provides a dedicated accelerator for a hierarchical greedy decoding algorithm, comprising: The transformation unit is configured to receive an original error symptom and multiply it by a transformation matrix through sparse matrix-vector multiplication to generate a transformed error symptom; The decoding core is used to execute a hierarchical greedy decoding algorithm, divide the error pattern into left errors and right errors, and perform calculations including guessing right errors and decoding left errors in parallel based on the transformed error symptoms to obtain a new value of the error pattern; The replacement unit is used for receiving a new value of the error pattern and multiplying it with the replacement matrix through sparse matrix-vector multiplication to calculate a final error pattern.
[0009] Preferably, the decoding core comprises: a plurality of hierarchical decoding units, a first comparison tree and an execution controller; A plurality of hierarchical decoding units are used to receive the transformed error symptoms to calculate the objective function value of each candidate error guess in parallel; The first comparison tree is used to receive the objective function values from the plurality of hierarchical decoding units and compare and identify the minimum objective function value, and store the minimum objective function value and the corresponding new value of the error mode after comparing with the current minimum objective function value; The execution controller is used to control the continuation or termination of the decoding program.
[0010] Preferably, each hierarchical decoding unit comprises: a first error symptom incremental updating unit, a plurality of greedy decoding cores and a first log-likelihood ratio calculation unit; The first error symptom incremental update unit is used to perform sparse matrix-vector multiplication and XOR operation to calculate the error symptom corresponding to the left error, and split the error symptom corresponding to the left error into K partial left error symptoms and send them to K greedy decoding cores respectively; The multiple greedy decoding cores are used to process K partial left error symptoms respectively, and splice the decoded partial left errors into a complete left error; The first log-likelihood ratio calculation unit is used to calculate the objective function value according to the complete left error.
[0011] Preferably, in the first error symptom incremental update unit, a sparse matrix is stored in a specific compression format, including a sparse matrix table and a non-zero row index table; during calculation, the row indices of non-zero elements are extracted from the sparse matrix table and the non-zero row index table according to the input column index, and an XOR operation is applied only to these non-zero elements; after performing the incremental update and sparse XOR operation to obtain the left error symptom, the value at the row index selected in the current iteration is operated with a specific value and written back; and the left error symptom corresponding to the error bit identified in the current iteration is stored so that the residual error symptom calculation data can be reused in subsequent iterations.
[0012] Preferably, each greedy decoding core comprises: a second error symptom incremental update unit, a second log-likelihood ratio calculation unit, and a second comparison tree; The second error symptom incremental updating unit is configured to receive a partial left error symptom, divide the left error to be decoded into a first candidate guess and a second candidate guess, calculate the second candidate guess based on the received partial left error symptom and the first candidate guess, and once the second candidate guess is obtained, concatenate the first candidate guess and the second candidate guess to form a left error; The second log-likelihood ratio calculation unit is used to calculate the objective function value according to the non-zero elements in the first candidate guess and the second candidate guess; The second comparison tree is used to identify the minimum value of the objective function value output by the second log-likelihood ratio calculation unit, and regard the left error corresponding to the minimum value as the best value of the current iteration and output it.
[0013] Preferably, the objective function value is to minimize the number of non-zero values in the left error, and use their indexes to retrieve the corresponding weights from the weight register file, and then apply the adder tree to parallel calculation to evaluate the possibility of the left error.
[0014] Preferably, in the second comparison tree, once it is found that the currently calculated objective function value is smaller than the existing minimum objective function value, the currently calculated objective function value is updated to the minimum objective function value, and guessing is stopped.
[0015] Preferably, in the first comparison tree, once it is found that the currently calculated objective function value is less than the existing minimum objective function value, the currently calculated objective function value is updated to the minimum objective function value, and the iteration is stopped.
[0016] Preferably, the transformation matrix and the permutation matrix are pre-loaded into a sparse matrix buffer before performing real-time decoding.
[0017] The accelerator hardware architecture designed in this invention fully utilizes the sparsity of the parity check matrix and the parallelism of the hierarchical greedy decoding algorithm. Furthermore, the error symptom incremental update unit avoids repeated calculations by storing and reusing error symptom data, further reducing decoding latency. Compared with the prior art, this invention has at least the following beneficial effects: (1) Extremely low decoding latency: The accelerator designed in this paper achieves a decoding latency of less than 1 μs. For example, for the largest BB code [[784,24,24]], its decoding latency is only 840 ns. This real-time performance is crucial for fault-tolerant quantum computing and can effectively prevent error symptoms from accumulating and delaying subsequent quantum operations. (2) Delay insensitive to code size: Compared to the BP decoder, the latency of the accelerator designed in this invention is insensitive to the size of the qLDPC code and exhibits a logarithmic growth trend. For example, for BB codes, when the parity check matrix column size increases by approximately 10 times, the latency only increases from 720ns to 840ns. This demonstrates that the accelerator architecture designed in this invention has good scalability and can handle larger-scale qLDPC codes.
[0018] (3) High efficiency and low resource usage: The accelerator designed in this invention has extremely low resource usage when implemented on an FPGA. For example, for a medium-sized BB code [[144,12,12]], the LUT and FF resource utilization rates are less than 10% and 1%, respectively. Even for a large BB code [[784,24,24]], only 31.26% of LUTs and 2.35% of FFs are required. This enables the accelerator designed in this invention to be seamlessly integrated into the FPGA control processor of current superconducting quantum platforms. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0020] Figure 1 Schematic diagram of the overall hardware architecture of a dedicated accelerator for a hierarchical greedy decoding algorithm provided by an embodiment of the present invention; Figure 2 is a pseudo code diagram of the hierarchical greedy decoding algorithm provided by an embodiment of the present invention; Figure 3 is a schematic diagram of a first error symptom incremental update unit provided by an embodiment of the present invention; Figure 4 Schematic diagram of the hardware architecture of the greedy decoding core provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0022] The inventive concept of the present invention is: in view of the limitations of low computational efficiency and high latency in the real-time performance of the qLDPC decoder in the prior art, the embodiment of the present invention provides a dedicated accelerator for a hierarchical greedy decoding algorithm. The accelerator aims to make full use of the sparsity characteristics of qLDPC codes, combined with the parallel computing advantages of the independently designed hierarchical greedy decoding algorithm, to reduce data transmission delays by carefully optimizing the data flow path. At the same time, a dedicated transformation unit, decoding core and permutation unit are designed, which are respectively responsible for the transformation processing of error symptoms, parallel decoding calculations and the permutation output of the final error pattern. The units work together to achieve efficient real-time decoding. This design not only improves the real-time performance of qLDPC decoding, but also ensures the accuracy and stability of decoding, providing strong support for high-speed communication systems.
[0023] like Figure 1 As shown, the embodiment provides a dedicated accelerator for the hierarchical greedy decoding algorithm, which is intended to support the hierarchical greedy decoding algorithm ( Figure 2 ) to achieve real-time decoding of qLDPC codes. This accelerator is designed directly based on the structure of the hierarchical greedy decoding algorithm, mapping each operation to dedicated logic on the FPGA. Through customized data flow and sparse computing design, it minimizes decoding latency. The dedicated accelerator mainly consists of three core components: a transform unit, a decoding core, and a permutation unit. The decoding core includes a hierarchical decoding unit (HDU), a first comparison tree, and an execution controller. The HDU includes a first error symptom incremental update unit, a greedy decoding core (GDC), and a first log-likelihood ratio (LLR) calculation unit.
[0024] Based on this hardware architecture, the decoding data flow of the dedicated accelerator consists of five main steps: (1) Error symptom transformation: The transformation unit receives the original error symptom and through sparse matrix-vector multiplication (MVM) with the transformation matrix Multiplying, generating the transformed error symptoms . Transformation Matrix It is pre-loaded into the sparse matrix buffer before real-time decoding is performed.
[0025] (2) Target value calculation: Error symptoms after transformation is sent to multiple hierarchical decoding units (HDUs), each of which is responsible for calculating the Candidate incorrect guesses The error pattern will be divided into left error and right error Each HDU first triggers the first error symptom incremental update unit to calculate the error symptom corresponding to the left error by performing sparse MVM and exclusive OR (XOR) operations. , expressed as , represents an arbitrary sparse matrix. Then, the error symptom corresponding to the left error is The blocks of the diagonal matrix in the decoupled check matrix Split into K parts Left Error Symptom , sent to K greedy decoding cores (GDC). Each GDC processes its corresponding block matrix and , and outputs a partial error ,in The outputs of all GDCs are combined into a complete left-hand error The first log-likelihood ratio (LLR) calculation unit then uses the complete error The weight corresponding to each error bit Calculate the final objective function value .
[0026] (3) Comparison: The comparison tree receives the objective function value from the HDU , and compare them in parallel to identify the solution with the minimum weight. The output is the best solution corresponding to the minimum weight error guess .
[0027] (4) Parameter update: The best solution output by the first comparison tree With the current minimum If the new solution reduces the minimum value, a parameter update is triggered and stored 、 and The execution controller monitors the enable signal to decide whether the decoding process should continue or terminate.
[0028] (5) Error pattern replacement: The replacement unit receives the error pattern and through sparse MVM with permutation matrix Multiply and calculate the final error pattern In order to speed up the decoding calculation process, before performing real-time decoding, the permutation matrix Also preloaded into the sparse matrix buffer.
[0029] In the embodiment, Figure 3 As shown in (a), in order to efficiently calculate the left error symptoms in the hierarchical greedy decoding algorithm , To represent the right error symptom, the present invention proposes a first error symptom incremental update unit with two key features: First, sparse MVM and XOR: by utilizing the right error The present invention converts the complex MVM process into a simple row selection. The columns are also sparse, and the present invention only selects non-zero elements to perform XOR operations with error symptoms, thereby accelerating the decoding process; second, incremental update: by storing and reusing the error bits identified in the previous iteration Corresponding left error symptoms , avoiding recalculation of the entire error symptom, thus significantly reducing redundant calculations.
[0030] In the embodiment, Figure 3 (b) in the figure shows an example of computing the second iteration using the first error symptom incremental update unit. First, sparse MVM is applied to compute , that is, in the second round of iteration The calculation result of . Sparse matrix The compressed format proposed by the present invention is used for storage. The format includes a sparse matrix table and a non-zero row index table. The sparse matrix table stores the column address and the number of non-zero elements in each column, while the non-zero row index table records the row index corresponding to these non-zero elements. In the example, the column index is input. , so directly get The first column of is retrieved and the row indices of the "1" in the first column (i.e., 1 and 4) are extracted from the non-zero row index table. This non-zero extraction prepares the way for the subsequent sparse XOR operation. Since any value remains unchanged when XORed with zero, only the non-zero elements need to be XORed. Next, performing the incremental update and sparse XOR operation yields the left error symptom. .from Retrieve the values at the selected row indices (i.e., 1 and 4) from the register file, XOR them with 1, and write the result back Register file. After all first error symptom incremental update units are completed, the error bits identified in the current iteration are updated. Corresponding Save to register file. This allows the reuse of residual error symptom computation data in subsequent iterations.
[0031] In the embodiment, Figure 4 The architecture and computational flow of the Greedy Decoding Core (GDC) are presented to effectively support Figure 2 Each GDC consists of three main components: the second error symptom incremental update unit, the second log-likelihood ratio (LLR) calculation unit and the second comparison tree. Divide into first candidate guess and the second candidate guess Two parts, each GDC receives part of the left error symptom and all first candidate guesses As input, multiple second error symptom incremental update units are calculated in parallel , using any matrix and the right side Once the left part is obtained , will and Splicing together to form the left error To evaluate the candidate left-side error The probability of occurrence, GDC calls the second LLR calculation unit to calculate the objective function value The LLR calculation unit is calculated by focusing on and The non-zero elements of the weight register file are obtained using their indices The corresponding weights are retrieved from the tree, and then the adder tree is applied to calculate the objective function values in parallel, thus performing this task efficiently. Finally, all the calculated objective function values are sent to the comparison tree, which identifies the minimum. The left error corresponding to this minimum is considered the best solution for the current iteration.
[0032] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A dedicated accelerator for hierarchical greedy decoding algorithm, characterized by: include: transform unit, decoding core, and permutation unit; The transformation unit is configured to receive an original error symptom and multiply it by a transformation matrix through sparse matrix-vector multiplication to generate a transformed error symptom; The decoding core is used to execute a hierarchical greedy decoding algorithm, divide the error pattern into left errors and right errors, and perform calculations including guessing right errors and decoding left errors in parallel based on the transformed error symptoms to obtain a new value of the error pattern; The replacement unit is used for receiving a new value of the error pattern and multiplying it with the replacement matrix through sparse matrix-vector multiplication to calculate a final error pattern.
2. The dedicated accelerator for the hierarchical greedy decoding algorithm according to claim 1, characterized in that: The decoding core includes: a plurality of hierarchical decoding units, a first comparison tree and an execution controller; A plurality of hierarchical decoding units are used to receive the transformed error symptoms to calculate the objective function value of each candidate error guess in parallel; The first comparison tree is used to receive the objective function values from the plurality of hierarchical decoding units and compare and identify the minimum objective function value, and store the minimum objective function value and the corresponding new value of the error mode after comparing with the current minimum objective function value; The execution controller is used to control the continuation or termination of the decoding program.
3. The dedicated accelerator for the hierarchical greedy decoding algorithm according to claim 2, characterized in that: Each hierarchical decoding unit includes: a first error symptom incremental update unit, a plurality of greedy decoding cores, and a first log-likelihood ratio calculation unit; The first error symptom incremental update unit is used to perform sparse matrix-vector multiplication and XOR operation to calculate the error symptom corresponding to the left error, and split the error symptom corresponding to the left error into K partial left error symptoms and send them to K greedy decoding cores respectively; The multiple greedy decoding cores are used to process K partial left error symptoms respectively, and splice the decoded partial left errors into a complete left error; The first log-likelihood ratio calculation unit is used to calculate the objective function value according to the complete left error.
4. The dedicated accelerator for the hierarchical greedy decoding algorithm according to claim 3, characterized in that: In the first error symptom incremental update unit, a sparse matrix is stored in a specific compression format, including a sparse matrix table and a non-zero row index table; during calculation, the row indices of non-zero elements are extracted from the sparse matrix table and the non-zero row index table according to the input column index, and an XOR operation is applied only to these non-zero elements; after performing the incremental update and the sparse XOR operation to obtain the left error symptom, the value at the row index selected by the current iteration is operated with a specific value and written back; and the left error symptom corresponding to the error bit identified in the current iteration is stored so that the residual error symptom calculation data can be reused in subsequent iterations.
5. The dedicated accelerator for hierarchical greedy decoding algorithm according to claim 3, characterized in that: Each greedy decoding core includes: a second error symptom incremental update unit, a second log-likelihood ratio calculation unit, and a second comparison tree; The second error symptom incremental updating unit is configured to receive a partial left error symptom, divide the left error to be decoded into a first candidate guess and a second candidate guess, calculate the second candidate guess based on the received partial left error symptom and the first candidate guess, and once the second candidate guess is obtained, concatenate the first candidate guess and the second candidate guess to form a left error; The second log-likelihood ratio calculation unit is used to calculate the objective function value according to the non-zero elements in the first candidate guess and the second candidate guess; The second comparison tree is used to identify the minimum value of the objective function value output by the second log-likelihood ratio calculation unit, and regard the left error corresponding to the minimum value as the best value of the current iteration and output it.
6. The dedicated accelerator for hierarchical greedy decoding algorithm according to claim 5, characterized in that: The objective function value is to minimize the number of non-zero values in the left error, and use their indexes to retrieve the corresponding weights from the weight register file, and then apply the adder tree to parallel calculation to evaluate the possibility of the left error.
7. The dedicated accelerator for hierarchical greedy decoding algorithm according to claim 5, characterized in that: In the second comparison tree, once it is found that the currently calculated objective function value is less than the existing minimum objective function value, the currently calculated objective function value is updated to the minimum objective function value and guessing is stopped.
8. The dedicated accelerator for hierarchical greedy decoding algorithm according to claim 1, characterized in that: In the first comparison tree, once it is found that the currently calculated objective function value is less than the existing minimum objective function value, the currently calculated objective function value is updated to the minimum objective function value and the iteration is stopped.
9. The dedicated accelerator for hierarchical greedy decoding algorithm according to claim 1, characterized in that: Before performing real-time decoding, the transformation matrix and permutation matrix are pre-loaded into a sparse matrix buffer.
Citation Information
Patent Citations
Quantum error correction decoding system and method, fault-tolerant quantum error correction system and chip
CN112988451A
Quantum error correction hardware decoder and chip
CN118863081A
Quantum error mitigation acceleration method and accelerator using sparsity in tensor product
CN119831067A
Accelerator and acceleration method based on Top-K sparse moment vector multiplication
CN120123629A
Sparse optimizatoins for a matrix accelerator architecture
US20210035258A1