A low-complexity high-speed LDPC decoder and control method, device for CCSDS near-earth application standard

CN120165704BActive Publication Date: 2026-09-15XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510235590.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2026-09-15
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

[0003]现有技术中,专利公开号为“CN115580309A”,名称为“一种提高译码器译码效率和吞吐量的LDPC译码器”的发明,提出了一种可以提高吞吐量和译码效率的CCSDS标准下的LDPC译码器,其针对“行重为2的循环矩阵需要循环两次译码才能完成更新,译码效率低”的问题,通过两路并行的操作仅在一次循环中实现了对行重为2的循环矩阵的伪后验概率信息的更新,提高了译码效率和吞吐量;该专利采用L-NMSA,其发明虽提高了译码效率和吞吐量,但没有利用原先每层码字已更新的信息,牺牲了纠错性能

Benefits of technology

1.本发明提供的LDPC译码器,通过增加输入缓冲模块、校验方程计算模块和输出模块,并改进译码处理方法,最终在低复杂度下,实现了更高吞吐量的译码器,LDPC译码器同时具备低功耗与灵活性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120165704B_ABST
    Figure CN120165704B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of satellite communication, and discloses a low-complexity high-speed LDPC decoder and control method and equipment facing CCSDS near-earth application standards; the LDPC decoder comprises an input buffer module, a channel initial information storage module, a control module, a check node external information storage module, a check node information updating module, a check equation calculation module, a variable node external information storage module, a variable node information updating module and an output module; the LDPC decoder effectively improves the utilization efficiency of storage resources by improving the decoding processing method and adopting overlapping double-frame processing; the LDPC decoder adopts a multi-factor correction approximate minimum sum algorithm and a tree structure minimum sum and approximate second minimum value generator, so that a great decrease in the realization complexity of the decoder is obtained with a slight loss in error correction performance; and a channel decoding processing equipment is used for realizing the low-complexity high-speed LDPC decoder facing the CCSDS near-earth application standards.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of satellite communication technology and relates to a low-complexity high-speed LDPC decoder and its control method and equipment for CCSDS near-ground application standard. Background Technology

[0002] Low-Density Parity-Check (LDPC) codes, as a high-performance error correction coding technique, feature high coding gain, low error flatness, and approximation of the Shannon limit, making them widely used in wireless communication, digital broadcasting, data storage, and deep space communication. The LDPC codes in the Consultative Committee for Space Data Systems (CCSDS) near-Earth application standard have high code rates, low required parity redundancy overhead, and coding gains of approximately 7 dB, making them highly favored in satellite communication. LDPC decoding typically employs soft-decision iterative decoding algorithms. In practical engineering applications, the Normalized Min-Sum Algorithm (NMSA) and Offset Min-Sum Algorithm (OMA) are specifically used. The LDPC decoder (OMSA) requires a large amount of data computation and storage for each iteration, resulting in high complexity. Furthermore, it requires multiple iterations to achieve excellent error correction performance. This means that implementing a high-speed and reliable LDPC decoder consumes a significant amount of hardware resources. Therefore, improving the resource utilization efficiency of this LDPC decoder and realizing a high-speed LDPC decoder with low complexity is an important problem that urgently needs to be solved in the field of satellite data transmission.

[0003] In the prior art, the invention with patent publication number "CN115580309A" entitled "An LDPC decoder for improving decoding efficiency and throughput" proposes an LDPC decoder under the CCSDS standard that can improve throughput and decoding efficiency. It addresses the problem that "a circular matrix with a row weight of 2 requires two decoding cycles to complete the update, resulting in low decoding efficiency." By using two parallel operations, it updates the pseudo-posterior probability information of the circular matrix with a row weight of 2 in only one cycle, thereby improving decoding efficiency and throughput. This patent uses L-NMSA. Although its invention improves decoding efficiency and throughput, it does not utilize the information that has already been updated in each layer of codewords, sacrificing error correction performance.

[0004] In the prior art, patent publication number "CN115664584A" entitled "A High-Efficiency LDPC Decoder for High-Speed ​​Satellite Links" improves the core processing modules (variable node and check node information update units) of the LDPC decoder. For the variable node update unit, the patent improves the utilization of FPGA slice resources by reducing the number of pipeline stages, while reducing decoding latency. For the check node update unit, the patent discards the second smallest value in the first comparison stage, which greatly reduces the implementation complexity of the circuit without significant performance loss. However, in order to reduce the bit width expansion of subtraction operations, the removal of subtraction operations increases the number of adders / subtractors in the circuit. At the same time, the space for operation optimization in this invention is limited. In the final implementation, the decoder still has high implementation complexity and resource consumption. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to propose a low-complexity, high-speed LDPC decoder and its control method and device for near-ground applications of the CCSDS standard. This LDPC decoder, by adding an input buffer module, a check equation calculation module, and an output module, adopts an improved partially parallel architecture and uses a multi-factor modified approximate minimum sum algorithm for the decoding algorithm, enabling the LDPC decoder to achieve higher throughput with low complexity, and possessing low power consumption and flexibility.

[0006] To achieve the above-mentioned technical objectives, the technical solution adopted by the present invention is as follows: In the first aspect, a low-complexity high-speed LDPC decoder for CCSDS near-ground application standard is provided. The LDPC decoder includes an input buffer module, a channel initial information storage module, a control module, a check node external information storage module, a check node information update module, a check equation calculation module, a variable node external information storage module, a variable node information update module, and an output module. The input buffer module is used to receive the channel initial information of the high-speed input data frame, and output it to the channel initial information storage module after channel number conversion. The high-speed input data frame includes odd data frames and even data frames. The input buffer module includes a first-level input buffer unit and a second-level input buffer unit, which are used to receive the input odd data frames and even data frames respectively, and perform channel number conversion. The channel initial information storage module receives and stores the channel initial information after the high-speed input data frame has undergone path number conversion, and finally outputs it to the variable node information update module. The channel initial information storage module includes a first-level channel initial information storage sub-module and a second-level channel initial information storage sub-module. The first-level channel initial information storage sub-module writes its internal data into the second-level channel initial information storage sub-module. The control module is used to coordinate the operation of different modules of the LDPC decoder; The external information storage module for the verification node is used to store the external information V2C (Variable Node to Check Node) passed from the variable node to the verification node and the hard decision information of the pseudo-posterior probability information of the variable node; The verification node information update module is used to calculate the external information passed from the verification node to the variable node; The verification equation calculation module is used to calculate the verification equation corresponding to the LDPC code verification matrix and to check whether the hard decision information of the pseudo-posterior probability information of the variable node satisfies the verification equation. The variable node external information storage module is used to store the external information C2V (Check Node to Variable Node) passed from the check node to the variable node. The variable node information update module is used to calculate the hard decision information of the external information passed from the variable node to the verification node and the pseudo-posterior probability information of the variable node. The LDPC decoder comprises an improved partially parallel architecture consisting of a channel initialization information storage module, a check node external information storage module, a check node information update module, a variable node information update module, a variable node external information storage module, and an output storage submodule within the output module. This improved partially parallel architecture divides the QC-LDPC code's check matrix according to the number of cyclic permutation matrices (CPMs), allocating independent memory to each CPM. Simultaneously, the improved partially parallel architecture... p row or p The extra-list information is stored at the same address in the corresponding memory; When the maximum throughput of the LDPC decoder is less than or equal to the rate of the high-speed input data frame, the improved partially parallel architecture adopts an overlapping double-frame processing strategy, wherein the row parallelism of the improved partially parallel architecture is... mp The column parallelism is np The first-level channel initial information storage submodule and the second-level channel initial information storage submodule respectively include n Each random access memory (RAM) includes a verification node external information storage module and a variable node external information storage module, respectively. m n w There are 1 RAM, each with an effective capacity of 2. L / p Qp, Q To quantize the number of bits; when the LDPC decoder's maximum throughput exceeds the rate of the high-speed input data frame, the improved partially parallel architecture adopts a non-overlapping double-frame processing strategy. Specifically, the improved partially parallel architecture merges the external information storage module for the verification node and the external information storage module for the variable node into a single external information storage module. This external information storage module only includes... m n w There are 1 RAM, each with an effective capacity of 1. L / p Qp .

[0007] Furthermore, the control module includes a channel initial information storage module read control unit, a routing unit, a check node external information storage module read control unit, a variable node external information storage module read control unit, a variable node information update module external information output enable control unit, and an output storage submodule write control unit. The channel initial information storage module read control unit is used to control the reading of the channel initial information stored in the channel initial information storage module; The routing unit is used to control the data write addresses of the external information storage module for the verification node and the external information storage module for the variable node; The read control unit of the external information storage module of the verification node is used to control the reading of the external information of the verification node stored in the external information storage module of the verification node; The variable node external information storage module read control unit is used to control the reading of variable node external information stored in the variable node external information storage module; The variable node information update module external information output enable control unit is used to control the output enable of external information of the variable node information update module. The output storage submodule write control unit is used to control the data writing of the output storage submodule.

[0008] Furthermore, the external information storage module for the verification node supports additional storage of hard decision information, including pseudo-posterior probability information of variable nodes. When the LDPC decoder is compatible with the early termination iteration operation mode, the external information storage module for the verification node selects a compact storage strategy or a distributed RAM storage strategy. The compact storage strategy stores the hard decision information and the external information passed from the variable nodes to the verification node in the same memory. The external information storage module for the verification node consists of... m n w It consists of 1 RAM, each RAM having an effective capacity of 2. L / p Q ( p +1).

[0009] Furthermore, the decoding algorithm of the LDPC decoder adopts the Improved Normalized Approximate Min-Sum Algorithm (IAMSA) with multi-factor correction for flooding scheduling. The IAMSA algorithm verifies the node information update formula as follows:

[0010] in, r ji Indicates the first j The verification node is passed to the first... i External information of each variable node Q j \ i Indicates except the first i Outside of the variable node, with the first variable node j The set of other variable nodes connected to a check node. sign Represents a symbolic function. q ij Indicates the first i The variable node is passed to the first j The external information of each verification node, where max represents the function for finding the maximum value. α 1 and α 2 represents the normalization factor, and amin represents the function for finding the minimum and approximate second minimum values. β 1 and β 2 represents the offset factor.

[0011] Furthermore, both the verification node information update module and the variable node information update module are implemented using a multi-stage pipeline. This multi-stage pipeline employs a tree-structured minimum and approximate second-minimum value generator (amin) to calculate the minimum and approximate second-minimum values ​​(M1AM2VG). TS The minimum and near-second minimum values ​​are generated in the tree structure. w r The minimum and near-second minimum values ​​of the number need to be inserted. log2w r Automated production line.

[0012] Furthermore, the minimum and near-minimum value generators of the tree structure are divided into a first stage and a second stage during computation; in the first stage, the input data is divided into... N In the first stage, each set of data is divided into pipeline stages according to a tree-structured minimum value generator, and the minimum value of each set of data is obtained by pairwise comparison. In the second stage, the set of minimum values ​​of each set of data is divided into pipeline stages according to a tree-structured minimum and second minimum value generator, and the minimum value and the approximate second minimum value of all input data are obtained by pairwise comparison.

[0013] Secondly, a control method for a low-complexity, high-speed LDPC decoder for near-ground applications of the CCSDS standard includes the following steps: S1: The first-level channel initial information storage submodule of the high-speed input data frame is input to the channel initial information storage module; S2: At the start of the first iteration of each data frame in the high-speed input data frame, the data stored in the first-level channel initial information storage submodule in step S1 is read and written to the second-level channel initial information storage submodule and then written to the external information storage module of the verification node after passing through the multiplexer. S3: When the output data of the external information storage module of the verification node described in step S2 is valid, the verification node information update module and the verification equation calculation module start the verification node information update work. The verification node information update module outputs the calculated external information to the external information storage module of the variable node, and the verification equation calculation module outputs the judgment flag of whether the codeword satisfies the verification equation to the control module. S4: When the output data of the variable node external information storage module in step S3 and the second-level channel initial information storage submodule in step S2 are valid, the variable node information update module starts the variable node information update work. The second-level channel initial information storage submodule outputs the original data frame corresponding to the data frame that is undergoing decoding iteration to the variable node information update module. The variable node information update module outputs the calculated external information to the verification node external information storage module. S5: The verification node information update in step S3 and the variable node information update in step S4 together constitute the iteration of the input data frame. When the input data frame reaches the maximum number of iterations or meets the early termination iteration criterion, the data frame terminates the iteration at the end of the current iteration cycle. Hard decision information. c Valid information is written to the output storage submodule under the control of the control module; when the data frame iteration has not terminated, the variable node passes the external information V2C and hard decision information to the verification node. cFirst, it passes through a multiplexer, and then is written to the external information storage module of the verification node in step S3 under the control of the routing unit. S6: Read the hard decision information from the output storage submodule described in step S5 in the order of data frame columns. c It converts the stored out-of-order bit data into decoded output in the correct order and removes invalid data.

[0014] Thirdly, a channel decoding processing device, based on the LDPC decoder, wherein the channel decoding processing device is used to implement the LDPC decoder.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The LDPC decoder provided by this invention, by adding an input buffer module, a check equation calculation module and an output module, and improving the decoding processing method, ultimately achieves a higher throughput decoder with low complexity. The LDPC decoder also has low power consumption and flexibility.

[0016] 2. The LDPC decoder provided by this invention adopts overlapping double-frame processing. By optimizing the time utilization of the check node and variable node information update modules, parallel processing of two frames of data is achieved. With only an increase in information storage resources and a small amount of control logic resources, and without an increase in computing resources, the decoder throughput is significantly improved with lower resource overhead.

[0017] 3. The LDPC decoder provided by this invention adopts an improved partially parallel architecture. The information update units of the parity check node and the variable node process the row and column groups of the parity check matrix in a time-division multiplexing manner. The improved partially parallel architecture further enhances the parallel processing capability of the partially parallel decoder by mapping the external information of multiple rows (or multiple columns) of the CPM to the same storage address, while effectively avoiding storage access conflicts.

[0018] 4. The LDPC decoder provided by this invention employs an efficient storage strategy, specifically including: a compact storage strategy that stores hard decision information and external information (V2C) transmitted from variable nodes to check nodes in the same memory to save storage and counter resources; and a storage strategy based on the minimum block RAM of digital circuits, which optimizes the bit width design of the block RAM to match as closely as possible to the maximum bit width of the minimum size block RAM or an integer multiple thereof, thereby effectively improving the utilization efficiency of storage resources. (Distributed RAM storage strategy, listed alongside the compact storage strategy) For FPGAs, when chip block RAM resources are scarce or internal utilization is low, the check node external information storage module will use distributed RAM to store hard decision information.

[0019] 5. The LDPC decoder provided by this invention employs a flooding-scheduling, multi-factor correction approximate minimum sum algorithm. During check node information updates, the minimum value function only calculates the accurate minimum and the approximate second smallest value to reduce the decoder's implementation complexity. Simultaneously, different normalization and offset factors are added to compensate for the minimum and approximate second smallest values, minimizing the decoder's error correction performance loss. In the decoder implementation, this invention uses a tree-structured minimum and approximate second smallest value generator. By reducing the number of comparators and selectors, the decoder's hardware resource overhead is reduced, resulting in a significant decrease in decoder implementation complexity at the cost of minimal error correction performance loss.

[0020] In summary, the LDPC decoder provided by this invention, through improvements such as an improved partially parallel architecture, overlapping double-frame processing, a multi-factor corrected approximate minimum sum algorithm, and a tree-structured minimum and approximate second-minimum value generator, can achieve higher throughput while reducing complexity, and incurs only a slight loss in error correction performance, thus saving costs. Attached Figure Description

[0021] Figure 1 This is a block diagram illustrating the implementation of a low-complexity, high-speed LDPC decoder for near-ground applications of the CCSDS standard, provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of a control method for a low-complexity, high-speed LDPC decoder for near-ground applications of the CCSDS standard, provided by an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure of a conventional partially parallel architecture LDPC decoder provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of the improved partially parallel architecture LDPC decoder provided in the embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of the LDPC decoder for overlapping double-frame processing provided in an embodiment of the present invention; Figure 6 This is a schematic diagram illustrating the principle of overlapping double-frame processing in the LDPC decoder provided in this embodiment of the invention; Figure 7 This is a schematic diagram of the structure of an LDPC decoder that supports early termination of iteration provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the pipeline structure of the verification node information update unit provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of the pipeline structure of the variable node information update unit provided in an embodiment of the present invention; Figure 10 This is a schematic diagram of the minimum and near-second minimum value generator for a tree structure provided in an embodiment of the present invention; Figure 11 This is the minimum value generator 2-MVG1 with two inputs and a single minimum value output in the tree structure minimum and near-minimum value generator provided in the embodiments of the present invention; Figure 12 This is the minimum value generator 2-MVG2 with two inputs and two minimum value outputs in the minimum and near-minimum value generator of the tree structure provided in the embodiments of the present invention; Figure 13 This is a schematic diagram of the structure of the tree structure minimum and approximate second minimum value generator with 32 input data provided in the embodiment of the present invention; Figure 14 This is a bit error rate curve of the (8160, 7136) LDPC code provided in an embodiment of the present invention. Detailed Implementation

[0022] The following is combined with Figures 1 to 14 The present invention will be further described in detail with reference to the embodiments.

[0023] Firstly, a low-complexity, high-speed LDPC decoder for near-ground applications of the CCSDS standard is provided. The LDPC decoder includes an input buffer module, a channel initialization information storage module, a control module, a check node external information storage module, a check node information update module, a check equation calculation module, a variable node external information storage module, a variable node information update module, and an output module. The input buffer module is used to receive the channel initial information of the high-speed input data frame, and output it to the channel initial information storage module after channel number conversion. The high-speed input data frame includes odd data frames and even data frames, which can improve the data throughput of the system. The input buffer module includes a first-level input buffer unit and a second-level input buffer unit, which are used to receive the input odd data frames and even data frames respectively, and perform channel number conversion. The channel initial information storage module receives and stores the channel initial information after the high-speed input data frame has undergone path conversion, and finally outputs it to the variable node information update module. The channel initial information storage module includes a first-level channel initial information storage submodule and a second-level channel initial information storage submodule, which are used for initializing the input data frame and providing the original input data frame required by the variable node information update module for each iteration, respectively. The first-level channel initial information storage submodule writes its internal data into the second-level channel initial information storage submodule. The control module is used to coordinate the operation of different modules of the LDPC decoder to ensure that the decoding process is performed correctly. The external information storage module of the verification node is used to store the external information V2C passed from the variable node to the verification node and the hard decision information of the pseudo posterior probability information of the variable node. The verification node information update module is used to calculate the external information passed from the verification node to the variable node; The verification equation calculation module is used to calculate the verification equation corresponding to the LDPC code verification matrix and to check whether the hard decision information of the pseudo-posterior probability information of the variable node satisfies the verification equation. The variable node external information storage module is used to store the external information C2V passed from the verification node to the variable node; The variable node information update module is used to calculate the hard decision information of the external information passed from the variable node to the verification node and the pseudo-posterior probability information of the variable node. The output module includes an output storage submodule, two selection merging units, a deletion unit, and a first-in-first-out (FIFO) data buffer. The output module processes and outputs the hard decision information of the pseudo-posterior probability information of the variable node in the order of the output storage submodule, the two selection merging units, the deletion unit, and the FIFO data buffer.

[0024] Specifically, the information stored in the first-level channel initial information storage submodule is written into the check node external information storage module and the second-level channel initial information storage submodule at the start of the first iteration of the corresponding input data frame. Simultaneously, for the CCSDS near-ground communication system, the length of the LDPC decoder input data frame is 8160, while the standard LDPC code length is 8176. Before the data frame iterative decoding, the first-level channel initial information storage submodule will pad the 18 all-zero bits added before the data frame information bit encoding and delete the last two bits. For these 18 all-zero bits, the first-level channel initial information storage submodule initializes their corresponding information to two's complement data with a sign bit of 0 and all other bits of 1, ensuring that the check node information update module ignores these bits and improves error correction performance.

[0025] Furthermore, the control module includes a channel initial information storage module read control unit, a routing unit, a check node external information storage module read control unit, a variable node external information storage module read control unit, a variable node information update module external information output enable control unit, and an output storage submodule write control unit. The channel initial information storage module read control unit is used to control the reading of the channel initial information stored in the channel initial information storage module; The routing unit is used to control the data write addresses of the external information storage module for the verification node and the external information storage module for the variable node; The read control unit of the external information storage module of the verification node is used to control the reading of the external information of the verification node stored in the external information storage module of the verification node; The variable node external information storage module read control unit is used to control the reading of variable node external information stored in the variable node external information storage module; The variable node information update module external information output enable control unit is used to control the output enable of external information of the variable node information update module. The output storage submodule write control unit is used to control the data writing of the output storage submodule.

[0026] Reference Figure 1 The control module of this invention includes a channel initial information storage module read control unit, a routing unit, a check node external information storage module read control unit, a variable node external information storage module read control unit, a variable node information update module external information output enable control unit, and an output storage submodule write control unit. These are respectively used to control the reading of channel initial information stored in the channel initial information storage module, the data writing address of the check node and variable node external information storage modules, the reading of check node and variable node external information stored in the check node and variable node external information storage modules, the output enable of external information of the variable node information update module, and the data writing of the output storage submodule, so as to ensure the normal operation of the decoder while improving the throughput of the decoder.

[0027] Specifically, for the first-level channel initial information storage submodule, when the stored information is valid and the read data frame will not conflict with the data frame being processed by the decoder or the decoder is idle, the channel initial information storage module read control unit ensures that the channel initial information it stores can be read immediately; for the second-level channel initial information storage submodule, the channel initial information storage module read control unit ensures that the channel initial information it stores is read when the variable node information update module works in each iteration of the corresponding data frame.

[0028] Specifically, the routing unit effectively avoids memory access conflicts by storing information from the same group of columns or rows in the parity check matrix at the same address location; the read control unit of the external information storage module for the parity check node and variable node improves the decoder throughput by minimizing the transmission latency of the external information storage module for the parity check node and variable node; the external information output enable control unit of the variable node information update module invalidates the external information and hard decision information output by the variable node to the parity check node during the last iteration of the data frame, thereby avoiding conflicts between the new input data frame and the data frame that is about to complete the iteration; the write control unit of the output storage submodule ensures that the valid hard decision information output by the variable node information update module is written to the output storage submodule in a timely manner when the data frame reaches the maximum number of iterations or meets the early termination iteration criterion.

[0029] Furthermore, the LDPC decoder comprises an improved partially parallel architecture consisting of a channel initial information storage module, a check node external information storage module, a check node information update module, a variable node information update module, a variable node external information storage module, and an output storage submodule in the output module. This improved partially parallel architecture divides the QC-LDPC code's check matrix according to the number of cyclic permutation matrices (CPMs) within it, and allocates independent memory to each CPM. Simultaneously, the improved partially parallel architecture stores the CPMs... p row or p The extra-list information is stored at the same address in the corresponding memory; in, p This represents the extended parallelism of CPM. For regular QC-LDPC codes, their parity-check matrix is ​​determined by... m n indivual L Composed of submatrices of order, with row and column weights of size . w .

[0030] Reference Figure 4 The decoder of this invention employs an improved partially parallel architecture, which is based on... Figure 3 This improved partially parallel architecture is derived from the traditional model. It divides the parity-check matrix of the QC-LDPC code according to the number of check millimeters (CPMs), and uses an equal number of memories to store the extrinsic information of each CPM. Furthermore, the improved partially parallel architecture further improves the extrinsic information of each CPM by... p row or p Out-of-line information is mapped to the same storage address, which further enhances the parallel processing capability of some parallel decoders, while effectively avoiding storage access conflicts. R This indicates the QC-LDPC code rate.

[0031] Specifically, p This represents the extended parallelism of CPM. For regular QC-LDPC codes, their parity-check matrix is ​​determined by... m n indivual L Composed of submatrices of order, with row and column weights of size . w The line parallelism of the decoder in a traditional partially parallel architecture is... m The column parallelism is n The channel initial information storage module includes 2 n There are 1 RAM, each with an effective capacity of 1. L Q , Q To quantize the number of bits, the off-node information storage unit includes m n w There are 1 RAM, each with an effective capacity of 1. L Q After the improvement, the decoder's line parallelism is... mp The column parallelism is np The channel initial information storage module includes 2 n There are 1 RAM, each with an effective capacity of 1. L / p Qp The external information storage unit includes m n w There are 1 RAM, each with an effective capacity of 1. L / p Qp .

[0032] Furthermore, when the maximum throughput of the LDPC decoder is less than or equal to the rate of the high-speed input data frame, the improved partially parallel architecture adopts an overlapping double-frame processing strategy, wherein the row parallelism of the improved partially parallel architecture is... mp The column parallelism is np The first-level channel initial information storage submodule and the second-level channel initial information storage submodule respectively include n Each random access memory (RAM) includes a verification node external information storage module and a variable node external information storage module, respectively. m n w There are 1 RAM, each with an effective capacity of 2. L / p Qp, Q To quantize the number of bits; when the LDPC decoder's maximum throughput exceeds the rate of the high-speed input data frame, the improved partially parallel architecture adopts a non-overlapping double-frame processing strategy. Specifically, the improved partially parallel architecture merges the external information storage module for the verification node and the external information storage module for the variable node into a single external information storage module. This external information storage module only includes... m n w There are 1 RAM, each with an effective capacity of 1. L / p Qp .

[0033] Reference Figure 5 and Figure 6 When the LDPC decoder employs an overlapped double-frame processing strategy, it achieves parallel processing of two frames of data by optimizing the time utilization of the check node and variable node information update modules. This results in a significant increase in decoder throughput with lower resource overhead, requiring only an increase in information storage resources and a small amount of control logic resources, without increasing computational resources. Theoretically, using overlapped double-frame processing can double the decoder throughput.

[0034] Specifically, in the LDPC decoder employing overlapping double-frame processing, apart from the input buffer module and the FIFO in the output module, all other functional units with storage capabilities utilize a depth of 2 times. L / p The memory is thus divided into two regions, front and back, for storing odd-numbered data frames and even-numbered data frames respectively, thereby enabling efficient dual-frame parallel processing.

[0035] Specifically, when L Unable to be p When divisible, the depth is 2 times. L / p The memory of the first L / p The remaining addresses rem This information will repeat the first address of the partition. rem This information simplifies access to the information; the next partition of the memory and its depth are... L / p The memory employs the same strategy; among which, rem express L / p The remainder.

[0036] Furthermore, the external information storage module for the verification node supports additional storage of hard decision information, including pseudo-posterior probability information of variable nodes. When the LDPC decoder is compatible with the early termination iteration operation mode, the external information storage module for the verification node selects a compact storage strategy or a distributed RAM storage strategy. The compact storage strategy stores the hard decision information and the external information passed from the variable nodes to the verification node in the same memory. The external information storage module for the verification node consists of... m n w It consists of 1 RAM, each RAM having an effective capacity of 2. L / p Q ( p +1).

[0037] Specifically, when adopting a compact storage strategy, refer to Figure 7 The present invention's external information storage module for verification nodes supports additional storage of hard decision information, including pseudo-posterior probability information of variable nodes. Specifically, when the verification equation calculation module (the Parity Check Update (PCU) is the processing unit within the verification equation calculation module) and the verification node information update module work synchronously, the hard decision information input by the corresponding internal unit and the external information passed from the variable node to the verification node are in the same position in the verification matrix. Therefore, the external information storage module for verification nodes can adopt a compact and efficient storage strategy, storing the hard decision information and the external information passed from the variable node to the verification node in the same memory to save storage and counter resources. Under this storage strategy, the external information storage module for verification nodes consists of... m n w It consists of 1 RAM, each RAM having an effective capacity of 2. L / p Q ( p +1).

[0038] Specifically, when adopting a distributed RAM storage strategy, for FPGAs, in cases where chip BRAM resources are scarce or internal utilization is low, distributed RAM will be used to store hard decision information.

[0039] Furthermore, the decoding algorithm of the LDPC decoder, through the cooperation of the check node external information storage module, check node information update module, check equation calculation module, variable node external information storage module, and variable node information update module, implements the IAMSA (Approximate Minimum Sum Algorithm) with multi-factor correction using flooding scheduling. This algorithm is an improvement on the Normalized Minimum Sum Algorithm (NMSA). First, in order to further reduce the computational complexity of the check node information update operation, this method proposes the Approximate Minimum Sum Algorithm (AMSA), using a tree-structured minimum and near-minimum value generator to transform the solution of the minimum value function during check node information update into the solution of the accurate minimum and near-minimum values. In the check node information update module, NAMSA only calculates the accurate minimum and near-minimum values ​​of the minimum function during check node information update by multiplying them by the same normalization factor for compensation. IAMSA adds different normalization factors and offset factors to compensate for the minimum and near-minimum values. The check node information update formula of the IAMSA algorithm is as follows:

[0040] in, r ji Indicates the first j The verification node is passed to the first... i External information of each variable node Q j \ i Indicates except the first i Outside of the variable node, with the first variable node j The set of other variable nodes connected to a check node. sign Represents a symbolic function. q ij Indicates the first i The variable node is passed to the first j The external information of each verification node, where max represents the function for finding the maximum value. α 1 and α 2 represents the normalization factor, and amin represents the function for finding the minimum and approximate second minimum values. β 1 and β 2 represents the offset factor.

[0041] Reference Figure 8 and Figure 9In this invention, both the check node information update module and the variable node information update module are implemented using a multi-stage pipeline. By reducing the critical path latency of the core processing module, the maximum clock frequency that the decoder can support is increased, thereby improving the decoding throughput. Specifically, the decoder algorithm of this invention uses F-IAMSA, which is an improvement upon NAMSA. NAMSA, when updating check node information, only seeks the exact minimum value and the approximate second smallest value of the minimum function to reduce the implementation complexity of the decoder, while adding the same normalization factor for compensation. IAMSA, on the other hand, adds different normalization factors and offset factors to compensate for the minimum value and the approximate second smallest value, thereby reducing the loss of the decoder's error correction performance.

[0042] The verification node information update module includes: mp Each processing unit has a pipeline structure divided into two parts: the upper part is responsible for computation and processing. w r The sign bit of the input data V2C is used, while the lower half is used for computation and processing. w r The absolute value of each input data V2C, w r and w c These represent the row weight and column weight of the parity check matrix, respectively. The processing unit is divided into a 5-stage pipeline based on its computational function. The first stage pipeline is responsible for extracting the sign bit and absolute value of the V2C; the upper part of the second stage pipeline calculates the sum of all sign bit data (sign) through an XOR operation. w r and delay w r The first half is the sign bit, while the second half is used for calculation. w r The minimum value among absolute values, min 1st Approximate second smallest value amin 2nd And the index idx corresponding to the minimum value; the first half of the third-level pipeline pairs sign. w r and w r Perform an XOR operation on each of the sign bits to obtain... w r The sum of all sign bits except itself, sign( w r-1), the lower half calculates the minimum value and the approximate second smallest value after compensation, and delays the minimum value index; when it reaches the 4th stage pipeline, the operation of the sign bit in the upper half has been completed, and it only needs to delay and wait for the absolute value part to be completed before selection and merging. The lower half truncates the result of the previous stage and continues to delay the minimum value index. When truncating, it needs to be rounded; the 5th stage pipeline will select and merge the processed sign bit and absolute value data. If the sign bit index is equal to the minimum value index, the absolute value bit output is the result after second smallest value compensation. Otherwise, the absolute value bit output is the result after minimum value compensation; during merging, the output of the sign bit operation is used as the high bit of the original code output, and the output of the absolute value bit operation is used as the low bit of the original code output; the calculation of converting the original code (for the original code in the previous step) to the two's complement is put into the variable node information update module for processing.

[0043] Furthermore, both the verification node information update module and the variable node information update module are implemented using a multi-stage pipeline. The operations in the second-stage pipeline are refined by inserting more pipeline stages. This multi-stage pipeline uses a tree-structured minimum and approximate second-minimum value generator (amin) to calculate the minimum and approximate second-minimum values ​​(M1AM2VG). TS The minimum and near-second minimum values ​​are generated in the tree structure. w r The minimum and near-second minimum values ​​of the number need to be inserted. log2 w r Automated production line.

[0044] Specifically, the verification node information update module includes mp Each processing unit has a pipeline structure divided into two parts: the upper part is responsible for computation and processing. w r The sign bit of the input data V2C is used, while the lower half is used for computation and processing. w r The absolute value of each input data V2C, w r and w c These represent the row weight and column weight of the parity check matrix, respectively. The processing unit is divided into a 5-stage pipeline based on its computational function. The first stage pipeline is responsible for extracting the sign bit and absolute value of the V2C; the upper part of the second stage pipeline calculates the sum of all sign bit data (sign) through an XOR operation. w r and delay wr The first half is the sign bit, while the second half is used for calculation. w r The minimum value among absolute values, min 1st Approximate second smallest value amin 2nd And the index idx corresponding to the minimum value; the first half of the third-level pipeline pairs sign. w r and w r Perform an XOR operation on each of the sign bits to obtain... w r The sum of all sign bits except itself, sign( w r -1), the lower half calculates the minimum value and the approximate second smallest value after compensation, and delays the minimum value index; by the fourth stage of the pipeline, the operation of the sign bit in the upper half has been completed, and only needs to be delayed to wait for the absolute value part to be completed before selection and merging. The lower half truncates the result of the previous stage and continues to delay the minimum value index. Rounding is required during truncation; the fifth stage pipeline will select and merge the processed sign bit and absolute value data. If the sign bit index is equal to the minimum value index, the absolute value bit output is the result after compensation for the second smallest value. Otherwise, the absolute value bit output is the result after compensation for the minimum value. During merging, the output of the sign bit operation is used as the high bit of the original code output, and the output of the absolute value bit operation is used as the low bit of the original code output.

[0045] Preferably, considering the simplicity of the pipeline structure of the Variable Node Update (VNU), in order to balance the pipeline delay and computational resources of the Variable Node Update (CNU) and Check Node Update (CNU), the calculation of converting the original code to the complement code is placed in the Variable Node Update (VNU) for processing.

[0046] Preferably, in the second-stage pipeline of the verification node information update unit, the lower half of the computational function is relatively complex, and the critical path delay is large. To further optimize the processing speed of the verification node information update unit, this part of the operation is refined by inserting more stages of the pipeline, and a tree-structured minimum and near-minimum value generator is used to find the minimum and near-minimum values. The tree-structured minimum and near-minimum value generator is described below. Figure 10 ,beg w r The minimum and near-second minimum values ​​of the number need to be inserted. log2 w r The system employs a multi-stage pipeline. Furthermore, the tree-structured minimum and near-minimum generators reduce the hardware resource overhead of the decoder by decreasing the number of comparators and selectors, trading a small loss in error correction performance for a significant reduction in decoder implementation complexity.

[0047] The variable node information update module includes: np Each processing unit employs a 4-stage pipeline structure, using... l i This represents the initial channel information corresponding to the current variable node, expressed as... w c The external information C2V passed from each verification node to the variable node represents... r ji , , Indicates the relationship with the first i The set of other check nodes connected to each variable node, the C2V data is in original code form; the first-stage pipeline first converts the external information C2V into two's complement and delays... l i The second-stage pipeline uses an array of adders to calculate the sum of all input data; the third-stage pipeline uses an array of subtractors to calculate the sum and then subtracts the sum from the input data. w c The difference between C2V; the fourth-stage pipeline needs to truncate the intermediate calculation results to obtain the final hard decision information. c and w c One V2C; Optionally, for consecutive addition operations of multiple data, the second-level pipeline can be refined; Preferably, when the decoder's operating frequency allows, the number of pipeline stages for the check node and variable node information update units can be reduced.

[0048] Specifically, the variable node information update module includes np Each processing unit employs a 4-stage pipeline structure, using... l i This represents the initial channel information corresponding to the current variable node, expressed as... w c The external information C2V passed from each verification node to the variable node represents... r ji , , Indicates the relationship with the first i The C2V data is a set of other check nodes connected to a variable node, and its form is in original code. Addition and subtraction operations are more convenient using two's complement; therefore, the first-stage pipeline first converts the external information C2V to two's complement and delays... l iThe second-stage pipeline uses an array of adders to calculate the sum of all input data; the third-stage pipeline uses an array of subtractors to calculate the sum and then subtracts the sum from the input data. w c The difference between C2V values; to prevent precision loss, the bit width of the intermediate calculation result must be wider than the bit width of the input and output data. Therefore, the fourth-stage pipeline needs to truncate the intermediate calculation result to obtain the hard decision information. c and w c One V2C; Optionally, for the continuous addition of multiple data, the second-level pipeline can be refined, that is, multiple pipelines can be inserted into the second-level pipeline; Preferably, when the decoder's operating frequency allows, the number of pipeline stages for the check node and variable node information update units can be reduced to decrease the number of registers used. For FPGAs, evenly distributing LUT resources and FF resources helps improve the utilization efficiency of FPGA slice resources.

[0049] Furthermore, the minimum and near-minimum value generators of the tree structure are divided into a first stage and a second stage during computation; in the first stage, the input data is divided into... N In the first stage, each set of data is divided into pipeline stages according to a tree-structured minimum value generator, and the minimum value of each set of data is obtained by pairwise comparison. In the second stage, the set of minimum values ​​of each set of data is divided into pipeline stages according to a tree-structured minimum and second minimum value generator, and the minimum value and the approximate second minimum value of all input data are obtained by pairwise comparison.

[0050] The index of the minimum value is generated by a minimum value index generator. 2-MVG1 represents a minimum value generator with 2 inputs and a single minimum value output, and 2-MVG2 represents a minimum value generator with 2 inputs and 2 minimum value outputs; both will output a comparison result. cp The set of comparison results CP It will be used as input to the minimum index generator. For 2 k Given 1 input data, the minimum and second minimum value generators in the tree structure require 2... k + N -3 two-input comparators and 2 k +2 N - 4 two-input selectors, excluding the two-input selector required by the minimum value index generator; among them, N If it is a power of 2, N =2 k Then the minimum and near-minimum value generator of the tree structure generates the exact minimum and second-minimum values.

[0051] The verification equation calculation module and the verification node information update module are executed in parallel, including mp Each processing unit corresponds to a check equation and calculates a syndrome component. A syndrome component that is all zero indicates a correct decoding result. The calculation of the syndrome component involves finding... w r The XOR sum of hard decision information; preferably, a counting operation is used to determine whether the adjoint vector is a zero vector; optionally, for the XOR sum of multiple data, a multi-stage pipeline can be inserted.

[0052] Reference Figure 7 The equation verification calculation module and the node verification information update module are executed in parallel, including... mp Each processing unit corresponds to a check equation and calculates a syndrome component. When all syndrome components are zero (i.e., the syndrome vector is zero), the decoding result is correct. The calculation of the syndrome component involves finding... w r The XOR sum of hard decision information.

[0053] Preferably, in LDPC decoding, the length of the syndrome vector is usually quite long, so a counting operation is used to determine whether the syndrome vector is a zero vector in order to save storage resources.

[0054] Optionally, for the XOR sum of multiple data, multi-stage pipelines can also be inserted to reduce the critical path latency of the module.

[0055] The output module is responsible for converting the stored out-of-order bit data into decoded output in the correct order. The bit data corresponding to each sub-matrix block is sequential. In decoders using regular QC-LDPC codes, when... L Cannot be p During integer division, the deletion cell takes effect to remove invalid data.

[0056] Secondly, referring to Figure 2 A control method for a low-complexity, high-speed LDPC decoder for near-ground applications of the CCSDS standard includes the following steps: S1, Data Input: High-speed input data frames are input to the first-level channel initial information storage submodule of the channel initial information storage module. If the number of parallel input data frames is higher than the extended parallelism of the decoder, the input data frames must first be converted through the input buffer module before being input to the first-level channel initial information storage submodule of the channel initial information storage module. S2, Initialization: The decoder uses overlapping double-frame processing to improve throughput by 2 times. Except for the input buffer module and the FIFO in the output module, all other modules with storage functions use a depth of 2 times. L / p The memory is divided into two partitions, one for storing odd-numbered data frames and the other for storing even-numbered data frames. The first-level channel initial information storage submodule stores the odd-numbered data frames from the input data frames in column order. n The preceding partition of the RAM block address space stores even-numbered data frames in the following partition; at the beginning of the first iteration of each data frame in the high-speed input data frame, the data stored in the first-level channel initial information storage submodule in step S1 is read and written to the second-level channel initial information storage submodule and written to the external information storage module of the check node after passing through the multiplexer; the decoder iteration count is initialized to 0, and the maximum iteration count is determined according to the error correction performance requirements; S3, Verification Node Update and Verification Equation Calculation: When the output data of the verification node external information storage module described in step S2 is valid, the verification node information update module and the verification equation calculation module begin the verification node information update process. The verification node information update module outputs the calculated external information C2V to the variable node external information storage module, and the verification equation calculation module outputs the judgment flag indicating whether the codeword satisfies the verification equation to the control module. The verification node information update module inputs the external information V2C passed from the variable node to the verification node to calculate the external information C2V passed from the verification node to the variable node. C2V is then written to the variable node external information storage module under the control of the routing unit. The verification equation calculation module inputs hard decision information. c , used to calculate the syndrome to check whether the codeword satisfies the check equation; S4, Variable Node Update and Pseudo-Posterior Probability Information Calculation and Decision: When the output data of the variable node external information storage module in step S3 and the second-level channel initial information storage submodule in step S2 are valid, the variable node information update module starts the variable node information update work. The second-level channel initial information storage submodule outputs the original data frame corresponding to the data frame undergoing decoding iteration to the variable node information update module. The variable node information update module outputs the calculated external information V2C to the check node external information storage module. The variable node information update module inputs the external information C2V passed from the check node to the variable node and the channel initial information of the data frame, and is used to calculate the hard decision information of the external information V2C passed from the variable node to the check node and the pseudo-posterior probability information of the variable node. c After updating the variable nodes and calculating and deciding the pseudo-posterior probability information, the decoder iteration count is incremented by 1. S5, Termination Decision: The update of the verification node information in step S3 and the update of the variable node information in step S4 together constitute the iteration of the input data frame. When the input data frame reaches the maximum number of iterations or meets the early termination iteration criterion, the data frame terminates the iteration at the end of the current iteration cycle. Hard decision information. cValid information is written to the output storage submodule under the control of the control module; when the data frame iteration has not terminated, V2C and c First, it passes through a multiplexer, and then is written to the external information storage module of the verification node in step S3 under the control of the routing unit. S6, Decoding result processing and output: Read the hard decision information from the output storage submodule described in step S5 in the order of data frame columns. c It converts the stored out-of-order bit data into decoded output in the correct order and removes invalid data.

[0057] when L Unable to be p When divisible, the depth is 2 times. L / p The memory of the first L / p The remaining addresses rem This information will repeat the first address of the partition. rem This information. Similarly, the next partition of the memory and its depth are... L / p The memory uses the same strategy; among them, rem express L / p The remainder.

[0058] The decoder is compatible with both fixed iteration and early termination iteration operation modes. In fixed iteration mode, the storage and computation modules associated with early termination iteration mode will stop working. If the decoder only needs to support fixed iteration mode, the storage and computation modules associated with early termination iteration mode will be removed.

[0059] Reference Figure 7 The decoder provided by this invention is compatible with both fixed iteration and early termination iteration operation modes. In the early termination iteration mode, when the input data frame reaches the maximum number of iterations or meets the early termination iteration criterion, the iteration will terminate at the end of the current iteration cycle; if the iteration does not terminate, the variable node transmits the external information V2C and the pseudo-posterior probability information of the variable node to the verification node as hard decision information. c It will pass through a multiplexer, and then be written to the external information storage module of the verification node under the control of the routing unit; if the iteration terminates, c The valid information is written to the output storage submodule under the control of the control module.

[0060] Preferably, for the fixed iteration mode, the storage and computing modules associated with the early termination iteration mode of the decoder will stop working to save power; preferably, if the decoder only needs to support the fixed iteration mode, the storage and computing modules associated with the early termination iteration mode will be removed to further save resources.

[0061] Specifically, the fixed iteration method has the advantages of simple implementation and predictable performance, while early termination of iteration can improve decoding efficiency, reduce decoding latency, and effectively save power consumption while avoiding additional computation. In practical applications, the appropriate strategy can be flexibly selected according to the specific needs of the system.

[0062] For FPGA (or ASIC implementation), the decoder adopts an efficient storage strategy based on the minimum block RAM of digital circuits. Specifically, for Xilinx 7 series FPGAs, the maximum bit width of the 18K RAM is 36. This decoder can match 18K RAM with a maximum bit width of 36 or an integer multiple of 36, thereby effectively improving the utilization efficiency of storage resources.

[0063] Thirdly, another objective of the present invention is to provide a channel decoding processing device for implementing the low-complexity high-speed LDPC decoder for CCSDS near-ground application standards.

[0064] Example: Reference Figure 1 This invention provides a low-complexity, high-speed LDPC decoder for the CCSDS near-ground application standard. It is implemented on an FPGA for the shortened (8160, 7136) LDPC code used in actual near-ground communication systems. The decoding algorithm uses F-IAMSA and a (6, 2) quantization scheme, i.e., 1 sign bit, 3 integer bits, and 2 fractional bits. The CCSDS near-ground application standard uses a (8176, 7154) regular QC-LDPC code with a code rate of 7 / 8, and its parity-check matrix consists of 2... It consists of 16 submatrices of order 511, with a row and column weight of 2 for each submatrix. The entire parity check matrix has a row weight of 32 and a column weight of 4.

[0065] Reference Figure 1 In this embodiment of the invention, the first-level channel initial information storage submodule supplements the 18 all-zero bits added before the data frame information bit encoding and deletes the last two bits to adapt to the shortened (8160, 7136) LDPC code used in actual near-ground communication systems. Simultaneously, this module initializes the information corresponding to these 18 all-zero bits as two's complement data with a sign bit of 0 and all other bits of 1, thereby ensuring that the check node information update module can ignore these bits and improve error correction performance.

[0066] Reference Figure 1 For FPGA implementation, this invention employs a high-efficiency storage strategy based on the minimum block RAM of digital circuits. By optimizing the bit width design of the block RAM to match the maximum bit width of 36 for the minimum size block RAM of the FPGA as closely as possible, the utilization efficiency of storage resources is effectively improved. The quantization bit count and the extended parallelism of CPM are both designed to be 6. Therefore, the decoder's row parallelism is 12, and its column parallelism is 96. The channel initial information storage module includes 32 pseudo-dual-port RAMs, each with an effective capacity of 256. 36. The external information storage modules for the verification node and variable node each include 64 pseudo-dual-port RAMs, with each RAM having an effective capacity of 256. 42 and 256 36. The output module includes two pseudo-dual-port RAMs, each with an effective capacity of 256. 84. Since 511 is not divisible by 6, the remaining 5 bits of information at the 86th address of a memory with a depth of 256 will repeat the first 5 bits of information at the first address of the partition to simplify access to the information; the next partition of the memory will use the same strategy.

[0067] Reference Figure 1 and Figure 8 The verification node information update module of this embodiment includes 12 processing units, and the pipeline structure of each processing unit is divided into upper and lower parts. The upper part is responsible for calculating and processing the sign bits of 32 input data V2C, while the lower part is used to calculate and process the absolute values ​​of 32 input data V2C. The processing units are divided into 5-stage pipelines according to their operation functions. Among them, the first stage pipeline is responsible for extracting the sign bits and absolute values ​​of V2C; the upper part of the second stage pipeline obtains the sum of all sign bit data sign32 through XOR operation and delays it by 32 sign bits, while the lower part is used to calculate the minimum value min among the 32 absolute values. 1st Approximate second smallest value amin 2ndThe minimum value is also calculated using the index idx. In the third-stage pipeline, the upper half performs an XOR operation on sign32 and the 32 sign bits to obtain the sum of all 32 sign bits except itself, sign31. The lower half calculates the minimum value and the approximate second smallest value after compensation, and delays the minimum value index. In the fourth-stage pipeline, the sign bit operations in the upper half are complete; only a delay is needed until the absolute value operations are finished before selection and merging. The lower half truncates the previous stage's result and continues to delay the minimum value index, rounding off the truncation. The fifth-stage pipeline selects and merges the processed sign bits and absolute value data. If the sign bit index equals the minimum value index, the absolute value output is the result after second smallest value compensation; otherwise, the absolute value output is the result after minimum value compensation. During merging, the output of the sign bit operation is used as the high-order bit of the original code output, and the output of the absolute value operation is used as the low-order bit of the original code output. Normalization factor. α 1 and α 2 is a floating-point number between 0 and 1, and the offset factor is... β 1 and β 2 is a floating-point number greater than or equal to 0. Appropriate values ​​for these four elements can improve decoding performance. MATLAB simulations show that for the (8160, 7136) LDPC code, when... α 1 and α 2 take 0.75, β 1 take 0, β When the bit error rate is set to 0.25, the F-IAMSA algorithm achieves performance similar to the floating-point F-NMSA algorithm. This choice also makes it suitable for hardware implementation. The bit error rate curve of the F-IAMSA algorithm is shown in the reference... Figure 14 Quadrature Phase Shift Keying (QPSK) modulation and Additive White Gaussian Noise (AWGN) channel are used. Under the (6,2) quantization scheme, the decoding threshold of the F-IAMSA algorithm after 10 iterations is approximately 4.14 dB (bit error rate BER of 10). -7 Compared to the floating-point F-NMSA algorithm, the performance loss is approximately 0.05 dB; compared to the fixed-point F-NMSA algorithm, the performance loss is approximately 0.02 dB; and compared to the fixed-point F-NMSA algorithm, the performance improvement is approximately 0.02 dB.

[0068] Considering the simplicity of the pipeline structure of the variable node information update unit, to balance the pipeline latency and computational resources of the variable node and check node information update units, the calculation of converting the original code to the two's complement is placed in the variable node information update unit. Furthermore, in the second-stage pipeline of the check node information update unit, the lower half of the computational function is relatively complex, and the critical path latency is large. To further optimize the processing speed of the check node information update unit, this part of the operation is refined, inserting more pipeline stages, and using a tree-structured minimum and near-minimum value generator to find the minimum and near-minimum values. The tree-structured minimum and near-minimum value generator and its two basic components are described below. Figures 10 to 13 Finding the minimum and near-second minimum values ​​of 32 numbers would require inserting a 5-stage pipeline. Furthermore, the tree-structured minimum and near-second minimum value generator reduces the hardware resource overhead of the decoder by reducing the number of comparators and selectors, trading a small loss in error correction performance for a significant reduction in decoder implementation complexity.

[0069] Reference Figure 10 and Figure 13 In this embodiment of the invention, the tree-structured minimum and near-minimum value generator operates in two stages. In the first stage, the input data is divided into four groups. Each group is pipelined according to the minimum value generator of the tree structure, and the minimum value of each group is obtained through pairwise comparisons. In the second stage, the set of minimum values ​​for each group is pipelined according to the minimum and near-minimum value generator of the tree structure, and the minimum and near-minimum values ​​of all input data are obtained through pairwise comparisons. The index of the minimum value is generated by the minimum value index generator. For 32 input data points, the tree-structured minimum and near-minimum value generator requires 33 two-input comparators and 36 two-input selectors; the two-input selectors required by the minimum value index generator are not included in this calculation. Where m... 10 and m 11 These represent the minimum values ​​of the first and second 2-MVG2 values ​​in this pipeline stage, respectively, m. 20 and m 21 These represent the second smallest values ​​of 2-MVG2 for the first and second stages of this pipeline, respectively.

[0070] Reference Figure 1 and Figure 9 The variable node information update module in this embodiment of the invention includes 96 processing units, each of which adopts a 4-stage pipeline structure. l i This represents the initial channel information corresponding to the current variable node, expressed as the external information (C2V) passed to the variable node by the four check nodes. r ji , , Indicates the relationship with the first iThe C2V data is a set of other check nodes connected to a variable node, and its form is in original code. Addition and subtraction operations are more convenient using two's complement; therefore, the first-stage pipeline first converts the external information C2V to two's complement and delays... l i The second-stage pipeline uses an adder array to calculate the sum of all input data (sum). The third-stage pipeline uses a subtractor array to calculate the sum and then subtracts the differences of the four C2V values. To prevent precision loss, the bit width of the intermediate calculation results must be wider than the bit width of the input and output data. Therefore, the fourth-stage pipeline needs to truncate the intermediate calculation results to obtain the hard decision information. c And 4 V2Cs.

[0071] Reference Figure 1 and Figure 7 In this embodiment of the invention, the verification equation calculation module and the verification node information update module are executed in parallel, including 12 processing units. Each processing unit corresponds to a verification equation and calculates a syndrome component. When all syndrome components are zero, i.e., the syndrome vector is a zero vector, it indicates that the decoding result is correct. The calculation of the syndrome component is to calculate the XOR sum of 32 hard decision information. Meanwhile, in LDPC decoding, the length of the syndrome vector is 1022, so a counting operation is used to determine whether the syndrome vector is a zero vector to save storage resources.

[0072] Reference Figure 1 The output module of this embodiment includes an output storage submodule, two selection and merging units, a deletion unit, and a first-in-first-out data buffer. The output module is responsible for converting the stored out-of-order bit data into 8-channel decoded output arranged in the correct order. The bit data corresponding to each sub-matrix block is sequential.

[0073] Reference Figure 1 This invention is implemented based on the Xilinx XC7VX485TFFG1761-2 FPGA, with a clock frequency constraint of 320MHz. The decoder resource usage is shown in Table 1. Table 1 Decoder Hardware Resource Usage Table LUT 26991 303600 8.89 LUTRAM 387 130800 0.30 FF 38882 607200 6.40 BRAM 117.50 1030 11.41 Slice 8623 75900 11.36 For an LDPC decoder based on the CCSDS near-ground application standard, this invention adds an input buffer module and an output module to shape the input and output data, thereby further improving the decoder's flexibility and adaptability.

[0074] Under the same clock frequency and quantization scheme, this invention employs F-IAMSA and extended parallelism of CPM. pOption 6 is chosen. The verification node information update unit adopts an 8-stage pipeline, the variable node information update unit adopts a 3-stage pipeline, and the external information storage module of the verification node uses distributed RAM to store hard decision information. The hardware implementation results of the decoder implemented are compared with those proposed in Dr. Kang Jing's thesis "Research on Encoding and Decoding Algorithm and Efficient Implementation Technology of LDPC Code for High-Speed ​​Data Transmission between Space and Ground" (Kang Jing. Research on Encoding and Decoding Algorithm and Efficient Implementation Technology of LDPC Code for High-Speed ​​Data Transmission between Space and Ground [D]. University of Chinese Academy of Sciences (National Space Science Center of Chinese Academy of Sciences), 2021.) as shown in Table 2. Table 2 Comparison of hardware implementation results of the decoder of this invention and the decoder proposed by Kang Jing

[0075] Compared to the high-speed LDPC decoder in Kang Jing's thesis, the decoder of this invention achieves an approximately 88.2% increase in throughput while saving approximately 38.1% of LUT resources, 13.5% of FF resources, and 2.8% of BRAM resources on the FPGA hardware, with only a 0.02dB loss in error correction performance. Furthermore, the decoder supports two operating modes and adds an input buffer module and an output module, further enhancing its flexibility and adaptability. Compared to the decoder proposed in Dr. Kang Jing's thesis, this invention not only improves the decoding algorithm but also employs superior decoding strategies: overlapping double-frame processing and an efficient storage strategy. Therefore, this invention enables the realization of a lower complexity and higher throughput CCSDS near-ground application standard LDPC decoder while increasing both the decoder's functionality and flexibility.

[0076] Under the same clock frequency, quantization scheme, and single iteration clock period, this invention employs F-IAMSA and extended parallelism of CPM. p The decoder is consistent with that proposed in Dr. Xie Tianjiao's dissertation "Research on Key Technologies of Adaptive Transmission and Networking in Integrated Space-Ground Networks" (Xie Tianjiao. Research on Key Technologies of Adaptive Transmission and Networking in Integrated Space-Ground Networks [D]. Northwestern Polytechnical University, 2020.), namely... p =12, the verification node information update unit adopts an 8-stage pipeline, and the variable node information update unit adopts a 3-stage pipeline. The comparison between the implemented 275MHz decoder and the decoder hardware implementation results proposed in Dr. Xie Tianjiao's dissertation is shown in Table 3: Table 3 Comparison of hardware implementation results of the decoder of this invention and the decoder proposed by Xie Tianjiao

[0077] In implementation, this invention removes the input buffer module and output module, retaining only the output storage submodule within the output module to facilitate fair comparison. Compared to the high-speed LDPC decoder in Xie Tianjiao's paper, this invention's 275MHz decoder, with a 0.02dB loss in error correction performance, saves approximately 9.5% of LUT resources and 25.7% of FF resources in FPGA hardware resources. Furthermore, increasing the number of pipeline stages in the core module can improve the maximum clock frequency supported by the decoder. When the verification node information update unit uses a 9-stage pipeline, the variable node information update unit uses a 4-stage pipeline, and an output register is added to the memory, the maximum clock frequency supported by this invention's decoder reaches 320MHz. Table 3 shows a comparison between the implemented 320MHz decoder and the decoder hardware implementation results proposed in Dr. Xie Tianjiao's paper. At this point, the theoretical throughput of this invention's decoder after 10 iterations can reach 4.477Gbps. Compared to the high-speed LDPC decoder in Xie Tianjiao's thesis, the decoder of this invention achieves a throughput increase of approximately 11.8% while saving approximately 6.6% of LUT resources and 13.7% of FF resources in FPGA hardware resources, with only a 0.02dB loss in error correction performance. Compared to the decoder proposed in Dr. Xie Tianjiao's thesis, this invention improves the decoding algorithm, employing a tree-structured minimum and near-minimum value generator and a multi-stage pipeline design. This achieves a higher-performance decoder design with minimal loss in error correction performance, better meeting the needs of high-speed satellite data transmission systems.

Claims

1. A low-complexity, high-speed LDPC decoder for near-ground applications of the CCSDS standard, characterized in that, Includes the following modules: Input buffer module: includes a first-level input buffer unit and a second-level input buffer unit, which respectively receive odd-numbered data frames and even-numbered data frames of the high-speed input data frame and perform channel number conversion. The input buffer module receives the channel initial information of the high-speed input data frame and outputs it to the channel initial information storage module after channel number conversion. Channel initial information storage module: outputs the received channel initial information after path number conversion to the variable node information update module; includes a first-level channel initial information storage submodule and a second-level channel initial information storage submodule, the first-level channel initial information storage submodule writes its internal data into the second-level channel initial information storage submodule; Control module: Coordinates the operation of different modules in the LDPC decoder; Verification Node External Information Storage Module: Stores hard decision information including external information (V2C) passed from variable nodes to verification nodes and pseudo-posterior probability information of variable nodes; Verification node information update module: Calculates the external information passed from the verification node to the variable node; The verification equation calculation module calculates the verification equation corresponding to the LDPC code verification matrix and checks whether the hard decision information of the pseudo-posterior probability information of the variable node satisfies the verification equation. External information storage module for variable nodes: stores the external information (C2V) passed from the verification node to the variable node; Variable node information update module: calculates the hard decision information of the external information passed from the variable node to the verification node and the pseudo-posterior probability information of the variable node; Output module; The LDPC decoder comprises an improved partially parallel architecture consisting of a channel initialization information storage module, a check node external information storage module, a check node information update module, a variable node information update module, a variable node external information storage module, and an output storage submodule within the output module. It divides the QC-LDPC code's check matrix according to the number of cyclic permutation matrices (CPMs) within it, allocating independent memory to each CPM. Simultaneously, it stores the cyclic permutation matrices of each CPM... p row or p The extra-list information is stored at the same address in the corresponding memory; p To extend the parallelism of the cyclic permutation matrix CPM, the parity-check matrix of the QC-LDPC code is formed by... m n indivual L Composed of submatrices of order, with row and column weights of size . w ; When the maximum throughput of the LDPC decoder is less than or equal to the rate of the high-speed input data frame, the improved partially parallel architecture adopts an overlapping double-frame processing strategy, with a line parallelism of [value missing]. mp The column parallelism is np The first-level channel initial information storage submodule and the second-level channel initial information storage submodule respectively include n Each random access memory (RAM) includes a verification node external information storage module and a variable node external information storage module, respectively. m n w Each random access memory (RAM) has an effective capacity of 2. L / p Qp, Q The number of quantization bits; When the LDPC decoder's maximum throughput exceeds the rate of the high-speed input data frame, the improved partially parallel architecture adopts a non-overlapping double-frame processing strategy. This strategy merges the external information storage module for the check node and the external information storage module for the variable node into a single external information storage module. This external information storage module only includes... m n w There are random access memory (RAM) units, each with an effective capacity of . L / p Qp .

2. The LDPC decoder according to claim 1, characterized in that, The control module includes a channel initial information storage module read control unit, a routing unit, a check node external information storage module read control unit, a variable node external information storage module read control unit, a variable node information update module external information output enable control unit, and an output storage submodule write control unit. The channel initial information storage module read control unit is used to control the reading of the channel initial information stored in the channel initial information storage module; The routing unit is used to control the data write addresses of the external information storage module for the verification node and the external information storage module for the variable node; The read control unit of the external information storage module for the verification node is used to control the reading of the external information of the verification node stored in the external information storage module for the verification node; The variable node external information storage module read control unit is used to control the reading of variable node external information stored in the variable node external information storage module; The variable node information update module external information output enable control unit is used to control the output enable of external information of the variable node information update module. The output storage submodule write control unit is used to control the data writing of the output storage submodule.

3. The LDPC decoder according to claim 1, characterized in that, The external information storage module for the verification node supports additional storage of hard decision information, including pseudo-posterior probability information of variable nodes. When the LDPC decoder is compatible with the early termination iteration operation mode, the external information storage module for the verification node selects a compact storage strategy or a distributed random access memory (RAM) storage strategy. The compact storage strategy stores the hard decision information and the external information passed from the variable nodes to the verification node in the same memory. The external information storage module for the verification node consists of... m n w It consists of 2 random access memories (RAMs), each with an effective capacity of 2. L / p Q ( p +1).

4. The LDPC decoder according to claim 1, characterized in that, The decoding algorithm of the LDPC decoder adopts the IAMSA (Infinite Minimum Sum Algorithm) with multi-factor correction based on flooding scheduling. The IAMSA check node information update formula is as follows: in, r ji Indicates the first j The verification node is passed to the first... i External information of each variable node Q j \ i Indicates except the first i Outside of the variable node, with the first variable node j The set of other variable nodes connected to a check node. sign Represents a symbolic function. q ij Indicates the first i The variable node is passed to the first j The external information of each verification node, where max represents the function for finding the maximum value. α 1 and α 2 represents the normalization factor, and amin represents the function for finding the minimum and approximate second minimum values. β 1 and β 2 represents the offset factor.

5. The LDPC decoder according to claim 1, characterized in that, Both the verification node information update module and the variable node information update module are implemented using a multi-stage pipeline. The multi-stage pipeline uses a tree-structured minimum and near-minimum value generator to calculate the minimum and near-minimum values. The tree-structured minimum and near-minimum value generator calculates... w r The minimum and near-second minimum values ​​of the number need to be inserted. log2 w r Multi-stage production line w r This indicates the row weight of the parity check matrix.

6. The LDPC decoder according to claim 5, characterized in that, The minimum and near-second minimum value generators of the tree structure are divided into a first stage and a second stage during computation; in the first stage, the input data is divided into... N The data is divided into groups, and each group of data is divided into pipeline stages according to the minimum value generator of the tree structure. The minimum value of each group of data is obtained by comparing each pair of data. In the second stage, the minimum value set of each data set is divided into pipeline stages according to the minimum and second minimum value generators of the tree structure, and the minimum and approximate second minimum values ​​of all input data are obtained by pairwise comparisons.

7. A control method for a low-complexity, high-speed LDPC decoder for CCSDS near-ground application standards, based on the LDPC decoder of claim 1, characterized in that, Includes the following steps: S1: The first-level channel initial information storage submodule of the high-speed input data frame is input to the channel initial information storage module; S2: At the start of the first iteration of each data frame in the high-speed input data frame, the data stored in the first-level channel initial information storage submodule in step S1 is read and written to the second-level channel initial information storage submodule and then written to the external information storage module of the verification node after passing through the multiplexer. S3: When the output data of the external information storage module of the verification node described in step S2 is valid, the verification node information update module and the verification equation calculation module start the verification node information update work. The verification node information update module outputs the calculated external information to the external information storage module of the variable node, and the verification equation calculation module outputs the judgment flag of whether the codeword satisfies the verification equation to the control module. S4: When the output data of the variable node external information storage module in step S3 and the second-level channel initial information storage submodule in step S2 are valid, the variable node information update module starts the variable node information update work. The second-level channel initial information storage submodule outputs the original data frame corresponding to the data frame that is undergoing decoding iteration to the variable node information update module. The variable node information update module outputs the calculated external information to the verification node external information storage module. S5: The verification node information update in step S3 and the variable node information update in step S4 together constitute the iteration of the input data frame. When the input data frame reaches the maximum number of iterations or meets the early termination iteration criterion, the data frame terminates the iteration at the end of the current iteration cycle. Hard decision information. c The valid information is written to the output storage submodule under the control of the control module; While the data frame iteration has not terminated, the variable node passes external information (V2C) and hard decision information to the verification node. c First, it passes through a multiplexer, and then is written to the external information storage module of the verification node in step S3 under the control of the routing unit. S6: Read the hard decision information from the output storage submodule described in step S5 in the order of data frame columns. c It converts the stored out-of-order bit data into decoded output in the correct order and removes invalid data.

8. A channel decoding processing apparatus, based on the LDPC decoder according to any one of claims 1 to 6, characterized in that, The channel decoding processing device is used to implement the LDPC decoder.

Citation Information

Patent Citations

  • LDPC (Low Density Parity Check) decoder for improving decoding efficiency and throughput of decoder

    CN115580309A

  • High-energy-efficiency LDPC decoder for high-speed satellite link

    CN115664584A

  • High-speed LDPC (Low Density Parity Check) decoder suitable for near-earth satellite communication and interleaving decoding method

    CN118316462A