Low-complexity high-speed LDPC decoder oriented to CCSDS near-earth application standard, control method and equipment
By adding an input buffer module and a verification equation calculation module to the LDPC decoder, an approximate minimum sum algorithm with improved partial parallel architecture and flood scheduling is adopted to solve the problems of high complexity and low resource utilization efficiency in the prior art, and high throughput and low power consumption under low complexity are achieved.
Patent Information
- Application Number
- CN202510235590.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-28
AI Technical Summary
Existing LDPC decoders have high complexity and low resource utilization efficiency when implementing high-speed decoding, making it difficult to achieve high throughput at low complexity.
By adding input buffer module, verification equation calculation module and output module, the improved partial parallel architecture and multi-factor correction approximate minimum sum algorithm for flood scheduling is used to optimize the time utilization of the check node and variable node information update module to realize overlapping double-frame processing.
It achieves higher throughput at low complexity, has low power consumption and flexibility, and significantly improves resource utilization efficiency through optimization of storage strategies and pipeline design.
Smart Images

Figure CN120165704A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of satellite communications and relates to a low-complexity high-speed LDPC decoder oriented to the CCSDS near-earth application standard and a control method and equipment. Background Art
[0002] As a high-performance error correction coding technology, Low-Density Parity-Check (LDPC) code has the characteristics of high coding gain, low error floor, and close to the Shannon limit. It is widely used in wireless communication, digital broadcasting, data storage, deep space communication and other fields. The LDPC code of the Consultative Committee for Space Data Systems (CCSDS) near-earth application standard has a high code rate, low required check redundancy overhead, and a coding gain of about 7dB. Therefore, it is very popular in the field of satellite communication. LDPC decoding usually adopts a soft decision iterative decoding algorithm. In actual engineering applications, the normalized minimum sum algorithm (NMSA) and the offset minimum sum algorithm (Offset Each iteration of the LDPC decoder requires a large amount of data calculation and storage, which has high implementation complexity and requires multiple iterations to obtain excellent error correction performance, which means that a large amount of hardware resources are needed to realize a high-speed and reliable LDPC decoder. Therefore, improving the resource utilization efficiency of the LDPC code decoder and realizing a high-speed LDPC decoder with low complexity are important issues to be solved in the field of satellite data transmission.
[0003] In the prior art, the invention with the patent publication number of "CN115580309A" and the name of "An LDPC decoder for improving decoding efficiency and throughput of decoder" proposes an LDPC decoder under the CCSDS standard that can improve throughput and decoding efficiency. It addresses the problem that "a circulant matrix with a row weight of 2 needs to be decoded twice in a cycle to complete the update, and the decoding efficiency is low". Through two parallel operations, the pseudo-posteriori probability information of the circulant matrix with a row weight of 2 is updated in only one cycle, thereby improving the decoding efficiency and throughput. The patent adopts L-NMSA. Although the invention improves the decoding efficiency and throughput, it does not utilize the information that has been updated in each layer of codewords, sacrificing the error correction performance.
[0004] In the prior art, the patent publication number is "CN115664584A", and the name is "A High-Energy-Efficient LDPC Decoder for High-Speed Satellite Links", which makes improvements to the core processing modules (variable node and check node information update units) of the LDPC decoder. For the variable node update unit, this patent improves the utilization rate of FPGA Slice resources by reducing the number of pipeline stages and reduces the decoding delay at the same time. For the check node update unit, this patent discards the second smallest value in the first-level comparison, which greatly reduces the implementation complexity of the circuit without obvious loss of performance. However, in order to reduce the subtraction operation and expand the bit width, the operation of removing the subtraction operation increases the number of circuit adders / subtractors, and at the same time, the space for optimizing the operation of this invention is limited. In the final implementation, the implementation complexity and resource occupancy of the decoder are still relatively high. Summary of the Invention
[0005] Aiming at the deficiencies in the prior art, the purpose of the present invention is to propose a low-complexity high-speed LDPC decoder, control method, and device for the CCSDS near-Earth application standard. The LDPC decoder adds an input buffer module, a check equation calculation module, and an output module, adopts an improved partially parallel architecture, and the decoding algorithm adopts a multi-factor corrected approximate min-sum algorithm, enabling the LDPC decoder to achieve higher throughput at low complexity and having effects such as low power consumption and flexibility.
[0006] To achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:
[0007] In a first aspect, a low-complexity high-speed LDPC decoder for the CCSDS near-Earth application standard, the LDPC decoder includes an input buffer module, a channel initial information storage module, a control module, a check node extrinsic information storage module, a check node information update module, a check equation calculation module, a variable node extrinsic information storage module, a variable node information update module, and an output module:
[0008] The input buffer module is used to receive the channel initial information of the high-speed input data frame, and after path conversion, output it to the channel initial information storage module. Among them, the high-speed input data frame includes odd data frames and even data frames. The input buffer module includes a first-level input buffer unit and a second-level input buffer unit, which are respectively used to receive the input odd data frames and even data frames and perform path conversion;
[0009] The channel initial information storage module receives the channel initial information after the conversion of the number of paths of the high-speed input data frame and stores it, and finally outputs it to the variable node information update module; the channel initial information storage module includes a first-level channel initial information storage sub-module and a second-level channel initial information storage sub-module, and the first-level channel initial information storage sub-module writes its internal data into the second-level channel initial information storage sub-module;
[0010] The control module is used to coordinate the operations of different modules of the LDPC decoder;
[0011] The check node extrinsic information storage module is used to store the extrinsic information (Variable Node to Check Node, V2C) passed from the variable node to the check node and the hard decision information of the pseudo a posteriori probability information of the variable node;
[0012] The check node information update module is used to calculate the extrinsic information passed from the check node to the variable node;
[0013] The check equation calculation module is used to calculate the check equations corresponding to the LDPC code check matrix and check whether the hard decision information of the pseudo a posteriori probability information of the variable node satisfies the check equations;
[0014] The variable node extrinsic information storage module is used to store the extrinsic information (Check Node to Variable Node, C2V) passed from the check node to the variable node;
[0015] The variable node information update module is used to calculate the extrinsic information passed from the variable node to the check node and the hard decision information of the pseudo a posteriori probability information of the variable node.
[0016] Further, the control module includes a channel initial information storage module read control unit, a routing unit, a check node extrinsic information storage module read control unit, a variable node extrinsic information storage module read control unit, a variable node information update module extrinsic information output enable control unit, and an output storage sub-module write control unit:
[0017] The channel initial information storage module read control unit is used to control the reading of the channel initial information stored in the channel initial information storage module;
[0018] The routing unit is used to control the data writing addresses of the check node extrinsic information storage module and the variable node extrinsic information storage module;
[0019] The check node extrinsic information storage module read control unit is used to control the reading of the check node extrinsic information stored in the check node extrinsic information storage module;
[0020] The read control unit of the extrinsic information storage module of the variable node is used to control the reading of the extrinsic information of the variable node stored in the extrinsic information storage module of the variable node;
[0021] The output enable control unit of the extrinsic information of the variable node information update module is used to control the output enable of the extrinsic information of the variable node information update module;
[0022] The write control unit of the output storage sub-module is used to control the data writing of the output storage sub-module.
[0023] Furthermore, the LDPC decoder consists of a channel initial information storage module, an extrinsic information storage module of the check node, a check node information update module, a variable node information update module, an extrinsic information storage module of the variable node, and an output storage sub-module to form an improved partially parallel architecture. The improved partially parallel architecture divides the parity-check matrix of the QC-LDPC code according to the number of cyclic permutation matrices (CPMs) therein, and assigns an independent memory to each CPM. At the same time, the improved partially parallel architecture stores the p rows (columns) of extrinsic information of the CPM at the same address of the corresponding memory.
[0024] Further, when the maximum throughput of the LDPC decoder is less than or equal to the rate of the high-speed input data frame, the improved partially parallel architecture adopts an overlapping double-frame processing strategy. Among them, the row parallelism of the improved partially parallel architecture is mp, the column parallelism is np, the first-level channel initial information storage sub-module and the second-level channel initial information storage sub-module each include n random access memories (RAMs), and the extrinsic information storage module of the check node and the extrinsic information storage module of the variable node each include m×n×w RAMs, and the effective capacity of each RAM is Q is the quantization bit number; when the maximum throughput of the LDPC decoder is greater than the rate of the high-speed input data frame, the improved partially parallel architecture adopts a non-overlapping double-frame processing strategy. Among them, the improved partially parallel architecture combines the extrinsic information storage module of the check node and the extrinsic information storage module of the variable node into an extrinsic information storage module, and the extrinsic information storage module only includes m×n×w RAMs, and the effective capacity of each RAM is
[0025] Further, the extrinsic information storage module of the check node supports additional storage of the hard decision information of the pseudo posterior probability information of the variable node. When the LDPC decoder is compatible with the operation mode of early termination of iteration, the extrinsic information storage module of the check node selects a compact storage strategy or a distributed RAM storage strategy. The compact storage strategy is to store the hard decision information and the extrinsic information passed by the variable node to the check node in the same memory. Among them, the extrinsic information storage module of the check node consists of m×n×w RAMs, and the effective capacity of each RAM is
[0026] Further, the decoding algorithm of the LDPC decoder adopts the Improved Normalized Approximate Min-Sum Algorithm (IAMSA) with flood scheduling and multi-factor correction. The check node information update formula of the IAMSA algorithm is as follows:
[0027]
[0028] Among them, r ji represents the extrinsic information passed from the j-th check node to the i-th variable node, Q j \i represents the set of other variable nodes connected to the j-th check node except the i-th variable node, sign represents the sign function, q ij represents the extrinsic information passed from the i-th variable node to the j-th check node, max represents the maximum value function, α1 and α2 represent the normalization factors, amin represents the maximum value and approximate sub-minimum value function, and β1 and β2 represent the offset factors.
[0029] Further, both the check node information update module and the variable node information update module are implemented by a multi-stage pipeline. The multi-stage pipeline uses a tree structure minimum and approximate sub-minimum value generator (amin) to calculate the minimum value and approximate sub-minimum value (Tree Structure Minimum and Approximate Sub Minimum Value Generator, M1AM2VG TS ). In the tree structure minimum and approximate sub-minimum value generator, inserting r numbers is required to calculate the minimum value and approximate sub-minimum value of stages of the pipeline.
[0030] Further, the minimum value and approximate second - minimum value generator of the tree structure is calculated in two stages; in the first stage, the input data is divided into N groups, and each group of data is divided into pipeline stages according to the minimum value generator of the tree structure, and the minimum value of each group of data is obtained through pairwise comparison; in the second stage, the set of minimum values of each group of data is divided into pipeline stages according to the minimum value and second - minimum value generator of the tree structure, and the minimum value and approximate second - minimum value of all input data are obtained through pairwise comparison.
[0031] In a second aspect, a control method for a low - complexity high - speed LDPC decoder for CCSDS near - earth application standards includes the following steps:
[0032] S1: The high - speed input data frame is input to the first - stage channel initial information storage sub - module of the channel initial information storage module.
[0033] S2: At the beginning of the first iteration of each data frame in the high - speed input data frame, the data stored in the first - stage channel initial information storage sub - module in step S1 is read and written into the second - stage channel initial information storage sub - module and written into the check node extrinsic information storage module through a multiplexer.
[0034] S3: When the output data of the check node extrinsic information storage module in step S2 is valid, the check node information update module and the check equation calculation module start the check node information update work. Among them, the check node information update module outputs the calculated extrinsic information C2V to the variable node extrinsic information storage module, and the check equation calculation module outputs the judgment flag of whether the codeword satisfies the check equation to the control module.
[0035] S4: When the output data of the variable node extrinsic information storage module in step S3 and the second - stage channel initial information storage sub - module in step S2 are valid, the variable node information update module starts the variable node information update work. The second - stage channel initial information storage sub - module outputs the original data frame corresponding to the data frame undergoing decoding iteration to the variable node information update module, and the variable node information update module outputs the calculated extrinsic information V2C to the check node extrinsic information storage module.
[0036] S5: The check node information update work in step S3 and the variable node information update work in step S4 together constitute the iteration of the input data frame. When the input data frame reaches the maximum number of iterations or meets the early termination iteration criterion, the data frame terminates the iteration at the end of this iteration cycle, and the valid information in the hard - decision information c is written into the output storage sub - module under the control of the control module; when the iteration of the data frame does not terminate, V2C and c first pass through a multiplexer and then are controlled by a routing unit to be written into the check node extrinsic information storage module in step S3.
[0037] S6: Read the hard decision information c from the output storage sub-module described in step S5 in the order of data frame columns, convert the stored out-of-order bit data into decoded output arranged in the correct order, and eliminate invalid data.
[0038] In a third aspect, a channel decoding processing device is provided, which is used to implement the low-complexity high-speed LDPC decoder for the CCSDS near-Earth application standard.
[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0040] 1. The LDPC decoder provided by the present invention realizes a decoder with higher throughput at low complexity by adding an input buffer module, a check equation calculation module and an output module, and improving the decoding processing method. The LDPC decoder also has low power consumption and flexibility.
[0041] 2. The LDPC decoder provided by the present invention adopts overlapping dual-frame processing, and realizes parallel processing of two frames of data by optimizing the time utilization rate of the check node and variable node information update module; with only an increase in information storage resources and a small amount of control logic resources, and without increasing computing resources, it achieves a significant improvement in the throughput of the decoder at a lower resource cost.
[0042] 3. The LDPC decoder provided by the present invention adopts an improved partially parallel architecture. The check node and variable node information update units process the rows and column groups of the check matrix in a time-division multiplexing manner. The improved partially parallel architecture further improves the parallel processing ability of the partially parallel decoder by mapping the extrinsic information of multiple rows (or columns) of the CPM to the same storage address, and effectively avoids storage access conflicts.
[0043] 4. The LDPC decoder provided by the present invention adopts an efficient storage strategy, specifically including: a compact storage strategy, which stores the hard decision information and the extrinsic information V2C passed from the variable node to the check node in the same memory to save storage and counter resources; a storage strategy based on the minimum block RAM of the digital circuit, which optimizes the bit width design of the block RAM to make it as close as possible to the maximum bit width or its integer multiple of the minimum size block RAM, thereby effectively improving the utilization efficiency of storage resources. (Distributed RAM storage strategy, juxtaposed with the compact storage strategy) For FPGAs, when the chip block RAM resources are scarce or the internal utilization rate is low, the check node extrinsic information storage module will use distributed RAM to store the hard decision information.
[0044] 5. The LDPC decoder provided by the present invention adopts a multi-factor corrected approximate min-sum algorithm with flooding scheduling. When updating the check node information, the minimum value function only calculates the exact minimum value and the approximate second minimum value to reduce the implementation complexity of the decoder. At the same time, different normalization factors and offset factors are added to compensate for the minimum value and the approximate second minimum value to reduce the loss of the decoder's error correction performance. When implementing the decoder, the present invention adopts a tree-structured minimum value and approximate second minimum value generator to reduce the hardware resource overhead of the decoder by reducing the number of comparators and selectors, and exchanges a small loss of error correction performance for a significant reduction in the implementation complexity of the decoder.
[0045] In summary, through improvements such as an improved partial parallel architecture, overlapping dual-frame processing, a multi-factor corrected approximate min-sum algorithm, and a tree-structured minimum value and approximate second minimum value generator, the LDPC decoder provided by the present invention can achieve higher throughput while reducing complexity, with a small loss of error correction performance and cost savings. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is an implementation block diagram of a low-complexity high-speed LDPC decoder for the CCSDS near-Earth application standard provided by an embodiment of the present invention;
[0047] Figure 2 is a schematic diagram of a control method of a low-complexity high-speed LDPC decoder for the CCSDS near-Earth application standard provided by an embodiment of the present invention;
[0048] Figure 3 is a schematic structural diagram of a traditional partial parallel architecture LDPC decoder provided by an embodiment of the present invention;
[0049] Figure 4 is a schematic structural diagram of an improved partial parallel architecture LDPC decoder provided by an embodiment of the present invention;
[0050] Figure 5 is a schematic structural diagram of an LDPC decoder with overlapping dual-frame processing provided by an embodiment of the present invention;
[0051] Figure 6 is a schematic diagram of the principle of overlapping dual-frame processing of an LDPC decoder provided by an embodiment of the present invention;
[0052] Figure 7 is a schematic structural diagram of an LDPC decoder supporting early termination of iteration provided by an embodiment of the present invention;
[0053] Figure 8 is a schematic diagram of the pipeline structure of a check node information update unit provided by an embodiment of the present invention;
[0054] Figure 9It is a schematic diagram of the pipeline structure of the variable node information update unit provided by an embodiment of the present invention;
[0055] Figure 10 It is a schematic diagram of the structure of the minimum value and approximate second minimum value generator of the tree structure provided by an embodiment of the present invention;
[0056] Figure 11 It is the minimum value generator 2-MVG1 with two inputs and a single minimum value output in the minimum value and approximate second minimum value generator of the tree structure provided by an embodiment of the present invention;
[0057] Figure 12 It is the minimum value generator 2-MVG2 with two inputs and two minimum value outputs in the minimum value and approximate second minimum value generator of the tree structure provided by an embodiment of the present invention;
[0058] Figure 13 It is a schematic diagram of the structure of the minimum value and approximate second minimum value generator of the tree structure with 32 input data provided by an embodiment of the present invention;
[0059] Figure 14 It is the bit error rate curve of the (8160, 7136) LDPC code provided by an embodiment of the present invention. Detailed implementation manners
[0060] Next, in combination with Figures 1 to 14 and the embodiments, the present invention will be further described in detail.
[0061] In a first aspect, a low-complexity high-speed LDPC decoder for the CCSDS near-Earth application standard, the LDPC decoder includes an input buffer module, a channel initial information storage module, a control module, a check node extrinsic information storage module, a check node information update module, a check equation calculation module, a variable node extrinsic information storage module, a variable node information update module, and an output module:
[0062] The input buffer module is configured to receive the channel initial information of the high-speed input data frame, and after path conversion, output it to the channel initial information storage module. Among them, the high-speed input data frame includes odd data frames and even data frames, which can improve the data throughput of the system. The input buffer module includes a first-level input buffer unit and a second-level input buffer unit, which are respectively configured to receive the input odd data frame and even data frame, and perform path conversion;
[0063] The channel initial information storage module receives the channel initial information after the conversion of the number of paths of the high-speed input data frame and stores it, and finally outputs it to the variable node information update module. The channel initial information storage module includes a first-level channel initial information storage sub-module and a second-level channel initial information storage sub-module, which are respectively used for the initialization of the input data frame and providing the original input data frame required by the variable node information update module for each iteration. The first-level channel initial information storage sub-module writes its internal data into the second-level channel initial information storage sub-module;
[0064] The control module is used to coordinate the operations of different modules of the LDPC decoder to ensure the correct progress of the decoding process;
[0065] The check node extrinsic information storage module is used to store the extrinsic information (Variable Node to Check Node, V2C) passed from the variable node to the check node and the hard decision information of the pseudo a posteriori probability information of the variable node;
[0066] The check node information update module is used to calculate the extrinsic information passed from the check node to the variable node;
[0067] The check equation calculation module is used to calculate the check equations corresponding to the LDPC code check matrix and check whether the hard decision information of the pseudo a posteriori probability information of the variable node satisfies the check equations;
[0068] The variable node extrinsic information storage module is used to store the extrinsic information (Check Node to Variable Node, C2V) passed from the check node to the variable node;
[0069] The variable node information update module is used to calculate the extrinsic information passed from the variable node to the check node and the hard decision information of the pseudo a posteriori probability information of the variable node;
[0070] The output module includes an output storage sub-module, two selection combining units, a deletion unit, and a first-in first-out data buffer (First Input First Output, FIFO). The output module jointly completes the processing and output of the hard decision information of the pseudo a posteriori probability information of the variable node in the order of the output storage sub-module, the two selection combining units, the deletion unit, and the first-in first-out data buffer.
[0071] Specifically, the information stored in the first-level channel initial information storage sub-module will be written into the check node extrinsic information storage module and the second-level channel initial information storage sub-module at the beginning of the first iteration of the corresponding input data frame; meanwhile, for the CCSDS near-Earth communication system, the length of the input data frame of the LDPC decoder is 8160, while the length of the standard LDPC code is 8176. The first-level channel initial information storage sub-module will complement the 18 all-0 bits added before the information bit encoding of the data frame and delete the last 2 bits before the iterative decoding of the data frame; for these 18 all-0 bits, the first-level channel initial information storage sub-module initializes the corresponding information to the complement code data with the sign bit being 0 and the rest of the bits being all 1, so as to ensure that the check node information update module ignores these bit information and improve the error correction performance.
[0072] Furthermore, the control module includes a read control unit for the channel initial information storage module, a routing unit, a read control unit for the check node extrinsic information storage module, a read control unit for the variable node extrinsic information storage module, an enable control unit for the output of the extrinsic information of the variable node information update module, and a write control unit for the output storage sub-module:
[0073] The read control unit for the channel initial information storage module is used to control the reading of the channel initial information stored in the channel initial information storage module;
[0074] The routing unit is used to control the data writing addresses of the check node extrinsic information storage module and the variable node extrinsic information storage module;
[0075] The read control unit for the check node extrinsic information storage module is used to control the reading of the check node extrinsic information stored in the check node extrinsic information storage module;
[0076] The read control unit for the variable node extrinsic information storage module is used to control the reading of the variable node extrinsic information stored in the variable node extrinsic information storage module;
[0077] The enable control unit for the output of the extrinsic information of the variable node information update module is used to control the output enable of the extrinsic information of the variable node information update module;
[0078] The write control unit for the output storage sub-module is used to control the data writing of the output storage sub-module.
[0079] Refer to Figure 1, the control module of the present invention includes a channel initial information storage module read control unit, a routing unit, a check node extrinsic information storage module read control unit, a variable node extrinsic information storage module read control unit, a variable node information update module extrinsic information output enable control unit, and an output storage sub-module write control unit, which are respectively used to control the reading of the channel initial information stored in the channel initial information storage module, the data writing addresses of the check node and variable node extrinsic information storage modules, the reading of the check node and variable node extrinsic information stored in the check node and variable node extrinsic information storage modules, the output enable of the variable node information update module extrinsic information, and the data writing of the output storage sub-module, so as to ensure the normal operation of the decoder while improving the throughput of the decoder.
[0080] Specifically, for the first-level channel initial information storage sub-module, when the information stored in it is valid and the read data frame does not conflict with the data frame being processed by the decoder or the decoder is idle, the channel initial information storage module read control unit ensures that the channel initial information stored in it can be immediately read; for the second-level channel initial information storage sub-module, the channel initial information storage module read control unit ensures that the channel initial information stored in it is read when the variable node information update module works in each iteration of the corresponding data frame.
[0081] Specifically, the routing unit effectively avoids memory access conflicts by storing the information of the same group of columns or rows in the parity-check matrix at the same address position; the check node and variable node extrinsic information storage module read control unit improves the throughput of the decoder by minimizing the transmission delay of the check node and variable node extrinsic information storage modules; the variable node information update module extrinsic information output enable control unit invalidates the extrinsic information and hard decision information output from the variable node to the check node at the last iteration of the data frame, thus avoiding conflicts between the newly input data frame and the data frame about to complete the iteration; the output storage sub-module write control unit ensures that when the data frame reaches the maximum number of iterations or meets the early termination iteration criterion, the valid hard decision information output by the variable node information update module is timely written into the output storage sub-module.
[0082] Furthermore, the LDPC decoder consists of a channel initial information storage module, a check node extrinsic information storage module, a check node information update module, a variable node information update module, a variable node extrinsic information storage module, and an output storage sub-module to form an improved partially parallel architecture. The improved partially parallel architecture divides the parity-check matrix of the QC-LDPC code according to the number of cyclic permutation matrices (CPMs) therein, and assigns an independent memory to each CPM. At the same time, the improved partially parallel architecture stores the p rows (columns) of extrinsic information of the CPM at the same address in the corresponding memory;
[0083] Among them, p is the extended parallelism of the CPM. For the regular QC-LDPC code, its parity-check matrix is composed of m×n sub-matrices of order L, and the row and column weights of the sub-matrices are of size w.
[0084] Refer to Figure 4 , the decoder of the present invention adopts an improved partially parallel architecture, which is improved based on the Figure 3 traditional partially parallel architecture. The parity-check matrix of the QC-LDPC code is divided according to the number of CPMs, and the external information of each CPM is stored in equal numbers of memories respectively. In addition, the improved partially parallel architecture maps the p rows (columns) of external information of the CPM to the same storage address, further improving the parallel processing ability of the partially parallel decoder and effectively avoiding storage access conflicts. R represents the code rate of the QC-LDPC code.
[0085] Specifically, p is the extended parallelism of the CPM. For the regular QC-LDPC code, its parity-check matrix is composed of m×n sub-matrices of order L, and the row and column weights of the sub-matrices are of size w. The row parallelism of the decoder of the traditional partially parallel architecture is m, and the column parallelism is n. The channel initial information storage module includes 2n RAMs, and the effective capacity of each RAM is L×Q, where Q is the number of quantization bits. The node external information storage unit includes m×n×w RAMs, and the effective capacity of each RAM is L×Q. After improvement, the row parallelism of the decoder is mp, the column parallelism is np, the channel initial information storage module includes 2n RAMs, and the effective capacity of each RAM is The node external information storage unit includes m×n×w RAMs, and the effective capacity of each RAM is
[0086] Furthermore, when the maximum throughput of the LDPC decoder is less than or equal to the rate of the high-speed input data frame, the improved partially parallel architecture adopts an overlapping dual-frame processing strategy. Among them, the row parallelism of the improved partially parallel architecture is mp, the column parallelism is np. The first-level channel initial information storage sub-module and the second-level channel initial information storage sub-module each include n random access memories (RAMs). The check node external information storage module and the variable node external information storage module each include m×n×w RAMs, and the effective capacity of each RAM is Q is the number of quantization bits; when the maximum throughput of the LDPC decoder is greater than the rate of the high-speed input data frame, the improved partially parallel architecture adopts a non-overlapping dual-frame processing strategy. Among them, the improved partially parallel architecture combines the check node external information storage module and the variable node external information storage module into a node external information storage module, and the node external information storage module only includes m×n×w RAMs, and the effective capacity of each RAM is
[0087] Referring to Figure 5 and Figure 6 When the LDPC decoder adopts the overlapping double-frame processing strategy, by optimizing the time utilization rate of the check node and variable node information update modules, parallel processing of two frames of data is achieved. With only an increase in information storage resources and a small amount of control logic resources, and without an increase in computing resources, a significant increase in the decoder throughput is achieved at a lower resource cost. Theoretically, the decoder throughput can be increased to twice the original by adopting the overlapping double-frame processing.
[0088] Specifically, in the LDPC decoder adopting the overlapping double-frame processing, except for the FIFOs in the input buffer module and the output module, all other functional units with storage effects use memories with a depth twice as large The memory is thus divided into two regions, the front and the back, which are used to store odd data frames and even data frames respectively, so as to achieve efficient parallel processing of double frames.
[0089] Specifically, when L cannot be divided evenly by p, the first rem pieces of information of the first address in the partition will be repeated for the rem pieces of information remaining at the th address of the memory with a depth twice as large to simplify the access to information; the same strategy is adopted for the latter partition of the memory and the memory with a depth of where rem represents the remainder of L / p.
[0090] Furthermore, the external information storage module of the check node supports additional storage of the hard decision information of the pseudo a posteriori probability information of the variable node. When the LDPC decoder is compatible with the operation mode of early termination of iteration, the external information storage module of the check node selects a compact storage strategy or a distributed RAM storage strategy. The compact storage strategy is to store the hard decision information and the external information passed by the variable node to the check node in the same memory. Among them, the external information storage module of the check node consists of m×n×w RAMs, and the effective capacity of each RAM is
[0091] Specifically, when adopting the compact storage strategy, referring to Figure 7, the external information storage module of the present invention for check nodes supports additional storage of the hard decision information of the pseudo posterior probability information of variable nodes. Specifically, when the parity check equation calculation module (the parity check update unit (Parity Check Update, PCU) is a processing unit inside the parity check equation calculation module) works synchronously with the check node information update module, the hard decision information input by the internal corresponding unit and the external information passed by the variable nodes to the check nodes are in the same position in the parity check matrix. Therefore, the external information storage module of the check nodes can adopt a compact and efficient storage strategy, storing the hard decision information and the external information passed by the variable nodes to the check nodes in the same memory to save storage and counter resources. Under this storage strategy, the external information storage module of the check nodes consists of m×n×w RAMs, and the effective capacity of each RAM is
[0092] Specifically, when adopting the distributed RAM storage strategy, for FPGA, in the case of shortage of on-chip BRAM resources or low internal utilization rate, the distributed RAM will be used to store the hard decision information.
[0093] Furthermore, for the decoding algorithm of the LDPC decoder, the decoding algorithm is realized by the cooperation of the external information storage module of the check nodes, the check node information update module, the parity check equation calculation module, the external information storage module of the variable nodes, and the variable node information update module to adopt the improved approximate min-sum algorithm (Improved Normalized Approximate Min-Sum Algorithm, IAMSA) with multi-factor correction using flood scheduling. This algorithm is improved from the normalized min-sum algorithm (NMSA); first, in order to further reduce the computational complexity of the check node information update operation, this method proposes the approximate min-sum algorithm (AMSA), which uses a tree-structured minimum value and approximate second minimum value generator to transform the solution of the minimum value function during the check node information update into the solution of the exact minimum value and the approximate second minimum value. In the check node information update module, NAMSA only compensates the exact minimum value and the approximate second minimum value obtained by multiplying the same normalization factor when solving the minimum value function during the check node information update. IAMSA adds different normalization factors and offset factors on this basis to compensate the minimum value and the approximate second minimum value. ), the check node information update formula of the IAMSA algorithm is as follows:
[0094]
[0095] where r ji represents the external information passed from the j-th check node to the i-th variable node, Q j \i represents the set of other variable nodes connected to the j-th check node except the i-th variable node, sign represents the sign function, q ijdenotes the extrinsic information passed from the \(i\)-th variable node to the \(j\)-th check node, max denotes the maximum function, \(\alpha_1\) and \(\alpha_2\) denote normalization factors, amin denotes the function of finding the maximum and approximate second minimum values, and \(\beta_1\) and \(\beta_2\) denote offset factors.
[0096] Referring to Figure 8 and Figure 9 In the present invention, both the check node information update module and the variable node information update module are implemented by a multi-stage pipeline. By reducing the critical path delay of the core processing module, the maximum clock frequency supported by the decoder is increased, and the decoding throughput is improved. Among them, the decoding algorithm of the decoder in the present invention adopts F-IAMSA, which is improved from NAMSA. When updating the check node information, NAMSA only finds the exact minimum value and the approximate second minimum value of the minimum value function to reduce the implementation complexity of the decoder, and at the same time adds the same normalization factor for compensation. IAMSA adds different normalization factors and offset factors on this basis to compensate for the minimum value and the approximate second minimum value to reduce the loss of the error correction performance of the decoder.
[0097] Among them, the check node information update module includes \(mp\) processing units, and the pipeline structure of each processing unit is divided into upper and lower parts. The upper part is responsible for calculating and processing the sign bits of \(w\) r input data V2C, and the lower part is used to calculate and process the absolute values of \(w\) r input data V2C. \(w\) r and \(w\) c respectively represent the row weight and column weight of the parity check matrix. The processing units are divided into 5-level pipelines according to the operation functions. Among them, the first-level pipeline is responsible for extracting the sign bits and absolute values of V2C; the upper part of the second-level pipeline obtains the sum signw r of all sign bit data through exclusive OR operation, and delays \(w\) r sign bits, and the lower part is used to calculate the minimum value min r of \(w\) 1st absolute values, the approximate second minimum value amin 2nd and the index idx corresponding to the minimum value; the upper part of the third-level pipeline performs exclusive OR operations on signw r and \(w\) r sign bits respectively to obtain the sum sign(w r of all other sign bit data except itself, r-1), for the lower part, the results after compensating the minimum value and the approximate second minimum value are calculated respectively, and the index of the minimum value is delayed; when reaching the 4th - stage pipeline, the operation of the sign bit in the upper part has ended, and only need to wait for the absolute - value part operation to be completed and then perform selection and merging. The lower part truncates the previous - stage result and continues to delay the minimum - value index. When truncating, rounding is required; the 5th - stage pipeline will perform selection and merging on the processed sign bit and absolute - value data. If the sign - bit index is equal to the minimum - value index, the absolute - value bit outputs the result after compensating the second minimum value, otherwise, the absolute - value bit outputs the result after compensating the minimum value; when merging, the output of the sign - bit operation serves as the high - order bit of the original - code output, and the output of the absolute - value bit operation serves as the low - order bit of the original - code output; the calculation of converting the original code (for the previous - step original code) to the complement code is processed in the variable - node information update module.
[0098] Further, both the check - node information update module and the variable - node information update module are implemented using multi - stage pipelines. The operations in the 2nd - stage pipeline are refined, and more stages of pipelines are inserted. The multi - stage pipeline uses a tree - structure minimum and approximate sub - minimum value generator (amin) to calculate the minimum value and the approximate sub - minimum value (Tree Structure Minimum andApproximate Sub Minimum Value Generator, M1AM2VG TS ) to calculate the minimum value and the approximate sub - minimum value of w r numbers, and stages of pipelines need to be inserted.
[0099] Specifically, the check - node information update module includes mp processing units, and the pipeline structure of each processing unit is divided into upper and lower parts. The upper part is responsible for calculating and processing the sign bits of w r input data V2C, and the lower part is used to calculate and process the absolute values of w r input data V2C. w r and w c represent the row weight and column weight of the parity - check matrix respectively. The processing units are divided into 5 - stage pipelines according to their operation functions. Among them, the 1st - stage pipeline is responsible for extracting the sign bits and absolute values of V2C; in the upper part of the 2nd - stage pipeline, the sum signw r of all sign - bit data is obtained through exclusive - OR operation, and w r sign bits are delayed. The lower part is used to calculate the minimum value min r , the approximate sub - minimum value amin 1st and the index idx corresponding to the minimum value among w 2nd absolute values; in the upper part of the 3rd - stage pipeline, signw r and wr The exclusive OR operation is performed on each sign bit to obtain w r The sum of all sign bit data except itself, sign(w r -1). For the lower part, the results after compensating the minimum value and the approximate second minimum value are calculated respectively, and the minimum value index is delayed; when reaching the 4th-stage pipeline, the operation of the upper part sign bits has ended, and only delay waiting is required until the absolute value part operation is completed and then selection and merging are performed. For the lower part, the results of the previous stage are truncated, and the minimum value index is continuously delayed. Rounding is required during truncation; at the 5th-stage pipeline, the processed sign bits and absolute value data are selected and merged. If the sign bit index is equal to the minimum value index, the absolute value bit outputs the result after compensating the second minimum value; otherwise, the absolute value bit outputs the result after compensating the minimum value; during merging, the output of the sign bit operation serves as the high bit of the original code output, and the output of the absolute value bit operation serves as the low bit of the original code output.
[0100] Preferably, considering that the pipeline structure of the Variable Node Update (VNU) is simple, in order to balance the pipeline delay and computing resources of the Variable Node Update (VNU) and the Check Node Update (CNU), the calculation of converting the original code to the complementary code is placed in the Variable Node Update unit for processing.
[0101] Preferably, in the 2nd-stage pipeline of the Check Node Update unit, the operation function of the lower part is relatively complex and the critical path delay is large. To further optimize the processing speed of the Check Node Update unit, the operations in this part are refined, more stages of pipelines are inserted, and a tree-structured minimum value and approximate second minimum value generator is used to find the minimum value and approximate second minimum value. The tree-structured minimum value and approximate second minimum value generator refers to Figure 10 to find the minimum value and approximate second minimum value of w r The number of stages of pipelines needs to be inserted to find the minimum value and approximate second minimum value of numbers; in addition, the tree-structured minimum value and approximate second minimum value generator reduces the hardware resource overhead of the decoder by reducing the number of comparators and selectors, and exchanges a small loss of error correction performance for a large reduction in the implementation complexity of the decoder.
[0102] Among them, the variable node information update module includes np processing units, each processing unit adopts a 4-stage pipeline structure, uses l i to represent the initial channel information corresponding to the current variable node, and uses w c the extrinsic information C2V passed from the ji check nodes to the variable node is represented as r i , j ∈ R iDenote the set of other check nodes connected to the \(i\)-th variable node, and the C2V data form is the original code; the first-level pipeline first converts the extrinsic information C2V into the two's complement and delays it by \(l\). i ; the second-level pipeline uses an adder array to calculate the sum \(sum\) of all input data; the third-level pipeline then uses a subtractor array to calculate the differences between \(sum\) and \(w\) C2Vs respectively. c ; the fourth-level pipeline needs to truncate the intermediate calculation results to finally obtain the hard decision information \(c\) and \(w\) c V2Cs;
[0103] Optionally, for the consecutive addition operation of multiple data, the second-level pipeline can be refined;
[0104] Preferably, when the working main frequency of the decoder permits, the number of pipeline stages of the check node and variable node information update units can be reduced.
[0105] Specifically, the variable node information update module includes \(np\) processing units, each processing unit adopts a 4-stage pipeline structure, uses \(l\) i to represent the initial channel information corresponding to the current variable node, and uses the extrinsic information C2V passed from \(w\) c check nodes to the variable node to represent \(r\) ji , \(j\in R\) i , \(R\) i denotes the set of other check nodes connected to the \(i\)-th variable node, and the C2V data form is the original code. It is more convenient to use the two's complement for addition and subtraction operations. Therefore, the first-level pipeline first converts the extrinsic information C2V into the two's complement and delays it by \(l\). i ; the second-level pipeline uses an adder array to calculate the sum \(sum\) of all input data; the third-level pipeline then uses a subtractor array to calculate the differences between \(sum\) and \(w\) c C2Vs; to prevent precision loss, the data bit width of the intermediate calculation results must be wider than that of the input and output data. Therefore, the fourth-level pipeline needs to truncate the intermediate calculation results to finally obtain the hard decision information \(c\) and \(w\) c V2Cs;
[0106] Optionally, for the consecutive addition operation of multiple data, the second-level pipeline can be refined, that is, multiple pipelines are inserted into the second-level pipeline;
[0107] Preferably, when the working main frequency of the decoder permits, the number of pipeline stages of the check node and variable node information update units can be reduced to reduce the number of register uses. For FPGAs, evenly distributing the LUT resources and FF resources helps to improve the utilization efficiency of the FPGA's Slice resources.
[0108] Further, the minimum value and approximate second minimum value generator of the tree structure is calculated in two stages; in the first stage, the input data is divided into N groups, and each group of data divides the pipeline stages according to the minimum value generator of the tree structure, and the minimum value of each group of data is obtained through pairwise comparison; in the second stage, the set of minimum values of each group of data divides the pipeline stages according to the minimum value and second minimum value generator of the tree structure, and the minimum value and approximate second minimum value of all input data are obtained through pairwise comparison.
[0109] The index of the minimum value is generated by the minimum value index generator. 2-MVG1 represents a minimum value generator with 2 inputs and a single minimum value output, and 2-MVG2 represents a minimum value generator with 2 inputs and 2 minimum value outputs. They both output a comparison result cp, and the set CP of comparison results will be used as the input of the minimum value index generator. For 2 k input data, the minimum value and second minimum value generator of the tree structure requires 2 k + N - 3 two-input comparators and 2 k + 2N - 4 two-input selectors, excluding the two-input selectors required by the minimum value index generator; where N is a power of 2. If N = 2 k , the minimum value and approximate second minimum value generator of the tree structure generates accurate minimum and second minimum values.
[0110] The parity check equation calculation module and the parity check node information update module are executed in parallel, including mp processing units. Each processing unit corresponds to a parity check equation and calculates a syndrome component. When all syndrome components are zero, that is, the syndrome vector is a zero vector, it indicates that the decoding result is correct. The calculation of the syndrome component is to find the exclusive OR sum of w r hard decision information; preferably, a counting operation is adopted to determine whether the syndrome vector is a zero vector; optionally, for the exclusive OR sum of multiple data, multi-stage pipelines can be inserted.
[0111] Refer to Figure 7 , the parity check equation calculation module and the parity check node information update module are executed in parallel, including mp processing units. Each processing unit corresponds to a parity check equation and calculates a syndrome component. When all syndrome components are zero, that is, the syndrome vector is a zero vector, it indicates that the decoding result is correct. The calculation of the syndrome component is to find the exclusive OR sum of w r hard decision information.
[0112] Preferably, in LDPC decoding, the length of the syndrome vector is usually long, so a counting operation is adopted to determine whether the syndrome vector is a zero vector to save storage resources.
[0113] Optionally, for the exclusive OR sum of multiple data, multi-stage pipelines can also be inserted to reduce the critical path delay of the module.
[0114] The output module is responsible for converting the stored out-of-order bit data into a decoded output arranged in the correct order. Among them, the bit data corresponding to each sub-matrix block is sequential. In a decoder using a regular QC-LDPC code, when L is not divisible by p, the deletion unit becomes effective to eliminate invalid data.
[0115] In a second aspect, referring to Figure 2 , a control method for a low-complexity high-speed LDPC decoder for the CCSDS near-Earth application standard includes the following steps:
[0116] S1, Data input: The high-speed input data frame is input to the first-level channel initial information storage sub-module of the channel initial information storage module. If the parallelism of the input data frame is higher than the decoder expansion parallelism, the input data frame first needs to pass through the input buffer module for channel conversion and then be input to the first-level channel initial information storage sub-module of the channel initial information storage module;
[0117] S2, Initialization: The decoder uses overlapping dual-frame processing to increase the throughput by 2 times. Except for the FIFOs in the input buffer module and the output module, the other modules with storage functions all use memories with a depth of 2 times of the memory. The memory thus has two front and back partitions, which are used to store odd and even data frames respectively; the first-level channel initial information storage sub-module stores the odd data frames in the input data frame in column order to the previous partition of the internal n RAM block address spaces, and stores the even data frames in the latter partition; at the beginning of the first iteration of each data frame in the high-speed input data frame, the data stored in the first-level channel initial information storage sub-module described in step S1 is read and then written to the second-level channel initial information storage sub-module and written to the check node extrinsic information storage module through a multiplexer; the decoder iteration count is initialized to 0, and the maximum iteration count is determined according to the error correction performance requirements;
[0118] S3, Check node update and check equation calculation: When the output data of the check node extrinsic information storage module described in step S2 is valid, the check node information update module and the check equation calculation module start the check node information update work. Among them, the check node information update module outputs the calculated extrinsic information C2V to the variable node extrinsic information storage module, and the check equation calculation module outputs the judgment flag of whether the codeword satisfies the check equation to the control module; the check node information update module inputs the extrinsic information V2C passed from the variable node to the check node, and is used to calculate the extrinsic information C2V passed from the check node to the variable node. C2V is then controlled by the routing unit and written to the variable node extrinsic information storage module; the check equation calculation module inputs the hard decision information c and is used to calculate the syndrome to check whether the codeword satisfies the check equation;
[0119] S4, variable node update and pseudo-posteriori probability information calculation and judgment: when the variable node external information storage module in step S3 and the second-level channel initial information storage submodule in step S2 output data valid, the variable node information update module starts the variable node information update work, the second-level channel initial information storage submodule outputs the original data frame corresponding to the data frame undergoing decoding iteration to the variable node information update module, and the variable node information update module outputs the calculated external information V2C to the check node external information storage module; the variable node information update module inputs the external information C2V passed by the check node to the variable node and the channel initial information of the data frame, which are used to calculate the external information V2C passed by the variable node to the check node and the hard decision information c of the pseudo-posteriori probability information of the variable node; after the variable node update and the pseudo-posteriori probability information calculation and judgment, the decoder iteration number is increased by 1;
[0120] S5, termination judgment: the update work of the check node information in step S3 and the update work of the variable node information in step S4 together constitute the iteration of the input data frame. When the input data frame reaches the maximum number of iterations or meets the early termination iteration criterion, the data frame terminates the iteration after the current iteration cycle ends, and the valid information in the hard decision information c is written into the output storage submodule under the control of the control module; when the data frame iteration is not terminated, V2C and c first pass through the multiplexer, and then are controlled by the routing unit to be written into the check node external information storage module in step S3;
[0121] S6, decoding result processing and output: read the hard decision information c from the output storage submodule described in step S5 in data frame column order, convert the stored disordered bit data into a decoding output arranged in the correct order, and remove invalid data.
[0122] When L is not divisible by p, the depth is doubled The memory of The remaining rem information of the address will repeat the first rem information of the first address of the partition. Similarly, the next partition of the memory and the depth are The same strategy is adopted for all memories; where rem represents the remainder of L / p.
[0123] The decoder is compatible with two operation modes: fixed iteration and early termination iteration. In the fixed iteration mode, the storage and computing modules related to the early termination iteration mode of the decoder will stop working. When the decoder only needs to support the fixed iteration mode, the storage and computing modules related to the early termination iteration mode will be removed.
[0124] Reference Figure 7, the decoder provided by the present invention is compatible with two operation modes: fixed iteration and early termination iteration. For the early termination iteration mode, when the input data frame reaches the maximum number of iterations or meets the early termination iteration criterion, the data frame will terminate the iteration after the end of the current iteration cycle; if the iteration does not terminate, the extrinsic information V2C passed from the variable node to the check node and the hard decision information c of the pseudo a posteriori probability information of the variable node will pass through the multiplexer and then be written into the check node extrinsic information storage module under the control of the routing unit; if the iteration terminates, the valid information in c will be written into the output storage sub-module under the control of the control module.
[0125] Preferably, for the fixed iteration mode, the storage and calculation modules related to the early termination iteration mode of the decoder will stop working to save power consumption; preferably, in the case where the decoder only needs to support the fixed iteration mode, the storage and calculation modules related to the early termination iteration mode will be removed to further save resources.
[0126] Specifically, the fixed iteration method has the advantages of simple implementation and predictable performance, while the early termination iteration can improve the decoding efficiency, reduce the decoding delay, and effectively save power consumption while avoiding additional calculations. In practical applications, appropriate strategies can be flexibly selected according to the specific requirements of the system.
[0127] For FPGA (or ASIC implementation), the decoder adopts an efficient storage strategy based on the minimum block RAM of digital circuits. Specifically, for Xilinx 7 series FPGAs, the maximum bit width of the 18K block RAM is 36, and this decoder can match the 18K block RAM with a maximum bit width of 36 or an integer multiple of 36, thereby effectively improving the utilization efficiency of storage resources.
[0128] In a third aspect, another object of the present invention is to provide a channel decoding processing device, and the channel decoding processing device is used to implement the low-complexity high-speed LDPC decoder for the CCSDS near-Earth application standard.
[0129] Embodiment:
[0130] Referring to Figure 1 , the embodiment of the present invention provides a low-complexity high-speed LDPC decoder for the CCSDS near-Earth application standard, and the FPGA implementation is carried out for the shortened (8160, 7136) LDPC code adopted by the actual near-Earth communication system. The decoding algorithm adopts F-IAMSA and adopts a (6, 2) quantization scheme, that is, 1 bit for the sign bit, 3 bits for the integer part, and 2 bits for the decimal part. The CCSDS near-Earth application standard adopts a (8176, 7154) regular QC-LDPC code with a code rate of 7 / 8, and its parity-check matrix is composed of 2 × 16 sub-matrices of order 511. The row and column weights of the sub-matrix are 2, the row weight of the entire parity-check matrix is 32, and the column weight is 4.
[0131] Reference Figure 1 Before the data frame iterative decoding, the first-level channel initial information storage sub-module of the embodiment of the present invention will complement the 18 all-0 bits added before the data frame information bit coding and delete the last 2 bits of information to adapt to the shortened (8160, 7136) LDPC code adopted in the actual near-earth communication system. At the same time, the module initializes the information corresponding to these 18 all-0 bits into a complement code data with a sign bit of 0 and the rest of the bits all being 1, so as to ensure that the check node information update module can ignore these bit information and improve the error correction performance.
[0132] Reference Figure 1 For the FPGA implementation, the embodiment of the present invention adopts an efficient storage strategy based on the smallest block RAM of the digital circuit. By optimizing the bit width design of the block RAM, it is made to match the maximum bit width of 36 of the smallest size block RAM of the FPGA as much as possible, thereby effectively improving the utilization efficiency of the storage resources. Among them, the quantization bit number and the extended parallelism of the CPM are both designed to be 6. Thus, the row parallelism of the decoder is 12 and the column parallelism is 96. The channel initial information storage module includes 32 pseudo-dual-port RAMs, and the effective capacity of each RAM is 256×36. The check node and variable node extrinsic information storage modules respectively include 64 pseudo-dual-port RAMs, and the effective capacity of each RAM is 256×42 and 256×36 respectively. The output module includes 2 pseudo-dual-port RAMs, and the effective capacity of each RAM is 256×84. Since 511 cannot be divided evenly by 6, the 5 bits of information remaining at the 86th address of the memory with a depth of 256 will repeat the first 5 bits of information of the first address of the partition to simplify the access to the information; the same strategy is adopted for the latter partition of the memory.
[0133] Reference Figure 1 and Figure 8 The check node information update module of the embodiment of the present invention includes 12 processing units, and the pipeline structure of each processing unit is divided into upper and lower parts. The upper part is responsible for calculating and processing the sign bits of 32 input data V2C, and the lower part is used for calculating and processing the absolute values of 32 input data V2C. The processing units are divided into 5-level pipelines according to the operation functions. Among them, the first-level pipeline is responsible for extracting the sign bits and absolute values of V2C; the upper part of the second-level pipeline obtains the sum sign32 of all sign bit data through exclusive OR operation and delays 32 sign bits, while the lower part is used for calculating the minimum value min 1st and the approximate second minimum value amin 2ndand the index idx corresponding to the minimum value; in the upper half of the third-level pipeline, exclusive OR operations are performed on sign32 and 32 sign bits respectively to obtain the sum sign31 of all other sign bit data except itself, and in the lower half, the results after compensating the minimum value and the approximate second minimum value are obtained respectively, and the minimum value index is delayed; when reaching the fourth-level pipeline, the operation of the sign bits in the upper half has ended, and only delay waiting is required until the absolute value part operation is completed and then selection and merging are performed. In the lower half, the results of the previous stage are truncated, and the minimum value index is continuously delayed. Rounding is required during truncation; in the fifth-level pipeline, selection and merging are performed on the processed sign bits and absolute value data. If the sign bit index is equal to the minimum value index, the absolute value bit outputs the result after compensating the second minimum value, otherwise, the absolute value bit outputs the result after compensating the minimum value; during merging, the output of the sign bit operation is used as the high bit of the original code output, and the output of the absolute value bit operation is used as the low bit of the original code output. The normalization factors α1 and α2 are floating-point numbers between 0 and 1, and the offset factors β1 and β2 are floating-point numbers greater than or equal to 0. Appropriate values of the four can improve the decoding performance. Through MATLAB simulation, it is found that for the (8160, 7136) LDPC code, when α1 and α2 take 0.75, β1 takes 0, and β2 takes 0.25, the F-IAMSA algorithm can obtain performance similar to that of the floating-point F-NMSA algorithm. At the same time, such a selection is also suitable for hardware implementation. The bit error rate curve of the F-IAMSA algorithm refers to Figure 14 , using Quadrature Phase Shift Keying (QPSK) modulation and Additive White Gaussian Noise (AWGN) channel. Under the (6, 2) quantization scheme, the decoding threshold of the F-IAMSA algorithm for 10 iterations is about 4.14 dB (bit error rate BER is 10 -7 ), the performance loss compared with the floating-point F-NMSA algorithm is about 0.05 dB, the performance loss compared with the fixed-point F-NMSA algorithm is about 0.02 dB, and the performance is improved by about 0.02 dB compared with the fixed-point F-NAMSA algorithm.
[0134] Considering that the pipeline structure of the variable node information update unit is simple, in order to balance the pipeline delay and computing resources of the variable node and the check node information update units, the calculation of converting the original code to the complement code is processed in the variable node information update unit. In addition, in the second-stage pipeline of the check node information update unit, the operation function of the lower half is relatively complex and the critical path delay is large. To further optimize the processing speed of the check node information update unit, the operations in this part are refined, more pipeline stages are inserted, and a tree-structured minimum and approximate second-minimum value generator is used to find the minimum and approximate second-minimum values. The tree-structured minimum and approximate second-minimum value generator and its two basic component units are referred to Figures 10 to 13 , inserting 5 pipeline stages is required to find the minimum and approximate second-minimum values of 32 numbers; in addition, the tree-structured minimum and approximate second-minimum value generator reduces the hardware resource overhead of the decoder by reducing the number of comparators and selectors, sacrificing a small loss in error correction performance in exchange for a significant reduction in the implementation complexity of the decoder.
[0135] Referred to Figure 10 and Figure 13 , in the embodiments of the present invention, the tree-structured minimum and approximate second-minimum value generator is divided into 2 stages during calculation. In the first stage, the input data is divided into 4 groups, and each group of data divides the pipeline stages according to the tree-structured minimum value generator, and the minimum value of each group of data is obtained through pairwise comparison; in the second stage, the set of minimum values of each group of data divides the pipeline stages according to the tree-structured minimum and second-minimum value generators, and the minimum and approximate second-minimum values of all input data are obtained through pairwise comparison. The index of the minimum value is generated by the minimum value index generator. For 32 input data, the tree-structured minimum and second-minimum value generator requires 33 two-input comparators and 36 two-input selectors, and the two-input selectors required by the minimum value index generator are not counted. Among them, m 10 and m 11 respectively represent the minimum values of the first and second 2-MVG2 of this stage of the pipeline, and m 20 and m 21 respectively represent the second-minimum values of the first and second 2-MVG2 of this stage of the pipeline.
[0136] Referred to Figure 1 and Figure 9 , in the embodiments of the present invention, the variable node information update module includes 96 processing units, each processing unit adopts a 4-stage pipeline structure, and l i represents the initial channel information corresponding to the current variable node, and the extrinsic information C2V transmitted from 4 check nodes to the variable node is represented by r ji , j ∈ R i , R iDenote the set of other check nodes connected to the \(i\)-th variable node. The C2V data form is in original code. It is more convenient to use two's complement for addition and subtraction. Therefore, the first-level pipeline first converts the extrinsic information C2V to two's complement and delays it by \(l\). i ; The second-level pipeline uses an adder array to calculate the sum \(sum\) of all input data; The third-level pipeline uses a subtractor array to calculate the differences between \(sum\) and 4 C2Vs respectively; To prevent precision loss, the data bit width of the intermediate calculation results must be wider than that of the input and output data. Therefore, the fourth-level pipeline needs to truncate the intermediate calculation results to finally obtain the hard decision information \(c\) and 4 V2Cs.
[0137] Refer to Figure 1 and Figure 7 In the embodiment of the present invention, the check equation calculation module and the check node information update module are executed in parallel, including 12 processing units. Each processing unit corresponds to a check equation and calculates a syndrome component. When all syndrome components are zero, that is, the syndrome vector is a zero vector, it indicates that the decoding result is correct. The calculation of the syndrome component is to find the exclusive OR sum of 32 hard decision information. At the same time, in LDPC decoding, the length of the syndrome vector is 1022. Therefore, a counting operation is adopted to determine whether the syndrome vector is a zero vector to save storage resources.
[0138] Refer to Figure 1 In the embodiment of the present invention, the output module includes an output storage sub-module, two selection combining units, a deletion unit, and a first-in-first-out data buffer. The output module is responsible for converting the stored out-of-order bit data into 8-way decoded outputs arranged in the correct order. Among them, the bit data corresponding to each sub-matrix block is sequential.
[0139] Refer to Figure 1 In the embodiment of the present invention, the synthesis is implemented based on the FPGA of Xilinx XC7VX485TFFG1761-2. The clock frequency constraint is 320 MHz, and the resource occupancy of the decoder is shown in Table 1:
[0140] Table 1 Hardware Resource Occupancy Table of the Decoder
[0141] Resource type Occupancy Total XC7VX485T resources Occupancy ratio % LUT 26991 303600 8.89 LUTRAM 387 130800 0.30 FF 38882 607200 6.40 BRAM 117.50 1030 11.41 Slice 8623 75900 11.36
[0142] For an LDPC decoder of the CCSDS near-Earth application standard, the present invention additionally adds an input buffer module and an output module to shape the input and output data, thereby further improving the flexibility and adaptability of the decoder.
[0143] Under the same clock frequency and quantization scheme, the present invention adopts F-IAMSA, selects 6 for the extended parallelism p of CPM, uses an 8-stage pipeline for the check node information update unit, uses a 3-stage pipeline for the variable node information update unit, and uses a distributed RAM in the extrinsic information storage module of the check node to store hard decision information. The comparison of the decoder implemented with the hardware implementation results of the decoder proposed in the doctoral thesis of Dr. Kang Jing, "Research on LDPC Coding and Decoding Algorithms and High-Efficiency Implementation Technologies for Space-Ground High-Speed Data Transmission" (Kang Jing. Research on LDPC Coding and Decoding Algorithms and High-Efficiency Implementation Technologies for Space-Ground High-Speed Data Transmission [D]. University of Chinese Academy of Sciences (National Space Science Center, Chinese Academy of Sciences), 2021.) is shown in Table 2:
[0144] Table 2 Comparison of the hardware implementation results of the decoder of the present invention and the decoder proposed by Kang Jing
[0145]
[0146] Compared with the high-speed LDPC decoder in Kang Jing's thesis, when the FPGA hardware resources occupied by the decoder of the present invention are respectively saved by about 38.1% of LUT resources, 13.5% of FF resources and 2.8% of BRAM resources compared with the former, the throughput is increased by about 88.2%, and only 0.02 dB is lost in terms of error correction performance. At the same time, the decoder of the present invention also supports two operation modes, and additionally adds an input buffer module and an output module, further improving the flexibility and adaptability of the decoder. Compared with the decoder proposed in Dr. Kang Jing's doctoral thesis, the present invention not only improves the decoding algorithm, but also adopts a better decoding strategy: overlapping dual-frame processing and efficient storage strategy. Therefore, the present invention can achieve a CCSDS near-Earth application standard LDPC decoder with lower complexity and higher throughput while increasing the functions and flexibility of the decoder.
[0147] Under the same clock frequency, quantization scheme and the same single-iteration clock cycle, the present invention adopts F-IAMSA, and the extended parallelism p of CPM is the same as that of the decoder proposed in the doctoral thesis of Dr. Xie Tianjiao, "Research on Key Technologies of Adaptive Transmission and Networking for Space-Ground Integrated Networks" (Xie Tianjiao. Research on Key Technologies of Adaptive Transmission and Networking for Space-Ground Integrated Networks [D]. Northwestern Polytechnical University, 2020.), that is, p = 12. The check node information update unit uses an 8-stage pipeline, and the variable node information update unit uses a 3-stage pipeline. The comparison of the 275 MHz decoder implemented with the hardware implementation results of the decoder proposed by Dr. Xie Tianjiao is shown in Table 3:
[0148] Table 3 Comparison of the hardware implementation results of the decoder of the present invention and the decoder proposed by Xie Tianjiao
[0149]
[0150] In implementation, the input buffer module and the output module of the present invention are removed, and only the output storage sub-module in the output module is retained for fair comparison. Compared with the high-speed LDPC decoder in Xie Tianjiao's thesis, the 275MHz decoder of the present invention saves approximately 9.5% of LUT resources and 25.7% of FF resources in terms of FPGA hardware resource occupancy while suffering a 0.02dB loss in error correction performance. In addition, increasing the pipeline stages of the core module can improve the maximum clock frequency supported by the decoder. When the check node information update unit of the present invention adopts a 9-stage pipeline, the variable node information update unit adopts a 4-stage pipeline, and output registers are added to the memory, the maximum clock frequency supported by the decoder of the present invention can reach up to 320MHz. The hardware implementation results of the 320MHz decoder implemented and the decoder proposed in Xie Tianjiao's doctoral thesis are compared as shown in Table 3. At this time, the theoretical throughput of the decoder of the present invention after 10 iterations can reach 4.477Gbps. Compared with the high-speed LDPC decoder in Xie Tianjiao's thesis, the decoder of the present invention saves approximately 6.6% of LUT resources and 13.7% of FF resources in terms of FPGA hardware resource occupancy compared to the former, while the throughput is increased by approximately 11.8%, and only a 0.02dB loss occurs in terms of error correction performance. Compared with the decoder proposed in Xie Tianjiao's doctoral thesis, the present invention improves the decoding algorithm, adopts a tree structure minimum value and approximate second minimum value generator, and a multi-stage pipeline design, and realizes a higher performance decoder design with a slight loss in error correction performance, which can better meet the requirements of the satellite high-speed data transmission system.
Claims
1. A low-complexity, high-speed LDPC decoder for CCSDS near-earth application standards, characterized in that: It includes an input buffer module, a channel initial information storage module, a control module, a check node external information storage module, a check node information update module, a check equation calculation module, a variable node external information storage module, a variable node information update module and an output module: The input buffer module is used to receive the channel initial information of the high-speed input data frame, and output it to the channel initial information storage module after the path conversion, wherein the high-speed input data frame includes an odd data frame and an even data frame, and the input buffer module includes a first-level input buffer unit and a second-level input buffer unit, which are respectively used to receive the input odd data frame and the even data frame, and perform the path conversion; The channel initial information storage module receives and stores the channel initial information after the high-speed input data frame is converted by the number of channels, and finally outputs it to the variable node information update module; the channel initial information storage module includes a first-level channel initial information storage submodule and a second-level channel initial information storage submodule, and the first-level channel initial information storage submodule writes its internal data into the second-level channel initial information storage submodule; The control module is used to coordinate the operations of different modules of the LDPC decoder; The check node external information storage module is used to store external information (Variable Node to Check Node, V2C) transmitted from the variable node to the check node and hard decision information of pseudo-posteriori probability information of the variable node; The check node information update module is used to calculate the external information that the check node transmits to the variable node; The check equation calculation module is used to calculate the check equation corresponding to the LDPC code check matrix, and to check whether the hard decision information of the pseudo-posteriori probability information of the variable node satisfies the check equation; The variable node external information storage module is used to store the external information transferred from the check node to the variable node (CheckNode to Variable Node, C2V); The variable node information updating module is used to calculate the hard decision information of the external information transmitted by the variable node to the check node and the pseudo-posteriori probability information of the variable node.
2. The LDPC decoder according to claim 1, characterized in that: The control module includes a channel initial information storage module read control unit, a routing unit, a check node external information storage module read control unit, a variable node external information storage module read control unit, a variable node information update module external information output enable control unit and an output storage submodule write control unit: The channel initial information storage module read control unit is used to control the reading of the channel initial information stored in the channel initial information storage module; The routing unit is used to control the data writing addresses of the check node external information storage module and the variable node external information storage module; The check node external information storage module read control unit is used to control the reading of the check node external information stored in the check node external information storage module; The variable node external information storage module read control unit is used to control the reading of the variable node external information stored in the variable node external information storage module; The variable node information update module external information output enable control unit is used to control the output enable of the variable node information update module external information; The output storage submodule write control unit is used to control the data writing of the output storage submodule.
3. The LDPC decoder according to claim 1, characterized in that: The LDPC decoder is composed of an improved partial parallel architecture through a channel initial information storage module, a check node external information storage module, a check node information update module, a variable node information update module, a variable node external information storage module and an output storage submodule. The improved partial parallel architecture divides the check matrix of the QC-LDPC code according to the number of cyclic permutation matrices (CPM) therein, and allocates an independent memory to each CPM. At the same time, the improved partial parallel architecture stores the p rows (columns) of the CPM external information at the same address of the corresponding memory.
4. The LDPC decoder according to claim 1, characterized in that: When the maximum throughput of the LDPC decoder is less than or equal to the rate of the high-speed input data frame, the improved partial parallel architecture adopts an overlapping double-frame processing strategy, wherein the row parallelism of the improved partial parallel architecture is mp, the column parallelism is np, the first-level channel initial information storage submodule and the second-level channel initial information storage submodule respectively include n random access memories (Random Access Memory, RAM), the check node external information storage module and the variable node external information storage module respectively include m×n×w RAMs, and the effective capacity of each RAM is Q is the number of quantization bits; when the maximum throughput of the LDPC decoder is greater than the rate of the high-speed input data frame, the improved partially parallel architecture adopts a non-overlapping double-frame processing strategy, wherein the improved partially parallel architecture merges the check node external information storage module and the variable node external information storage module into a node external information storage module, and the node external information storage module only includes m×n×w RAMs, and the effective capacity of each RAM is 5. The LDPC decoder according to claim 1, characterized in that: The check node external information storage module supports additional storage of hard decision information of pseudo-posteriori probability information of variable nodes. When the LDPC decoder is compatible with the operation mode of early termination of iteration, the check node external information storage module selects a compact storage strategy or a distributed RAM storage strategy. The compact storage strategy is to store the hard decision information and the external information passed to the check node by the variable node in the same memory, wherein the check node external information storage module is composed of m×n×w RAMs, and the effective capacity of each RAM is 6. The LDPC decoder according to claim 1, characterized in that: The decoding algorithm of the LDPC decoder adopts the Improved Normalized Approximate Min-Sum Algorithm (IAMSA) of flood scheduling. The IAMSA algorithm verifies the node information update formula as follows: Among them, r ji represents the external information passed from the jth check node to the ith variable node, Q j \i represents the set of other variable nodes connected to the jth check node except the i-th variable node, sign represents the sign function, q ij It represents the external information passed from the i-th variable node to the j-th check node, max represents the maximum value function, α1 and α2 represent the normalization factors, amin represents the maximum value and approximate minimum value functions, β1 and β2 represent the offset factors.
7. The LDPC decoder according to claim 1, characterized in that: The check node information update module and the variable node information update module are both implemented by a multi-stage pipeline, and the multi-stage pipeline uses a tree structure minimum value and approximate sub-minimum value generator (amin) to calculate the minimum value and the approximate sub-minimum value (Tree Structure Minimum and Approximate Sub Minimum Value Generator, M1AM2VG TS ), the minimum value and approximate sub-minimum value generator of the tree structure are used to find w r The minimum and approximate minimum of the number of values need to be inserted Level pipeline.
8. The LDPC decoder according to claim 7, characterized in that: The tree structured minimum value and approximate sub-minimum value generator is divided into a first stage and a second stage during calculation; in the first stage, the input data is divided into N groups, each group of data is divided into pipeline stages according to the tree structured minimum value generator, and the minimum value of each group of data is obtained by pairwise comparison; In the second stage, the minimum value set of each group of data is divided into pipeline stages according to the minimum value and second minimum value generator of the tree structure, and the minimum value and approximate second minimum value of all input data are obtained by pairwise comparison.
9. A control method for a low-complexity, high-speed LDPC decoder for CCSDS near-earth application standards, based on the LDPC decoder of claim 1, characterized in that: The following steps are involved: S1: high-speed input data frame is input into the first-level channel initial information storage submodule of the channel initial information storage module; S2: At the beginning of the first iteration of each data frame in the high-speed input data frame, the data stored in the first-level channel initial information storage submodule in step S1 is read and written into the second-level channel initial information storage submodule and written into the check node external information storage module through the multiplexer; S3: When the output data of the check node external information storage module in step S2 is valid, the check node information update module and the check equation calculation module start the check node information update work, wherein the check node information update module outputs the calculated external information C2V to the variable node external information storage module, and the check equation calculation module outputs the judgment flag of whether the codeword satisfies the check equation to the control module; S4: When the output data of the variable node external information storage module in step S3 and the second-level channel initial information storage submodule in step S2 are valid, the variable node information update module starts the variable node information update work, the second-level channel initial information storage submodule outputs the original data frame corresponding to the data frame undergoing decoding iteration to the variable node information update module, and the variable node information update module outputs the calculated external information V2C to the check node external information storage module; S5: the update work of the check node information in step S3 and the update work of the variable node information in step S4 together constitute the iteration of the input data frame. When the input data frame reaches the maximum number of iterations or meets the early termination iteration criterion, the data frame terminates the iteration after the current iteration cycle ends, and the valid information in the hard decision information c is written into the output storage submodule under the control of the control module; When the data frame iteration is not terminated, V2C and c first pass through the multiplexer, and then are controlled by the routing unit to be written into the check node external information storage module in step S3; S6: Read the hard decision information c from the output storage submodule in step S5 in data frame column order, convert the stored disordered bit data into a decoded output arranged in the correct order, and remove invalid data.
10. A channel decoding processing device, based on the LDPC decoder according to claim 1, characterized in that: The channel decoding processing device is used to implement the low-complexity high-speed LDPC decoder oriented to the CCSDS near-earth application standard.
Citation Information
Patent Citations
LDPC (Low Density Parity Check) decoder for improving decoding efficiency and throughput of decoder
CN115580309A
High-energy-efficiency LDPC decoder for high-speed satellite link
CN115664584A
Realization method for QC-LDPC (Quasi-Cyclic Low-Density Parity-Check) decoder for improving node processing parallelism
CN103220003A
Implementation method for partially-parallel LDPC decoder
CN104702292A
LDPC high-speed coding device based on CCSDS deep space communication standard
CN116318176A