LDPC encoder based on DVB-S2 protocol, transmitting terminal and test platform

By adopting the cyclic shift matrix parallelization algorithm, multi-port RAM storage architecture and check matrix compatibility mechanism in the DVB-S2 transmitting end system, the problems of high processing delay and low resource utilization in traditional systems in high data rate and large-capacity transmission scenarios are solved, and efficient data transmission and low resource consumption are achieved.

CN120185622AActive Publication Date: 2025-06-20FUDAN UNIVERSITY

Patent Information

Application Number
CN202510670382.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-20
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

In high data rate and large-capacity transmission scenarios, traditional DVB-S2 transmitter system has problems such as high processing delay and low hardware resource utilization, which is difficult to meet the needs of satellite communication systems for data transmission rate and reliability.

Method used

The optimized parallelization algorithm based on cyclic shift matrix is ​​adopted, combined with the multi-port RAM storage architecture and parity address blocking mechanism, compatible with ATSC 3.0 and DVB-S2 verification matrix, a complete DVB-S2 transmitting end system is designed, and an FPGA-upload computer verification platform is built to achieve comprehensive testing of system performance.

Benefits of technology

It improves the throughput of the encoder, reduces hardware resource usage, improves bit error rate performance, and achieves high data transmission rate and low resource utilization rate, which is suitable for high data rate and large-capacity transmission scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120185622A_ABST
    Figure CN120185622A_ABST
Patent Text Reader

Abstract

The invention discloses an LDPC encoder based on a DVB-S2 protocol, a transmitting terminal and a test platform, and the LDPC encoder comprises an information bit matrix multi-port input RAM which is used for storing input information bits; the check matrix multi-port ROM is used for storing a check matrix; the accumulative sum matrix RAM is used for storing the accumulative sum matrix generated by the accumulative sum matrix calculation module; the check bit matrix RAM is used for storing a check bit matrix; the check matrix calculation module reads a check matrix from the check matrix multi-port ROM, and the check matrix is converted into a cyclic shift matrix through preprocessing; the accumulative sum matrix calculation module is used for reading information bits from the information bit matrix multi-port input RAM, executing loop right shift and accumulative operation, and calculating to obtain an accumulative sum matrix; and the check bit matrix initialization module is used for calculating a check bit matrix initialization vector according to a calculation result of the accumulation and matrix calculation module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure belongs to the field of satellite communication technology, and particularly relates to an LDPC encoder, a transmitting end, and a test platform based on the DVB-S2 protocol. Background Art

[0002] With the rise of low-earth orbit satellite communication technology, satellite digital video broadcasting has brought new application scenarios and market demands. In this context, the DVB-S2 (Digital Video Broadcasting - Satellite Second Generation) protocol, as the core of satellite broadcast communication, plays a crucial role.

[0003] The DVB-S2 protocol was officially released by the European Telecommunications Standards Institute, aiming to improve the spectral efficiency and data transmission rate of satellite communication and replace the first-generation DVB-S standard. The DVB-S2 introduced a number of key technologies, including LDPC (Low-Density Parity-Check) coding, BCH (Bose-Chaudhuri-Hocquenghem) coding, high-order modulation (such as 8PSK, 16APSK, 32APSK), and Adaptive Coding and Modulation (ACM) technology. The application of these technologies has made significant progress in spectral utilization and system performance of DVB-S2, especially in high data rate and high-reliability transmission. Among them, LDPC coding, as a key technology with unique advantages, is mainly reflected in the following three points: The first point is that the frame length is relatively long. The LDPC coding frame length in the DVB-S2 protocol can reach up to 64,800 bits at most, while the maximum code block length of LDPC coding in the 5G NR protocol is usually 8,448 bits. The relatively long frame length improves the spectral utilization rate and data transmission efficiency; The second point is the unique ACM mechanism. The DVB-S2 protocol supports adaptive coding and modulation technology, which can dynamically adjust the LDPC code rate and modulation method according to the channel conditions, improving the reliability and efficiency of the system. Some other protocols, such as the ATSC 3.0 protocol, do not have such a flexible ACM mechanism; The third point is the special structure of the parity-check matrix. The LDPC-coded parity-check matrix in the DVB-S2 protocol adopts a cyclic structure, which is convenient for hardware implementation, simplifies the hardware circuit design, and reduces the hardware complexity. In contrast, the LDPC-coded parity-check matrix in the 5G NR protocol adopts a polygon-based structure, and the hardware implementation is relatively complex. These unique advantages enable the DVB-S2 protocol to exhibit higher performance and adaptability in satellite communication.

[0004] Considering that the Field Programmable Gate Array (FPGA) has significant advantages in scenarios such as high-throughput data processing and hardware acceleration, it has become an ideal choice for implementing the DVB-S2 transmitter system. The advantages of FPGA are mainly reflected in the following aspects: (1) High degree of customization: FPGA allows designers to customize the hardware architecture according to specific application requirements and enables targeted optimization. For example, in the DVB-S2 system, designers can flexibly adjust the hardware structure according to different coding rates and modulation methods to achieve the best performance.

[0005] (2) Rapid prototyping: FPGA supports rapid prototyping. Designers can implement and test the hardware design in a short time, accelerating the product development cycle. This is particularly important when verifying new technologies and updating standards, and can quickly respond to market demands.

[0006] (3) Parallel processing ability: FPGA inherently supports large-scale parallel processing and can execute multiple tasks simultaneously, which is particularly crucial for applications that require high throughput (such as LDPC encoders) and can significantly improve the data processing speed.

[0007] (4) Reconfigurability: FPGA supports online reconfiguration, allowing dynamic adjustment of hardware functions during system operation, providing flexibility and scalability for satellite communication systems, and enabling adaptation to different task requirements and channel conditions. Summary of the Invention

[0008] The present disclosure relates to a high-speed data transmission transmitter based on the DVB-S2 protocol and its implementation using FPGA. The purpose of the present disclosure is to improve the data transmission rate and reliability in satellite communication systems. By optimizing the parallel LDPC encoder and the DVB-S2 transmitter system, problems such as high processing delay and low hardware resource utilization in traditional solutions in high data rate and large-capacity transmission scenarios are solved.

[0009] The content implemented in the present disclosure includes: An optimized parallelization algorithm based on a cyclic shift matrix to improve the throughput of the encoder; Multi-port RAM storage architecture and parity address block mechanism reduce the occupancy of hardware resources; Compatible with the check matrices of ATSC 3.0 and DVB-S2, improving the bit error rate performance; Complete DVB-S2 transmitter system design, supporting multiple modulation methods and code rate configurations; Build an FPGA-host computer verification platform to achieve a comprehensive test of the system performance.

[0010] After verification, the transmitter of the present disclosure can achieve a throughput of 23.4 Gb / s at a clock frequency of 100 MHz, with low resource utilization, and the bit error rate performance is improved at different code rates. Brief Description of the Drawings

[0011] By referring to the accompanying drawings and reading the following detailed description, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, wherein: Figure 1 Check matrix diagram of the IRA type LDPC code according to one embodiment of the present disclosure.

[0012] Figure 2 Check matrix diagram of the MET type LDPC code according to one embodiment of the present disclosure.

[0013] Figure 3 Schematic diagram of the preprocessing of the check matrix according to one embodiment of the present disclosure.

[0014] Figure 4 Schematic diagram of the appendix of the LDPC check matrix of the 2 / 5 code rate of the DVB-S2 ordinary frame according to one embodiment of the present disclosure.

[0015] Figure 5 Schematic diagram of the coe file obtained after preprocessing the check matrix of the 6 / 15 code rate of the ATSC3.0 ordinary frame according to one embodiment of the present disclosure.

[0016] Figure 6 Flowchart of the preprocessing algorithm of the check matrix according to one embodiment of the present disclosure.

[0017] Figure 7 Hardware architecture diagram of the LDPC parallel encoder according to one embodiment of the present disclosure.

[0018] Figure 8 Flowchart of the implementation of the check matrix calculation module according to one embodiment of the present disclosure.

[0019] Figure 9 Flowchart of the implementation of the sum matrix calculation module according to one embodiment of the present disclosure.

[0020] Figure 10 Overall architecture diagram of the DVB-S2 transmitter system and interface driver according to one embodiment of the present disclosure.

[0021] Figure 11 Architecture diagram of the DVB-S2 transmitter system according to one embodiment of the present disclosure.

[0022] Figure 12 Structure diagram of the linear feedback shift register according to one embodiment of the present disclosure.

[0023] Figure 13 Structure diagram of the data buffer FIFO module according to one embodiment of the present disclosure.

[0024] Figure 14 Comparison experimental diagram of FPGA output and Matlab output at different code rates according to one embodiment of the present disclosure.

[0025] Figure 15 Comparison diagram of output data constellation diagram and IQ data under different modulation methods and code rates according to one embodiment of the present disclosure.

[0026] Figure 16 Comparison diagram of bit error rate performance of DVB-S2 and ATSC parity check matrices at different code rates according to one embodiment of the present disclosure. Specific implementation manners

[0027] After studying the existing solutions, the present disclosure finds that in practice, there are still the following challenges in implementing an FPGA transmitter based on the DVB-S2 protocol: 1) Limitations of current coding algorithms.

[0028] Although the DVB-S2 protocol supports multiple code rates and modulation methods, in high data rate and large-capacity transmission scenarios, traditional serial coding schemes usually have limitations such as high processing delay and low hardware resource utilization. For example, in low-earth orbit satellite communication, the high-speed data transmission between the satellite and the ground station poses extremely high requirements on the throughput and efficiency of the coding algorithm, while traditional coding schemes often fail to meet these requirements. In addition, existing parallel LDPC encoder schemes still have room for optimization in terms of parallelism, throughput, and bit error rate. For example, the highest parallelism of the latest parallelization methods in the literature is only 360, and there is no research on the parity check matrix to achieve a lower bit error rate. In addition, most of these research algorithms have not been fully tested in a complete DVB-S2 transmitter system, which also limits their effectiveness in practical applications.

[0029] 2) Challenges in FPGA engineering implementation.

[0030] Implementing a DVB-S2 transmitter system based on FPGA, especially a high-throughput LDPC encoder, also faces some engineering challenges.

[0031] Firstly, the satellite hardware resources are limited. The hardware resources on the satellite (such as storage units, logic gates, etc.) are very limited, while high-throughput LDPC encoders usually require a large number of parallel computing units and storage resources. How to optimize the encoder design under resource constraints, achieve high performance while reducing hardware costs, is a key issue. It is necessary to maximize the resource utilization efficiency through ingenious architecture design and resource reuse techniques. For example, although high throughput is achieved in some literature, the resource utilization rate is too high and it is not convenient to be deployed into the FPGA.

[0032] Secondly, there are requirements for circuit design at high clock frequencies. In the FPGA implementation, the increase in clock frequency can significantly improve data throughput, but at the same time, it also poses higher requirements for the timing constraints of circuit design. Especially at high clock frequencies above 100 MHz, the operations within each clock cycle must be completed within an extremely short time. This requires careful optimization of the circuit path, reduction of critical path delay, and ensuring the stable operation of the system. In addition, high clock frequencies may also lead to increased power consumption and heat dissipation problems, further increasing the complexity of the design.

[0033] Therefore, the purpose of the present disclosure is to provide a design and FPGA implementation method for a high-speed data transmission transmitter based on the DVB-S2 protocol to solve problems such as high processing delay and low hardware resource utilization rate in existing solutions under high data rate and large-capacity transmission scenarios.

[0034] According to one or more embodiments, a design and FPGA implementation method for a high-speed data transmission transmitter based on the DVB-S2 protocol specifically includes the following steps: (1) An optimized parallelization algorithm based on a cyclic shift matrix to improve the throughput of the encoder; (2) A multi-port RAM storage architecture and an odd-even address block mechanism to reduce the occupancy of hardware resources; (3) Compatibility with the ATSC 3.0 and DVB-S2 parity-check matrices to improve the bit error rate performance; (4) A complete DVB-S2 transmitter system design supporting multiple modulation methods and code rate configurations; (5) Building an FPGA-host computer verification platform to achieve a comprehensive test of the system performance.

[0035] The optimized parallelization algorithm based on the cyclic shift matrix improves the throughput of the encoder by optimizing the construction of the parity-check matrix and the encoding process. The specific steps include: (1) Preprocess the parity-check matrix and convert it into the form of a cyclic shift matrix; (2) Utilize the properties of the cyclic shift matrix to reduce the clock cycles required for parallelized encoding calculations; (3) In hardware implementation, support parallel operations through multi-port read / write RAM and multi-port read ROM.

[0036] The multi-port RAM storage architecture and the odd-even address block mechanism avoid redundant storage by constructing a multi-port RAM and solve the parallel read / write conflict by using the odd-even address block mechanism. The specific design includes: (1) Customize a multi-port read single-port write RAM and a multi-port read ROM; (2) Use an array of storage units to enable simultaneous access to corresponding data by different address lines; (3) Store the sum-accumulation matrix and the parity-check matrix into two RAMs according to the parity of the addresses respectively.

[0037] The compatible ATSC 3.0 and DVB-S2 parity-check matrix further optimizes the error correction ability of the encoder by replacing the parity-check matrix of some code rates in the DVB-S2 protocol with the parity-check matrix of the ATSC protocol. The specific implementation includes: (1) Compare the performance of LDPC encoding in the DVB-S2 and ATSC 3.0 protocols; (2) Select the parity-check matrix corresponding to the code rate in the ATSC 3.0 protocol for replacement; (3) In hardware implementation, generate the adapted parity-check matrix data through software preprocessing.

[0038] The complete DVB-S2 transmitter system design includes a BB Frame module, an FEC module, a bit interleaving and mapping module, and a PL Frame module, supporting multiple modulation methods such as QPSK, 8PSK, 16APSK, and 32APSK and flexible code rate configurations. The specific design includes: (1) The BB Frame module realizes the encapsulation of the baseband frame, including adding a baseband frame header, CRC encoding, and baseband scrambling; (2) The FEC module realizes forward error correction encoding, including BCH encoding and LDPC encoding; (3) The bit interleaving and mapping module realizes bit interleaving and constellation mapping; (4) The PL Frame module realizes the generation of the physical layer frame, including adding a PL frame header, PL scrambling, header modulation, and pilot insertion.

[0039] The construction of the FPGA-host computer verification platform conducts implementation verification and performance testing on the system through hardware loopback testing. The specific steps include: (1) Implement an interface driver module in the FPGA, which is responsible for transmitting the encoded data to the host computer in real time through UDP packets; (2) Generate random information bits in the host computer and send the data to the encoder in the FPGA through the interface driver module; (3) Transmit the encoded data back to the host computer through UDP packets and compare it with the output result of the Matlab toolbox function to verify the correctness of the system implementation; (4) Test key indicators such as the throughput, bit error rate performance, and resource utilization rate of the system to evaluate the system performance.

[0040] According to one or more embodiments, to address the above challenges of improving parallelism and throughput, the present disclosure proposes an LDPC encoder based on the DVB-S2 protocol, which realizes throughput improvement through an optimized parallelization algorithm based on a cyclic shift matrix. The specific method is as follows: 1) DVB-S2 protocol LDPC encoding algorithm: LDPC encoding is roughly divided into two types: IRA type (Irregular Repeat Accumulate) and MET type (Multi-Edge Type). Respectively as Figure 1 and Figure 2 shown. Among them, represents the information bit length , , respectively represent check bits of different lengths, matrix is mainly used for encoding information bits, matrix's main function is for rate matching, matrix and matrix are responsible for the association between information bits at different positions, matrix is a zero matrix used to introduce cyclic characteristics, matrix is an identity matrix, and the above matrices are all extended from the base matrix specified by the protocol. IRA type LDPC encoding is mainly applied to the DVB-S2 protocol. Due to its regular construction characteristics, this encoding method is convenient for hardware implementation. MET type LDPC encoding is mainly applied to the 5G NR protocol, and this encoding method has high flexibility and adaptability in mobile communication. The ATSC 3.0 protocol uses both of these LDPC encodings at the same time. The MET type is mainly used for low code rates, while the high code rate uses the IRA type.

[0041] The LDPC encoding in the DVB-S2 protocol uses the IRA type, and the encoded codeword is a systematic code. Therefore, the input information bit I can be set as , and the encoded check bit P is , the finally obtained codeword C is:

[0042] Among them, k is the length of the information bits, r is the length of the parity bits, and n = k + r represents the total length of the codeword after encoding. According to the encoding principle, the parity bits can be calculated by formula (2):

[0043] where H is the parity-check matrix of dimension and consists of two parts:

[0044] is constructed according to the protocol matrix, is a lower echelon matrix of dimension Expanding the two gives:

[0045] At the same time, since x + x = 0 in the Galois field GF(2), formula (4) can be changed to:

[0046] Observing the pattern in equation (5), the cumulative sum can be set as where :

[0047] Therefore, equation (5) can be simplified to:

[0048] According to the iterative calculation of equation (7), the final parity-bit calculation formula can be obtained:

[0049] This algorithm can be summarized as Algorithm 1:

[0050] Therefore, the LDPC encoding algorithm can be divided into two main steps: First, calculate the value of each parity bit, and second, accumulate these parity bits. Therefore, the DVB-S2 LDPC code is called an irregular repeat-accumulate (IRA) code. However, this simple encoding method has obvious deficiencies in efficiency because it can only output one parity bit each time, resulting in a slow speed. In addition, from the perspective of hardware complexity and resource utilization, this method is also infeasible because it is necessary to store a matrix of size and perform operations on the entire message of k bits with the matrix Perform multiplication between one row of k bits. It can be seen that a major drawback of this algorithm is that it does not fully utilize the sparse structure of the matrix Therefore, in order to better utilize the sparsity of the parity-check matrix, some scholars have proposed an IRA-type parallel algorithm

[0051] 2) IRA-type parallel encoding algorithm: a) Cumulative sum matrix calculation In the parallel implementation of encoding technology, it can be mainly divided into two strategies: one is the parallelization using multiple encoders. By configuring multiple independent encoder modules on the FPGA, different data streams can be processed simultaneously, thus improving the overall encoding rate. The other is to achieve parallelization through formula theory derivation. By analyzing the mathematical model of the encoding algorithm, the parts that can be executed in parallel are identified, and then the corresponding parallel structure is constructed in the hardware. This method not only optimizes the execution efficiency of the algorithm but also effectively reduces the resource occupancy and is applicable to complex encoding standards. Therefore, the IRA-type parallel encoding algorithm mainly performs parallel hardware implementation based on the mathematical model of the encoding algorithm. Reading the DVB-S2 protocol, it can be known that when constructing the parity-check matrix of LDPC encoding, it is grouped according to 360, so corresponding grouping can also be considered for the input data:

[0052] Among them, Group by the length of the information bits. For each item ( ) represents a row vector of length , specifically:

[0053] According to the LDPC encoding principle, it can be known that the key to encoding lies in solving the value of the cumulative sum s. Therefore, s can also be grouped and constructed as follows:

[0054] Here That is, group according to the length of the parity bits. At the same time, for each row vector in the matrix, its bit width is 360 bits, where : :

[0055] According to formula (6), it can be known that is the cumulative sum of the products of the information bits and the corresponding row elements of the parity-check matrix. And since the parity-check matrix is a sparse matrix, a large number of cumulative sum terms with a product of 0 can be ignored, and only the elements of the parity-check matrix with a value of 1, that is The corresponding information bits can be accumulated and calculated. At the same time, the appendix of the protocol stipulates the LDPC check matrix information corresponding to different frame lengths and different code rates, and specifies the row index of the element with the first column value of 1 in each group of the check matrix. The corresponding row indices of the other columns can be obtained by offsetting the row index of the previous column by q bits. At this time, it is noted that the row vector is exactly composed of the accumulated sum of each item offset by q bits multiple times. Therefore, the row vector can be constructed in the way specified by the protocol. However, it should be noted that the first row index of each group specified in the appendix does not necessarily match the subscript of the first item of the row vector. Therefore, the input data needs to be circularly shifted to the right for alignment during calculation. Define the function to represent circularly shifting the row vector I to the right by z bits, and set as the set of row indices of the first column in the x-th group specified in the appendix. Then the above algorithm can be summarized as Algorithm 2:

[0056] b) Calculation of the check bit matrix After calculating the accumulated sum matrix, according to Equation (8), it can be seen that the check bits can be obtained through the accumulated sum vector. Therefore, in order to match the parallelization requirement, the check bits also need to be grouped to construct a matrix :

[0057] where each item of the row vector can be deduced from the accumulated sum matrix:

[0058] Here , for its first row vector, it needs to be calculated separately. It is noted that:

[0059] Therefore, the initialization row vector needs to be solved, and this vector can be obtained from the accumulated sum matrix. Let the row vector T be:

[0060] where L is a lower triangular matrix with a dimension of , as shown in Equation (17), is the transpose of the accumulated sum vector of all rows of the accumulated sum matrix.

[0061]

[0062] At this time, let the function be to logically shift the row vector I to the right by z bits. Then it can be observed that shifting the row vector T to the right by one bit is the initialization vector:

[0063] Therefore, the calculation algorithm of the parity-check matrix can be summarized as Algorithm 3:

[0064] After calculating the parity-check matrix, all encoded parity-check bits can be obtained. However, there are still two deficiencies in the above parallelization algorithm: The first is that there is still room for optimization in the clock cycle consumption of the accumulation sum matrix during encoding calculation. Since the accumulation sum matrix needs to be stored in RAM when implemented in hardware, and when calculating its row vectors, it needs to be calculated according to the row index values in the appendix of the DVB-S2 protocol, and the positions of these index values are not necessarily arranged in order. Therefore, the row vectors stored in the same RAM address may be calculated multiple times. So, two clock cycles are required for reading and writing in one calculation, but only one clock cycle is required for writing to the RAM storing the parity-check matrix when calculating the parity-check matrix. Let the number of row index values of the single code rate in the appendix be N. Therefore, the minimum number of clock cycles required to calculate the entire parity-check matrix and the parity-check matrix is at least ; The second is that the parity-check bits stored in the calculated parity-check matrix are arranged column by column, which is not conducive to the RAM outputting the parity-check bits in the order of row addresses. Therefore, it is necessary to reorder the results to meet the requirements of parallel output. Therefore, to solve the above deficiencies, the present disclosure proposes a cyclic shift parallelization encoding algorithm.

[0065] c) Cyclic shift parallelization encoding algorithm Preprocessing of the parity-check matrix: According to the properties of cyclic shift, it can be seen that it is convenient for hardware implementation, that is, when calculating the product, only the subscripts of the non-zero elements in the first row need to be known, and then the corresponding cyclic right shift of the vector can be performed. Therefore, it can be considered to perform corresponding processing on the LDPC parity-check matrix to transform it into a cyclic shift matrix to reduce the clock cycles required for parallelization encoding calculation. Since the LDPC parity-check matrix in the DVB-S2 protocol is of the IRA type and has the characteristics of block structure, the parity-check matrix is spliced every q rows to construct a submatrix , where :

[0066] At this time, according to the block length of 360, it is divided into t groups of submatrices by columns , where :

[0067] It can be seen that by constructing the above transformation, the IRA-type parity-check matrix is converted into a matrix composed of submatrices , where each sub - matrix is a cyclic shift matrix of dimension:

[0068] It can be seen more intuitively from Figure 3 .

[0069] In equation (9), the input information bits have been grouped, and an accumulation sum matrix has been constructed through equation (11). At this time, it is noted that the first - row vector in the accumulation sum matrix, the grouped information bits, and the transformed parity - check matrix satisfy the expression:

[0070] Therefore, it can be deduced that other row vectors in the accumulation sum matrix also satisfy the expression:

[0071] Note that is a cyclic shift matrix of dimension, and each matrix has multiple non - zero elements in the first row. At the same time, due to the sparsity of the parity - check matrix, the transformed matrix will have many zero matrices. Calculating zero matrices still results in zero matrices. Therefore, it is necessary to find the that are not zero matrices for calculation. Let the set represent the non - zero elements contained in the cyclic matrix corresponding to the calculation of and its corresponding th group of input data . Then, according to the properties of the cyclic shift matrix:

[0072] Observing equation (24), it can be seen that if the information bits and the information of the transformed parity - check matrix are pre - stored, the calculation of the row vectors can be executed in parallel. And since the parity - check matrix is also calculated based on the accumulation sum matrix, according to the above derivation, the parity - check matrix can be calculated in parallel at the same time. Here, it is assumed that the parallelism parameter is m = 2. Then, the LDPC optimized encoding algorithm in the DVB - S2 protocol can be summarized as Algorithm 4:

[0073] It can be seen from the flow of this Algorithm 4 that the total number of clock cycles required to complete encoding is , because when calculating the accumulation sum matrix, since the set is obtained through matrix pre - processing, each row vector can be calculated in sequence, and it can be directly written into the RAM after calculation. So only one write clock cycle is required. In addition, since it is sequential calculation and different row vectors and can be computed in parallel. Therefore, the final number of clock cycles is reduced from to . If the parallelism parameter m = 2 is taken here, the speed of calculating the cumulative sum matrix can be increased by about four times, and the speed of calculating the parity bits can be increased by about two times. However, the speed improvement of this algorithm comes at the cost of space, that is, all information bits of the current frame need to be stored during encoding. In addition, for parallel computing and RAM with multi-port read / write and ROM with multi-port read need to be implemented in hardware.

[0074] Furthermore, in order to reduce the bit error rate, the embodiments of the present disclosure propose to establish an ATSC 3.0 and DVB-S2 parity check matrix compatibility mechanism to reduce the bit error rate. The specific method is as follows: a) Parity check matrix replacement: In different communication protocols, the implementation and application of LDPC coding are different, such as the DVB-S2 protocol and the ATSC3.0 protocol. These two protocols are respectively for satellite TV and terrestrial digital TV, and different LDPC coding schemes are adopted to adapt to their specific application scenarios and requirements. The ATSC 3.0 standard adopts two different structures, namely the IRA type used at high code rates and the MET type structure used at low code rates, while the DVB-S2 protocol only uses one, that is, the IRA structure. When researching and experimentally comparing the LDPC coding performance, it is found that the ATSC 3.0 protocol is better than the DVB-S2 protocol in terms of the SNR required when reaching FER = and the gap from the capacity limit in the AWGN channel.

[0075] Compared with the LDPC coding in the DVB-S2 protocol, the LDPC coding in the ATSC 3.0 protocol requires a lower signal-to-noise ratio and is closer to the channel capacity limit when reaching a specific bit error rate under the same conditions. Therefore, in the hardware implementation of DVB-S2, the parity check matrix can be replaced with the LDPC coding parity check matrix of the ATSC 3.0 protocol, so as to achieve more efficient signal transmission and stronger anti-interference ability. However, due to the differences in LDPC coding between the two protocols, only the same parts are selected for replacement. It should be noted that the code rates 6 / 15, 9 / 15, 10 / 15, and 12 / 15 in the ATSC 3.0 protocol are all IRA structures and just correspond to the code rates 2 / 5, 3 / 5, 2 / 3, and 4 / 5 in the DVB-S2 protocol respectively. Therefore, these code rates are selected for corresponding replacement.

[0076] Further, in order to implement the above method on FPGA and reduce resource utilization, the present disclosure implements the parallel LDPC algorithm proposed in the above Algorithm 4 on FPGA and designs a multi-port RAM storage architecture and an odd-even address block mechanism to reduce resource utilization. The specific method is as follows: 1) Software architecture design The software architecture part mainly preprocesses the LDPC parity-check matrix by writing code in Python in advance and replaces the parity-check matrix with the corresponding code rate using the parity-check matrix of the ATSC 3.0 protocol. After the calculation is completed, a coe file is generated and stored in the ROM. The parity-check matrix stored in the ROM can be quickly accessed during encoding, avoiding the delay caused by real-time calculation.

[0077] a) Design of parity-check matrix data structure: The structure of the LDPC encoding parity-check matrix is specified in detail in the appendix of the DVB-S2 protocol document. According to different frame lengths and code rates, corresponding parameters are set in the appendix to meet various application requirements. According to the introduction of the encoding principle in the DVB-S2 protocol, the LDPC parity-check matrix is divided into t groups with 360 columns in each group. This grouping method makes the management and processing of the parity-check matrix more efficient. Specifically, each row of data in the appendix represents the index of the first column value of 1 in each group. These indexes not only indicate the key positions of the parity-check bits but also provide the necessary basis for the subsequent calculation of the parity-check bits. The indexes of the remaining 360 - 1 column values of 1 are obtained by offset calculation of the index value of the first column, and the offset is This structured design effectively reduces the computational complexity and improves the operating efficiency of the encoder. The specific form of the appendix table is as Figure 4 shown in the data instance.

[0078] However, this representation method will increase the number of clock cycles consumed during parallel hardware implementation. To optimize this problem, it is necessary to preprocess the parity-check matrix to reduce the number of clock cycles required during encoding. Through effective preprocessing, the complexity of hardware implementation can be significantly reduced, and the overall efficiency of the system can be improved. The specific preprocessing method is as Figure 3 shown. This figure shows how to convert the original parity-check matrix into a structure more suitable for parallel calculation. Finally, the goal of preprocessing is to construct the set , which will be used to store the indexes of non-zero elements in the transformed cyclic shift matrix and the corresponding number of input information bit groups, and specifically needs to be processed into the form of Table 3. In this way, non-zero elements can be effectively managed and accessed, thus accelerating the subsequent calculation process.

[0079] To achieve hardware storage, the data in the table needs to be stored in the ROM. Therefore, the data must be processed, converted into hexadecimal format, and the corresponding.coe file is generated. First, represents the number of groups of input information bits, and its maximum value will not exceed the number of groups t. Therefore, it can be stored using 7-bit binary numbers. Similarly, represents the number of circular right shifts required for the information bit row vector of the current group. Since the data is grouped by 360 bits, the maximum value of will not exceed 360 either, which means it can be stored using 9-bit binary numbers. It should be noted that the number of binary tuples in each row may be different, but the maximum number of groups does not exceed 6 groups. Therefore, the total number of bits required to store one row of data is ( bits. For data with less than 6 groups, it will be filled with all 0s at the end to ensure that the length of each row of data is consistent. In addition, since the data in the ROM is stored in the order of row addresses, the y value does not require additional storage space and can be directly represented by the row address. Finally, all the processed data will be converted into hexadecimal format, and the result of the generated.coe file is as shown in Figure 5 the data instance.

[0080] b) Implementation of the parity-check matrix preprocessing system The implementation of the parity-check matrix preprocessing system mainly depends on two core functions: GetHMatrixMap and GetHMatrixMapTable. The design of these two functions aims to efficiently process and generate parity-check matrix mappings with different code rates. The GetHMatrixMap function reads the parity-check matrix data from a specified file, constructs the corresponding file name according to the input code rate and type (e.g., ATSC or DVB-S2), and then reads the file content. The GetHMatrixMapTable function uses the GetHMatrixMap function to obtain the parity-check matrix mapping, thereby re-arranging the matrix and obtaining the valid information index to achieve parity-check matrix preprocessing. The specific algorithm flow of the parity-check matrix preprocessing is as shown in Figure 6 shown. Figure 6Describes the flowchart of the parity-check matrix preprocessing algorithm, which explains how to generate and process the parity-check matrix mapping (matrix_map) based on the input code rate (code_rate) and protocol type (type). Among them, matrix_map represents the parity-check matrix construction information read from the protocol, n represents the total codeword length, k is the information bit length, r is the parity bit length, t is the number of information bit groups, q is the number of parity bit groups, matrix_group and rearranged_matrix_group represent the constructed parity-check matrix and the rearranged parity-check matrix respectively, group, i, and idx are used to count the number of different groups, and finally, the subscripts i and j of the matrix are stored in matrix_map_table, that is, the data set after the parity-check matrix preprocessing 。

[0081] 2) Hardware architecture design The overall design of the hardware architecture is as Figure 7 shown, including the information bit matrix multi-port RAM, the parity bit matrix multi-port ROM, the parity-check matrix calculation module, the sum matrix calculation module, the sum matrix RAM, the parity bit matrix initialization module, the parity bit matrix calculation module, the parity bit matrix RAM, and the data flow control module. First, the information bits are input through data_in and stored in the multi-port read-write RAM. Then, the parity bit matrix calculation module reads multiple parity-check matrix data from the parity bit matrix multi-port read ROM at the same time for parsing to obtain the corresponding values. Then, the sum matrix calculation module reads the information bit vectors from the information bit matrix RAM for circular right shift and accumulation operations to obtain the sum matrix and stores it in the sum matrix RAM. At the same time, the row vectors obtained each time are accumulated to calculate the parity bit matrix initialization vector. Then, the parity bit matrix calculation module reads the sum matrix values at the corresponding addresses to calculate the parity bits and rearranges them and stores them in the parity bit matrix RAM. Finally, the data flow control module controls the output of the information bits and the parity bits through the enable signal en to obtain data_out.

[0082] Figure 7 Details the hardware architecture design of a low-density parity-check (LDPC) encoder. The design includes multiple key modules, each of which performs a specific function to implement the LDPC encoding process. The hardware architecture includes the following modules: (1) Input - RAM (information bit matrix multi-port RAM) - responsible for storing the input information bits (data_in), supporting multi-port access, allowing multiple data to be read from the RAM simultaneously for parallel processing.

[0083] (2) Hmatrix- ROM (Check Matrix Multi-Port ROM) - Stores the check matrix. The multi-port ROM allows multiple check matrix data to be read simultaneously for subsequent calculations.

[0084] (3) SMatrix Calc Core (Sum Matrix Calculation Module) - Reads information bits from the Input RAM, performs circular right shift and accumulation operations, and stores the resulting sum matrix in the SMatrix RAM.

[0085] (4) SMatrix RAM (Sum Matrix RAM) - Stores the intermediate results generated by the sum matrix calculation module.

[0086] (5) HMatrix Calc Core (Check Matrix Calculation Module) - Reads check matrix data from the HMatrix ROM, parses the corresponding generator polynomial (ϕ) and check polynomial (ψ) values for check matrix calculation.

[0087] (6) PMatrix Init Core (Parity Bit Matrix Initialization Module) - Calculates the parity bit matrix initialization vector based on the calculation results of the sum matrix calculation module each time.

[0088] (7) PMatrix Calc Core (Parity Bit Matrix Calculation Module) - Reads the values from the sum matrix RAM, calculates the parity bits, and reorders them. The calculation results are stored in the PMatrix RAM.

[0089] (8) PMatrix RAM (Parity Bit Matrix RAM) - Stores the parity bits generated by the parity bit matrix calculation module.

[0090] (9) Encoder Control Core (Data Flow Control Module) - Responsible for controlling the data flow of the entire encoder, controlling the output of information bits and parity bits through the enable signal (en), and generating the final encoded output (data_out) and valid output signal (valid_out).

[0091] The working processes of these modules include: The input information bits are first stored in the Input-RAM; the parity-check matrix calculation module reads the parity-check matrix data from the HMatrix ROM and parses it; the sum-accumulation matrix calculation module reads the information bits from the Input RAM, performs circular right shift and sum-accumulation operations, and stores the results in the SMatrix RAM. Each calculation result of the sum-accumulation matrix is used to calculate the initialization vector of the parity-check bit matrix; the parity-check bit matrix calculation module reads the sum-accumulation matrix values in the SMatrix RAM, calculates the parity-check bits and reorders them, and stores the results in the PMatrix RAM; finally, the data stream control module controls the output of the information bits and the parity-check bits through the control signal en to obtain the encoded data (data_out).

[0092] Furthermore, multi-port RAM / ROM structures are adopted in the hardware architecture. Although the IP cores provided by Vivado by default only support dual-port RAM, in the method proposed in this disclosure, since the selected parallelism m supports multiple options, for example, it is necessary to read data from 4 addresses simultaneously, while dual-port RAM cannot achieve this function. If multiple single-port RAMs are used to store the same data to achieve simultaneous reading, multiple data copies will be stored, thus wasting storage space. Therefore, in order to meet the design requirements of high-performance systems, custom multi-port RAM and ROM are the keys to achieving efficient data processing capabilities. The design of multi-port RAM first requires clarifying several key parameters, including the address width ADDR_WIDTH, the data width DATA_WIDTH, the storage depth DEPTH, and the number of read ports NUM_READ_PORTS. During the implementation process, Verilog language is used to construct a custom multi-port read single-port write RAM and a multi-port read ROM. By using an array of storage cells, different address lines can be used to access the corresponding data simultaneously, which greatly improves the efficiency of data access.

[0093] In addition, for the internal implementation of the memory cell array, the statement (* ram_style = "block" *) is used to guide the synthesis tool to synthesize this RAM into a Block RAM (BRAM) as much as possible. Since BRAM is a dedicated storage resource inside the FPGA, it has the advantages of high bandwidth and low latency, and is suitable for application scenarios that require fast access. Compared with the Dynamic Random Access Memory (DRAM), BRAM has a faster read / write speed and a lower access latency because it is located in the logic hierarchy of the FPGA and can directly interact with the logic units. DRAM is usually used for large-capacity storage, and its structure requires row and column selection during access, resulting in higher latency. Although DRAM has a higher storage density and is suitable for large-scale data storage, in high-performance applications, the low-latency characteristic of BRAM makes it a better choice. Therefore, choosing BRAM as the implementation method of the multi-port RAM can significantly improve the data processing speed and meet the requirements of high-performance systems.

[0094] In the design of the write operation, the write enable signal is set to control the data to be written to the specified address at a specific time. However, the setting of read / write conflict is not particularly considered in this design because this algorithm will not have the situation of reading and writing the same address simultaneously during the implementation process. But in practical applications, additional control logic usually needs to be introduced to ensure the integrity and consistency of the data. In terms of the construction of the ROM, a structure similar to that of the RAM is adopted. However, since the ROM is usually used to store fixed data, a coe file is used to define its content during the design. This method ensures that the system can quickly access the required data after power-on without additional write operations. At the same time, by controlling the clock signal, multiple read ports are designed so that the system can parallelly read data from different addresses in the same clock cycle to facilitate the simultaneous parsing of multiple preprocessed check matrix data.

[0095] The algorithm flows of other sub-modules such as the check bit matrix calculation module and the sum matrix calculation module are respectively as Figure 8 and Figure 9As shown in the figure. Among them, state is used to indicate different states, matrix_rd_re is the read enable, matrix_rd_addr is the read address, matrix_rd_addr_mem is used to store the read data, SM_PM_DEPTH is used to represent the depth of the memory for accumulating and summing the matrix RAM, calc_idx is used to count which group of data is being processed, m_mem_[0 / 1] and alpha_mem_[0 / 1] respectively represent the data in the set, and i is used to count which set it is, input_ram_rd_data[0 / 1] represents the data read from the information bit matrix, shifted_data_mem[0 / 1] is used to store the data after circular right shift, and sm_ram_wr_data[0 / 1] represents the data written to the accumulating and summing matrix.

[0096] In the above processing procedures, the use of multi-port RAM and ROM is the key to achieving efficient data processing. Through the custom multi-port RAM and ROM, data at multiple addresses can be read in parallel within one clock cycle, improving the efficiency of data access for parity-check matrix calculation in the LDPC encoder. Among them, the multi-port RAM accesses multiple data simultaneously by using different address lines, improving the parallelism of data reading, and the multi-port ROM is used to store fixed data, such as the parity-check matrix, and enables fast data access through multiple read ports.

[0097] In order to verify the implemented parallel LDPC encoder, the present disclosure has built a test platform. The present disclosure is a hardware test platform for the DVB-S2 transmitter system based on FPGA, and its specific design is as follows.

[0098] 1) Overall design of the test platform architecture.

[0099] The overall hardware implementation design of the DVB-S2 transmitter system is carried out and an interface driver module is built to verify the above-designed LDPC parallel encoder and DVB-S2 transmitter system. The specific architecture is as Figure 10 shown. This architecture includes, (1) Host computer - an external computer or control device that communicates with the FPGA system. The host computer sends UDP packets to the FPGA for testing and verifying the functions of the LDPC encoder and DVB-S2 transmitter system.

[0100] (2) UDP communication interface module - used to process the UDP packets sent by the host computer. It receives data from the host computer and distributes the data to other modules through the interface inside the FPGA. At the same time, it is also responsible for encapsulating the processed data into UDP packets and sending them back to the host computer.

[0101] (3) DDR3 Memory Module - The main storage unit in the system, used to temporarily store data and programs, and plays a role in buffering and data management during data flow through the system.

[0102] (4) DDR Input FIFO - Input First-In-First-Out (FIFO) buffer, used to temporarily store data transmitted from the UDP communication interface module to alleviate the problem of mismatched data transfer rates between different modules.

[0103] (5) UDP Packet FIFO - Used to store complete UDP data packets.

[0104] (6) Baseband Processing Module DVB-S2 Transmitter - This module converts data into the DVB-S2 standard format and prepares the data for transmission; it reads data from the DDR3 memory module, processes it, and then sends the data to the DVBS2 FIFO.

[0105] (7) DDR Output FIFO - The output FIFO is used to temporarily store the data processed by the baseband processing module and control the data flow to the DVB-S2 transmitter.

[0106] (8) DVBS2 FIFO - Is the transmission buffer, used to store data to be transmitted through the physical medium (such as radio frequency).

[0107] Figure 10 The architecture design shown is used to verify the performance and functionality of the parallel LDPC encoder, test the data processing ability of the encoder and the compatibility of the transmitter in a real hardware environment, and ensure that the encoder can work properly in actual applications. This design utilizes the reconfigurable characteristics of the FPGA to achieve a flexible data processing flow and is controlled and monitored by an external host computer.

[0108] Further, the baseband processing module design, that is, the hardware implementation design of the DVB-S2 transmitter system, has the overall architecture as Figure 11 shown, and this architecture covers the baseband processing module design for signal processing in digital video broadcasting. Specifically, BB Frame Module (Baseband Frame Module), and it further includes, CRC Encoder - Cyclic Redundancy Check Encoder, used to detect errors in data transmission; BB Scrambler - Baseband Scrambler, used to reduce the DC component and long strings of 0 / 1 in the data for easy transmission; Baseband Signaling - Baseband signal processing, conditioning the signal to adapt to the baseband transmission requirements; FEC Module (Forward Error Correction Coding Module), and it further includes, BCH Encoder - A block code encoder that introduces additional redundancy into the data for error detection and correction at the receiving end; LDPC Encoder - A low-density parity-check encoder; The bit interleaving and mapping module, which further includes, Bit Mapper - A bit mapper that maps the encoded bit stream onto a constellation diagram for modulation; Bit Interleaver - A bit interleaver that is used to scatter the data to combat burst errors and enhance the robustness of the signal; PL Frame module (Physical Layer Frame module), which further includes, FIR Filter - A finite impulse response filter that is used for pulse shaping of the signal to reduce inter-symbol interference (ISI); PL Scrambler - A physical layer scrambler that further randomizes the data to increase the reliability of transmission; PL signaling & Pilot insertion - Physical layer signal processing and pilot insertion for channel estimation and synchronization.

[0109] The output IQ_out is the processed in-phase / quadrature signal.

[0110] This hardware architecture design implements a complete DVB-S2 transmitter system, from error detection coding (CRC), scrambling (BB Scrambler), forward error correction coding (FEC module), to mapping and interleaving before signal modulation, and then to physical layer signal processing (PL Frame module). Each module works together to ensure the integrity and reliability of the data during transmission, while meeting the technical requirements of the DVB-S2 standard. This hardware implementation method is applicable to digital video broadcast scenarios that require high reliability and high-quality transmission. The system includes four modules, namely the BB (Baseband) Frame module, the FEC (Forward Error Correction Coding) module, the bit interleaving and mapping module, and the PL (Physical layer) Frame module. The main function is to add a baseband frame header to the information bit to be sent, then perform BCH and LDPC encoding, then perform bit interleaving and map the bits into a constellation diagram corresponding to the selected modulation method to obtain the IQ value, and finally add a physical layer frame header and pass through a root raised cosine filter to obtain the final IQ value for output. Subsequently, it can be connected to a DAC to convert it into an analog signal and then be transmitted through radio frequency. Next, the design of each sub-module will be introduced: a) The main function of the BB Frame module is to encapsulate the input bitstream into a baseband frame, including adding a baseband frame header and performing baseband signal processing. The addition of the baseband frame header is to enable the correct identification and parsing of data frames at the receiving end, and the baseband signal processing includes operations such as CRC encoding and baseband scrambling. Among them, the parameter DFL represents the effective data length, which is determined by looking up the table specified in the protocol document after selecting the code rate and frame type. The parameter UPL represents the user data packet length. Here, MPEG - format data packets are used, so the fixed value is 1504 bits. TSorGS represents whether to use the transport stream or the generic stream. Here, the transport stream is selected, and the value is 0. SYNC represents the synchronization byte between data packets, and the fixed value is 47H. MODCOD refers to the selected modulation mode and code rate scheme, and the value range in the protocol document is [0, 31]. For example, when the value is 6, it represents the selection of QPSK and a 2 / 3 code rate. The specific hardware implementation of this module: Data is input from pkt_bits_in, and pkt_valid_in indicates whether the data is valid. After counting to a data packet, CRC encoding is performed on it. This encoding is implemented using a linear feedback shift register (LFSR), and its polynomial is ( ) and the CRC check code is generated through shift and exclusive - OR operations. The specific hardware structure is as shown in Figure 12 . Among them, EXOR represents the exclusive - OR operation, bit_in represents the input information bit, and crc_out represents the CRC - encoded output data.

[0111] b) The main function of the FEC module is to perform forward error correction encoding, including BCH encoding as the inner code and LDPC encoding as the outer code. BCH encoding is used to detect and correct single - bit errors, while LDPC encoding has a stronger error - correction ability. After cascading with BCH encoding, it can avoid error floors. BCH encoding can correct 8, 10, and 12 - bit errors, which are calculated by multiplying the first 8, 10, or 12 polynomials specified in the document to obtain the generating polynomial. When this module is implemented in hardware, the BCH encoder generates the check code through shift and exclusive - OR operations using a linear feedback shift register (LFSR). Specifically, according to the code rate code_rate_idx, determine and the length, then count the input bb_frame_out data, and use the first bits of data to find the BCH check bits using the LFSR and the calculated generating polynomial. Then, perform LDPC encoding on the bits of data with the added BCH check bits. The implementation of the LDPC encoder uses the parallel LDPC encoder described in detail in Chapters 3 and 4, which will not be elaborated here. Finally, the data after LDPC encoding is output through the interface fec_out, and the fec_valid_out signal indicates that the output data is valid.

[0112] c) The main functions of the bit interleaving and mapping module are to perform bit interleaving and bit mapping. Bit interleaving is used to scramble the bit order to improve the anti-interference ability of the signal. Bit mapping maps the bits onto the constellation diagram corresponding to the modulation method to generate IQ values. There are four modulation methods supported in the DVB-S2 protocol, namely QPSK, 8PSK, 16APSK, and 32APSK. According to the selected modulation method, the IQ values obtained from bit interleaving and constellation mapping are different, and the specific parameter settings are specified in the protocol document. When implemented in hardware, the bit interleaver is realized through a memory and an address generator. Specifically, the interleaving depth and interleaving pattern are determined according to the modulation method index mod_idx, the input fec_out data is counted and stored, and then the bit order is rearranged through the address generator to achieve bit interleaving. Then, the bit mapper uses a pre-stored look-up table (LUT) for one-to-one mapping to map the interleaved bits onto the constellation diagram to generate IQ values. The look-up table (LUT) stores the mapping relationship between bits and constellation points, and through the look-up table, bits can be quickly mapped to the corresponding constellation points. For example, when using QPSK, 2 bits correspond to a group of IQ values, and when using 16APSK, 4 bits correspond to a group of IQ values. When implemented in hardware, an I value or a Q value is stored with a length of 11 bits including 1 sign bit. Finally, the I-channel and Q-channel data are output through the interfaces i_out and q_out, and the iq_valid_out signal indicates that the output data is valid.

[0113] d) The main function of the PL Frame module is to generate the physical layer frame, including adding the PL frame header, performing PL scrambling, using BPSK modulation on the frame header, and generating dummy frames. The addition of the PL frame header is to correctly identify and parse the data frame at the receiving end, and the PL scrambling is to improve the anti-interference ability of data transmission. Generating dummy frames is used to fill the idle period to ensure continuous occupation of the channel. When implementing this module, a state machine is used to control the switching of each state. First, when the data of the previous module has not been calculated yet, the state will be switched to DUMMY. In this state, a dummy frame with a fixed length of symbol lengths is generated to fill the idle period and ensure continuous occupation of the channel, where the I value and Q value of each symbol are both . When the data is ready, the state is switched to PL_FRAME_GENERATE. At this time, a PL header with a length of 90 symbols is generated for each frame. The PL header contains a constant sequence with a length of 26 symbols (18D2E82H) to detect the start position of the frame header at the receiving end, and the remaining length is encoded into 64 symbols by the circuit specified in the protocol document for MODCOD. Then, if the PL header is used with the sequence If it is indicated, the sequence will be BPSK modulated. The specific formula is as follows, where :

[0114] After calculating the frame header, the state will be switched to PILOT_INSERTION. In this state, the length of the symbols in the current frame will be counted. After each symbol, a pilot with a fixed length of 36 and IQ values all being will be inserted. The insertion of the pilot facilitates synchronization at the receiving end. After completing the pilot insertion, the state will be switched to PL_SCRAMBLE. In this state, the circuit diagram specified in the protocol document is used to scramble the entire frame sequence at the symbol level. After the scrambling process, a PL frame will be obtained. However, to send the frame data, it still needs to be filtered through a root raised cosine filter. At this time, the state will also be switched to FILTERING. This filter is implemented using a FIR filter in hardware. First, the filter coefficients need to be generated through matlab's fdatool according to the roll-off factors of 0.35, 0.25, and 0.20 of the root raised cosine filter specified in the protocol. Then, the coefficient values are pre-stored in registers. Finally, the designed pipeline is used to perform convolution operations to calculate the filtering result and output it through the pl_frame_i_out and pl_frame_q_out interfaces. The pl_frame_valid_out signal indicates the validity of the output data.

[0115] Furthermore, the three core modules of the interface driver module design: the DDR3 storage interface module, the UDP communication interface module, and the data buffer FIFO module. The interface driver design is a key part of the FPGA implementation of the DVB-S2 encoder, responsible for high-speed storage, transmission, and buffering of data. The DDR3 storage interface module provides a large-capacity and high-bandwidth data storage ability. The UDP communication interface module realizes efficient data interaction with external devices. The data buffer FIFO module solves the clock domain and rate matching problems between different modules through a multi-level caching mechanism. Through the collaborative work of these three modules, the system can meet the strict requirements of the DVB-S2 standard for data throughput and real-time performance.

[0116] a) The DDR3 memory interface module is responsible for implementing high-speed data interaction between the FPGA and the external DDR3 memory. The design of the DDR3 memory interface module is mainly based on the requirements of the DVB-S2 encoder for high bandwidth and large-capacity storage. The DDR3 memory provides a theoretical bandwidth of up to 12.8 Gbps, which can meet the needs of the encoder when processing high data rates. In addition, the large-capacity storage capacity of DDR3 (supporting 2GB of storage space) enables the system to cache complete satellite frame data (such as a frame length of 64,800 bits), thus avoiding data loss or overflow. Through burst transfer optimization, the DDR3 interface can continuously transfer multiple data blocks in one operation, reducing the overhead of address switching and further improving the access efficiency. This design not only increases the system throughput but also reduces power consumption, enabling the FPGA to efficiently process complex encoding tasks. The implementation logic of the DDR3 memory interface module mainly revolves around the state machine. The state machine triggers read and write operations by detecting the FIFO threshold (such as ddr_input_fifo_rd_count > FDMA_BURST_LEN). In the write operation phase, the module requests to write through the fdma_wreq signal, with a fixed burst length of 512 (FDMA_BURST_LEN). After the data is aligned by 128 bits, it is transferred to the DDR3 through the fdma_wdata. The DDR3 address is specified by DDR3_0_addr. The column address strobe signal DDR3_0_cas_n and the row address strobe signal DDR3_0_ras_n are pulled low in sequence. After the write enable signal DDR3_0_we_n becomes valid, the data is written to the specified address through the DDR3 data bus DDR3_0_dq. After the write operation is completed, the fdma_wvalid signal is pulled high to confirm that the data has been successfully written to the DDR3. In the read operation phase, it is triggered by the fdma_rreq signal. The DDR3 address is specified by DDR3_0_addr. The column address and row address strobe signals are pulled low in sequence. The read data is returned to fdma_rdata through DDR3_0_dq, and the fdma_rvalid signal indicates that the data is valid. The arbitration state (ARBIT) switches the dual buffers (MEM1 / MEM2) through the fifo_switch signal, with address offsets of ADDR_MEM1_OFFSET (1000) and ADDR_MEM2_OFFSET (0x100000) respectively, to avoid read-write conflicts. This design ensures efficient data transmission and storage, and further improves the parallel processing ability of the system through the dual-buffer strategy.

[0117] b) The UDP communication interface module is responsible for implementing efficient data transmission between the FPGA and external devices. The design of the UDP communication interface module is to meet the requirements of real-time and efficiency in the data transmission process of the DVB-S2 encoder. In the DVB-S2 encoder, data needs to be transmitted to external devices quickly and accurately to ensure the real-time nature of the encoding process and the integrity of the data. UDP (User Datagram Protocol), as a connectionless transport layer protocol, has the characteristics of low latency and high throughput, and is very suitable for scenarios that require fast data transmission. Compared with TCP (Transmission Control Protocol), UDP does not need to establish a connection, reducing the overhead in the data transmission process and being able to transmit data from the sender to the receiver more quickly. This feature gives UDP obvious advantages in real-time data transmission. Especially in a system like the DVB-S2 encoder that has high requirements for data transmission real-time, the application of UDP can significantly improve the performance of the system. In addition, the simplicity and flexibility of the UDP protocol also bring convenience to the design and implementation of the system, reducing the complexity and power consumption of the system and improving the overall performance and reliability of the system. Therefore, choosing UDP as the communication protocol in the DVB-S2 encoder can better meet the requirements of the system for real-time and efficiency. In the specific hardware implementation, the UDP communication interface module controls the data reception and transmission process through a state machine. When the module detects that the data input request signal (udp_data_in_req) is pulled high, it indicates that there is data input, and the module starts to read data from the udp_data_in port. The read data will be temporarily stored and output to the target device through the udp_data_out port. During the data output process, the module indicates the validity of the output data through the udp_valid_out signal to ensure that the receiving end can correctly process the data. At the same time, the module requests the output data through the udp_data_out_req signal to ensure that the data can be sent to the target device in a timely manner. During the data transmission process, the module performs synchronization operations according to the system clocks (clk_15_625 and clk_200) to ensure the stability and accuracy of data transmission. The clk_15_625 clock signal is used to synchronize low-speed operations, while the clk_200 clock signal is used to synchronize high-speed operations. This multi-clock design enables the module to flexibly handle data transmission requirements at different rates, improving the adaptability and reliability of the system. In addition, the module also provides a core reset signal (core_reset) to reset the module state when the system starts or an error occurs, ensuring that the module can work from the initial state. Through the above design, the UDP communication interface module can efficiently support the real-time data transmission requirements of the DVB-S2 encoder, ensuring that data is transmitted to the target device quickly and accurately.

[0118] c) The design of the data buffer FIFO module is to meet the requirements of real-time and efficiency in the data transmission process of the DVB-S2 encoder. In the DVB-S2 encoder, the data transfer rates and clock domains between different modules may not be consistent. Therefore, it is necessary to buffer and synchronize through the FIFO. When implemented in hardware, each FIFO is also controlled by a state machine for data read and write operations. The state machine dynamically controls the read and write enable signals of each FIFO according to the status of the FIFO (such as full flag and empty flag) and the requirements of data transmission. It also monitors the counters of each FIFO (such as wr_count and rd_count) to ensure the correctness and timeliness of data transmission. The specific structure is as Figure 13 shown. Figure 13 Figure Figure 13 shows a data buffer FIFO (First In First Out) module, which is used in the DVB-S2 encoder. Specifically, it includes (1) PC (Personal Computer) - Sends UDP data packets to the DVB-S2 encoder; (2) DDR Input FIFO - Used to temporarily store the UDP data packets received from the PC; (3) DDR (Double Data Rate Synchronous Dynamic Random Access Memory) - Serves as the main data storage unit for data transmission between different modules, including storing the data received from the DDR Input FIFO and providing it to the DDROutput FIFO; (4) State (State Machine) - Controls the data read and write operations of the FIFO module; (5) UDP Packet FIFO - Stores the complete UDP data packets; (6) DVBS2 FIFO - Buffers the processed data and sends it to the DVBS2 transmission module; (7) DVBS2 TX (Transmission Module) - Transmits the data through radio frequency signals.

[0119] Figure 13 The data flow path through each module in

[0119] is: receive UDP data packets from the PC and store them in the DDR InputFIFO; read the data from the DDR Input FIFO and store it in the DDR memory; according to the control of the state machine, transfer the data between the DDR memory and each FIFO; read the data from the DDR Output FIFO, pass it through the DVBS2 FIFO, and finally transmit it to the DVBS2TX transmission module for transmission.

[0120] The state machine plays a core control role in this process, ensuring that the data transfer between each FIFO is synchronous and orderly, preventing data loss or overflow. By monitoring the counters of the FIFOs, the state machine can dynamically adjust the data stream to adapt to different transfer rates and processing requirements, thus meeting the requirements of the DVB-S2 encoder for real-time and efficient data transfer.

[0121] Furthermore, the present disclosure tests the implemented LDPC parallel encoder and the DVB-S2 transmitter system. The present disclosure builds a hardware loopback test platform for implementation verification and performance testing. The specific situation is as follows: 1) Introduction to the test platform To comprehensively evaluate the performance of the system, the present disclosure adopts a joint test scheme based on FPGA-host computer. The interface driver module designed in Chapter 4 is implemented on the FPGA side, and this module is responsible for real-time transmitting the encoded data to the host computer through UDP packets. During the test process, we focus on the following performance indicators: (1) The throughput of the encoder, which is evaluated by counting the amount of data transferred per unit time; (2) The bit error rate (BER) performance, which is evaluated by comparing the LDPC decoding situations using different parity check matrices; (3) The resource utilization rate, which is evaluated by obtaining the resource occupancy of each module of the encoder through the FPGA development tool.

[0122] Experimental process: When designing the parallel LDPC encoder, first, the python code is used to preprocess its parity check matrix, and then the verilog language is used to write the code for hardware implementation in the vivado software. After that, the bitstream file is written into the development board of the model Zynq-7000 XC7Z035-2FFG900I. Finally, the host computer sends data to the FPGA through UDP packets by writing python code, and the FPGA processes the data and then sends it back to the host computer through UDP packets. After receiving the data, the host computer stores it in a file. Then, the toolbox functions dvbs2WaveformGenerator and ldpcEncode of Matlab are used to generate the data of the DVB-S2 protocol and the data after LDPC encoding for comparison.

[0123] 2) System implementation verification, specifically including, a) System implementation verification of the parallel LDPC encoder When implementing the test, first generate random information bits in the host computer using Python, and then send the information bits to the LDPC encoder in the FPGA via UDP packets through the interface driver module. Next, transmit the data after encoding to the host computer via UDP packets and store it in the file fpga_ldpc_data.txt. Then, use the function ldpcEncoderConfig in the Matlab toolbox to configure the LDPC encoding. The main configuration is to use the same code rate and the parity-check matrix of the same protocol as the LDPC encoder in the FPGA. After the configuration is completed, use the function ldpcEncode to encode the same input data and store the result in matlab_ldpc_data.txt. Finally, compare the data in the two txt files to check whether the bit data in each row is the same, and randomly sample 50 values from the data to draw a heat map for easy observation. Specifically, as shown in Figure 14 shown. The results of bit-by-bit comparison are all the same, and it can also be seen from the figure that when the code rates are 2 / 5, 3 / 5, 2 / 3, and 4 / 5, the output results of the LDPC encoder implemented by the FPGA are consistent with the output results of the Matlab toolbox functions, which proves that the encoder can correctly encode at low, medium, and high code rates.

[0124] b) Verification of the DVB-S2 transmitter system implementation This test process is basically the same as that in the previous section. It also generates random bit information through Python code and then sends UDP packets through the interface driver module. After being received and processed by the FPGA, it is sent back to the host computer via UDP packets. The difference is that the data is stored in fpga_dvbs2_data.txt. Then, use the function dvbs2WaveformGenerator in Matlab to configure the same parameters as the transmitter system in the FPGA and run to generate DVB-S2 signal data and store it in the file matlab_dvbs2_data.txt. Finally, compare the data in the two files using RelativeMSE (Relative Mean Squared Error). MSE is a commonly used metric to measure the similarity between two signals. The smaller its value, the more similar the two signals are. In a digital communication system, MSE is often used to evaluate the error between the transmitted signal and the received signal. Its calculation formula is:

[0125] This formula divides MSE by the power of the reference signal (i.e., ), and take the logarithm to obtain a value in decibels (dB). The advantage of this is that it can more intuitively represent the magnitude of the error, especially in the case of large signal power variations. In addition, a constellation diagram is plotted for the output signal of the FPGA to facilitate the observation of the output results under various modulation methods. The specific test situation is as Figure 15 shown. After conducting experimental comparisons for four modulation methods and different code rates, it can be seen that the constellation diagram plotted from the output result data of the transmitter system in the FPGA conforms to the current modulation method. From Table 1, it can be seen that the MSE of the real and imaginary parts under the four MODCODs is basically around -79 dB. This data indicates that the error power is approximately a certain multiple of the reference signal power, that is, the error power is very small relative to the reference signal power. Therefore, it can be considered that the output results of the FPGA transmitter system under the four MODCODs are basically the same as the output results of the Matlab function. Therefore, the hardware implementation of the DVB-S2 transmitter system is correct.

[0126] Table 1 MSE results under different modulation methods and code rates

[0127] 3) Performance testing a) Throughput performance testing Table 2 Comparison between the latest existing research and the proposed solution of the present invention

[0128] Latest existing research literature: NANNIPIERI P, BARTOLACCI G, BERTOLUCCI M, et al. Design and Implementation of a Configurable Fully Compliant DVB-S2 LDPC Encoder for High Data-Rate Downlink Payload[J]. IEEE Access, 2024, 12: 39204-39220.

[0129] As shown in Table 2 for the comparison of throughput performance, the encoder designed by the method proposed in this paper uses Algorithm 4 for hardware experiments, and the parallelism parameter m = 2 is set, that is, 720 parallelism, which is twice that of the latest research. In addition, when designing the data flow control module in Chapter 3, because the information bits of bits are output first, and then the re-arranged bit parity bits are output after the information bits are output, the total throughput can be calculated by formula (27):

[0130] Therefore, when the method proposed in this disclosure is used in an FPGA with the model number Zynq-7000 XC7Z035-2FFG900I and a clock frequency of 100 MHz, the throughput can reach 23.4 Gb / s, which is approximately twice the throughput of 12 Gb / s in the latest research.

[0131] b) Bit error rate performance test The bit error rate performance of the LDPC encoder after replacing the parity check matrix using the ATSC protocol at code rates of 2 / 5, 3 / 5, 2 / 3, and 4 / 5 was tested and compared with that without replacement. The results are as Figure 16 shown. Here, BER represents the bit error rate and SNR represents the signal-to-noise ratio. It can be seen that when using the QPSK modulation method, the bit error rate performance of replacing the parity check matrix with the ATSC protocol at code rates of 2 / 5 and 4 / 5 is 0.1 dB better than that using the parity check matrix in the original DVB-S2 protocol. When the code rate is 3 / 5 and 2 / 3, the bit error rate performance of replacing the parity check matrix with the ATSC protocol is 0.2 dB better than that using the parity check matrix in the original DVB-S2 protocol. Therefore, the bit error rate performance can be improved by 0.1 - 0.2 dB after replacing with the ATSC protocol parity check matrix.

[0132] c) Resource utilization test Table 3 Comparison of hardware resource utilization between the latest research and the proposed solution of the present invention

[0133] The comparison of resource utilization is shown in Table 3. It can be seen from the table that the encoder designed by the method proposed in this paper is superior to the latest research method in terms of the resource utilization of BRAM (block RAM) and LUT (look-up table). Specifically, the encoder designed by the method in this paper uses 34 BRAM tiles, while the latest research method uses 47 BRAM tiles; in terms of LUT resources, the encoder designed by the method in this paper uses 19,878 LUT tiles, while the latest research method uses 40,354 LUT tiles. This shows that the encoder designed by the method in this paper is more efficient in using BRAM and LUT resources. However, in terms of the resource utilization of FFs (flip-flops), the encoder designed by the method in this paper is slightly higher than the latest research method. Specifically, the encoder designed by the method in this paper uses 42,185 FF tiles, while the latest research method uses 40,526 FF tiles. Although it is slightly higher in terms of FF resource utilization, generally speaking, the encoder designed by the method in this paper and the encoder designed by the latest research method have little difference in FPGA resource utilization.

[0134] In summary, the technical effects achieved by this disclosure include: (1) An optimized parallelization algorithm based on a cyclic shift matrix is proposed to improve throughput. By preprocessing the parity-check matrix and constructing an accumulation sum matrix parallel computing architecture, the encoding throughput is increased to 23.4 Gb / s at a 100 MHz clock, which is approximately twice that of the latest parallelization method at the same clock frequency. This algorithm improves the processing speed of the encoder while maintaining low hardware resource utilization.

[0135] (2) A multi-port RAM storage architecture and an odd-even address block mechanism are designed to achieve low resource utilization. By constructing a multi-port RAM to avoid redundant storage and using an odd-even address block mechanism to solve parallel read-write conflicts, a compact hardware design with low resource utilization is implemented on the Xilinx Zynq-7035 development board. This design realizes efficient storage and data processing under limited hardware resources.

[0136] (3) A compatibility mechanism for the parity-check matrices of ATSC 3.0 and DVB-S2 is established to improve the bit error rate performance. By replacing the parity-check matrix, the bit error rate is reduced by 0.1 - 0.2 dB at different code rates, expanding the protocol adaptability of the encoder and improving the bit error rate performance of the traditional DVB-S2 system without the need to reconstruct a new parity-check matrix.

[0137] (4) A DVB-S2 transmitter system based on FPGA is designed and implemented. This system realizes the BB Frame module, forward error correction coding module, bit interleaving and mapping module, and PL Frame module specified in the DVB-S2 protocol. In addition, by integrating the designed parallel LDPC encoder and supporting multiple modulation methods and code rates, this system meets the requirements of high data rate and large-capacity transmission.

[0138] (5) A hardware loopback test platform is built for implementation verification and performance testing. An FPGA-host computer test platform is built, and the interface driver module is implemented so that the host computer can send UDP packet data to the FPGA and then send the output data generated by the FPGA back to the host computer via UDP packets for system implementation verification and performance testing. The test content includes key indicators such as throughput, bit error rate performance, and resource utilization, proving the correctness and reliability of the disclosed system.

[0139] It should be noted that although the foregoing has described the spirit and principle of the present disclosure with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the disclosed specific embodiments, and the division of each aspect does not mean that the features in these aspects cannot be combined. This division is only for the convenience of expression. The present disclosure aims to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. An LDPC encoder based on the DVB-S2 protocol, characterized in that, The encoder includes, an information bit matrix multi-port input RAM for storing the input information bits; a parity-check matrix multi-port ROM for storing the parity-check matrix; a sum matrix RAM for storing the sum matrix generated by the sum matrix calculation module; a parity-check bit matrix RAM for storing the parity-check bit matrix calculated by the parity-check bit matrix calculation module; a parity-check matrix calculation module that reads the parity-check matrix from the parity-check matrix multi-port ROM, and the parity-check matrix is preprocessed and converted into a cyclic shift matrix; a sum matrix calculation module that reads the information bits from the information bit matrix multi-port input RAM, performs cyclic right shift and accumulation operations, and calculates the sum matrix; a parity-check bit matrix initialization module that calculates the parity-check bit matrix initialization vector according to the calculation result of the sum matrix calculation module; a parity-check bit matrix calculation module that reads the value in the sum matrix RAM, combines the parity-check bit matrix initialization vector, and calculates the parity-check bit matrix; a data flow control module for controlling the data flow of the LDPC encoder, controlling the output of the information bits and parity-check bits through an enable signal, and generating an LDPC encoded output.

2. The encoder according to claim 1, characterized in that, The storage unit of the information bit matrix multi-port input RAM adopts an odd-even address block mechanism, and stores the sum matrix and the parity-check bit matrix into the sum matrix RAM and the parity-check bit matrix RAM respectively according to the address parity.

3. The encoder according to claim 1, characterized in that, The parity-check matrix is a parity-check matrix based on the ATSC protocol.

4. The encoder according to claim 1, characterized in that, The working process of the encoder includes that the input information bits are first stored in the information bit matrix multi-port RAM; the parity-check matrix calculation module reads the parity-check matrix data from the parity-check bit matrix multi-port ROM and parses it; the sum matrix calculation module reads the information bits from the information bit matrix multi-port RAM, performs cyclic right shift and accumulation operations, and stores the result in the sum matrix RAM. Each calculation result of the sum matrix calculation module is used to calculate the parity-check bit matrix initialization vector; the parity-check bit matrix calculation module reads the sum matrix value in the sum matrix RAM, calculates the parity-check bits and reorders them, and stores the result in the parity-check bit matrix RAM; Finally, the data flow control module controls the output of the information bits and parity-check bits by controlling the enable signal to obtain the encoded data output.

5. A DVB-S2 transmitter, characterized in that, It includes a baseband frame module, a forward error correction coding module, a physical layer frame module, and a bit interleaving and mapping module, where the baseband frame module further includes: a CRC encoder, a baseband scrambler, and a baseband signal processing module; the forward error correction coding module further includes: a block code encoder and the LDPC encoder as described in claim 1; the physical layer frame module further includes: a finite impulse response filter, a physical layer scrambler, a physical layer signal processing and pilot insertion module; the bit interleaving and mapping module further includes: a bit mapper and a bit interleaver.

6. The transmitter according to claim 5, characterized in that, The working process of the transmitter includes firstly, the baseband frame module encapsulates the baseband frame and adds a baseband frame header; then, the forward error correction coding module performs BCH and LDPC coding; next, the bit interleaving and mapping module performs bit interleaving and maps the bits into a constellation diagram corresponding to the selected modulation method to obtain IQ values. Finally, the physical layer frame module generates physical layer frames, including adding a PL frame header, PL scrambling, header modulation, and pilot insertion, and passing through a root raised cosine filter to obtain the final IQ values and output them.

7. The transmitter according to claim 6, characterized in that, The obtained output bit stream is converted into an analog signal through digital-to-analog conversion for radio frequency transmission.

8. A DVB-S2 transmitter test platform, characterized in that, The platform includes a UDP communication interface module, a DDR3 storage module, a DVB-S2 transmitter as described in claim 5, and a data buffer FIFO module.

9. The platform according to claim 8, characterized in that, The data buffer FIFO module includes: DDR input FIFO, UDP packet FIFO, DDR output FIFO, and DVBS2 FIFO.

10. The platform according to claim 8, wherein The platform further includes a state machine for controlling the data read and write operations of the data buffer FIFO module.

Citation Information

Patent Citations

  • LDPC (Low Density Parity Check) encoder

    CN102684707A

  • LDPC (low density parity check) coder capable of generating matrix and check matrix jointly and coding method

    CN103023516A

  • Encoding method and encoder for RA-LDPC-CC in communication modulation system

    CN110324048A

  • Method for quickly generating LDPC (Low Density Parity Check) decoder

    CN118138058A

  • Low-density parity check code generation method and device, equipment and medium

    CN119814045A

Cited By

  • LDPC (Low Density Parity Check) code check matrix construction method, encoder, transceiving end and experimental system

    CN121217152A