Standard-oriented qc-ldpc encoder architecture, encoding method and chip

By adopting a standards-oriented QC-LDPC encoder architecture and utilizing switchable matrix operation modes and timing alignment, the problem of high hardware overhead and high latency in existing QC-LDPC encoders under diverse application scenarios is solved, achieving low latency, high throughput and low power consumption encoding performance, suitable for B5G and 6G communication systems.

CN122293099APending Publication Date: 2026-06-26PURPLE MOUNTAIN LAB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PURPLE MOUNTAIN LAB
Filing Date
2026-03-25
Publication Date
2026-06-26

Smart Images

  • Figure CN122293099A_ABST
    Figure CN122293099A_ABST
Patent Text Reader

Abstract

This application relates to the field of encoder technology and discloses a standard-oriented QC-LDPC encoder architecture, encoding method, and chip. In the encoder architecture, a first cyclic shift network module is configured to switch between the operation modes of matrix A and matrix C1, generating intermediate variables and first values ​​for extended parity bits, respectively. A timing buffer module is used to perform timing alignment buffering on the intermediate variables. A core transformation module receives the timing-aligned and buffered intermediate variables and solves for the core parity bit. A second cyclic shift network module adapts to the operation mode of matrix C2 and generates second values ​​for extended parity bits. The first and second values ​​for extended parity bits are XORed to obtain the extended parity bit, thus acquiring the encoding result. The technical solution provided by this application can provide an automated QC-LDPC encoder architecture with low hardware overhead, low latency, and high performance, while meeting the requirements of high throughput and extremely low power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of encoder technology, and in particular to a standards-oriented QC-LDPC encoder architecture, encoding method and chip. Background Technology

[0002] As communication technologies evolve towards B5G and 6G, communication systems face diverse and extremely differentiated demands. Especially in scenarios requiring high throughput, low latency, and low power consumption, existing QC-LDPC encoder designs face numerous challenges. General-purpose architectures retain a large amount of redundant resources to accommodate a wide range of parameter sets, leading to hardware inefficiency. Manually customized designs are time-consuming, lack flexibility, and struggle to adapt to future standard changes. Existing automated tools do not fully utilize the cyclic shift parallelism of QC-LDPC, making it difficult to optimize circuit timing and energy efficiency. Furthermore, most solutions lack tape-out verification at advanced process nodes, resulting in uncertainties regarding practical feasibility and reliability. In addition, the trade-off between hardware overhead and coding performance is difficult to reconcile; high throughput or full-mode compatibility often comes with increased area and power consumption, while saving resources may sacrifice speed or flexibility.

[0003] Therefore, how to provide an automated QC-LDPC encoder architecture that can achieve low hardware overhead, low latency and high performance in different application scenarios, while taking into account the requirements of high throughput and extremely low power consumption, is a technical problem that urgently needs to be solved. Summary of the Invention

[0004] This application provides a standard-oriented QC-LDPC encoder architecture, encoding method, and chip, which achieves the technical effect of providing an automated QC-LDPC encoder architecture with low hardware overhead, low latency, and high performance in different application scenarios, while taking into account the requirements of high throughput and extremely low power consumption.

[0005] To achieve the above objectives, the main technical solutions adopted in this application include: In a first aspect, embodiments of this application provide a standards-oriented QC-LDPC encoder architecture, comprising a first cyclic shift network module, a timing buffer module, a core transform module, and a second cyclic shift network module connected in sequence; wherein... The first cyclic shift network module is configured to switch between the operation modes of the adaptation matrix A or matrix C1, and is used to perform corresponding matrix operations on the bit sequence of the information to be encoded, respectively generating intermediate variables and the first operation value of the extended check bit; The timing buffer module is connected between the first cyclic shift network module and the core transformation module, and is used to perform timing alignment buffering on the intermediate variables to match the processing delay of the first cyclic shift network module. The core transformation module is adapted to a pre-defined standard double diagonal matrix structure, receives intermediate variables after time alignment buffering, solves for the core check bit, and transmits the core check bit to the second cyclic shift network module. The second cyclic shift network module adapts to the operation mode of matrix C2 and is used to perform matrix operations on the core parity bit to generate the extended parity bit second operation value; The extended check bit is obtained by XORing the first and second operation values ​​of the extended check bit to obtain an encoding result that conforms to the preset standard, including the sequence of information bits to be encoded, the core check bit, and the extended check bit.

[0006] In one embodiment, the encoder architecture is configured based on a five-dimensional parameter model, which includes the verification matrix extension dimension, base graph type, inter-layer folding factor, first-layer inner folding factor, and second-layer inner folding factor.

[0007] In one embodiment, the first cyclic shift network module and / or the second cyclic shift network module include multiple parallel cyclic shift networks, each of which includes multiple parallel cyclic shifters, an XOR reduction tree, and an output control accumulator; wherein the number of parallel paths of the first cyclic shift network module and the second cyclic shift network module is determined by the inter-layer folding factor; the number of parallel paths of the cyclic shifters within each cyclic shift network in the first cyclic shift network module is determined by the first intra-layer folding factor; and the number of parallel paths of the cyclic shifters within each cyclic shift network in the second cyclic shift network module is determined by the second intra-layer folding factor.

[0008] In one embodiment, the timing buffer module includes a timing matching buffer, the depth of which is determined based on the number of rows in the base map matrix corresponding to the preset standard and the interlayer folding factor.

[0009] In one implementation, the data bit width of the first cyclic shift network module and the second cyclic shift network module is consistent with the value of the extended dimension of the parity check matrix.

[0010] In one implementation, the base graph type corresponds to the base graph matrix parameters of the preset standard, and the diagonal matrix structure dimension of the core transformation module matches the base graph matrix parameters corresponding to the base graph type.

[0011] Secondly, embodiments of this application provide an encoding method for the standard-oriented QC-LDPC encoder architecture described above, the encoding method comprising: The information bit sequence to be encoded is input into the first cyclic shift network module, and the first cyclic shift network module is controlled to switch to the operation mode of the adaptation matrix A, and matrix A operation is performed on the information bit sequence to be encoded to generate intermediate variables and output them to the timing buffer module. The timing buffer module performs timing alignment buffering on the received intermediate variables, matches the processing delay of the first cyclic shift network module, and then transmits the intermediate variables to the core transformation module; wherein, the core transformation module performs operations on the intermediate variables based on a preset standard double diagonal matrix structure to obtain the core check bit, and transmits the core check bit to the second cyclic shift network module. The first cyclic shift network module is controlled to switch to the operation mode of adaptation matrix C1, and performs matrix C1 operation on the information bit sequence to be encoded to generate the first operation value of extended parity bit; and the second cyclic shift network module performs matrix C2 operation on the core parity bit based on the operation mode of adaptation matrix C2 to generate the second operation value of extended parity bit. The first and second values ​​of the extended check bit are XORed to obtain the extended check bit. The extended check bit is then integrated with the preset standard, which includes the sequence of information bits to be encoded, the core check bit, and the extended check bit, to output an encoding result that conforms to the preset standard.

[0012] In one embodiment, the encoding method further includes: The encoding operation cycle is determined based on the received base map type, inter-layer folding factor, first-layer inner folding factor, and second-layer inner folding factor. The hardware control logic and corresponding timing scheduling strategy for generating a finite state machine based on the encoding operation cycle are used by the finite state machine to perform coordinated operations of the encoder architecture by outputting matrix selection signals, data enable signals, and output enable signals according to the hardware control logic and the timing scheduling strategy; wherein, The matrix selection signal is output by the finite state machine to the mode control terminal of the first cyclic shift network module. When the finite state machine is in matrix A operation, it triggers the first cyclic shift network module to adapt to matrix A operation; when the finite state machine is in matrix C1 operation, it triggers the first cyclic shift network module to switch to adapt to matrix C1 operation. The data enable signal is output by the finite state machine in a time-division manner to the first cyclic shift network module, the timing buffer module, the core transformation module and the second cyclic shift network module, and is triggered to receive input data and perform operations only when the corresponding module enters the operation; The output enable signal is output by the finite state machine in a time-division manner to the first cyclic shift network module, the timing buffer module, the core transformation module, and the second cyclic shift network module. Data transmission is triggered only when the corresponding module completes the corresponding operation and generates valid data.

[0013] Thirdly, embodiments of this application provide a chip that is implemented using the standard-oriented QC-LDPC encoder architecture described above.

[0014] In one embodiment, the chip operates at a maximum frequency of 1.28 GHz at 1.0 V, with a peak throughput of 144.94 Gbps, an energy efficiency of 1.4 Tbit / J, and a normalized area efficiency of 426.3 Gbps / mm². 2 .

[0015] The technical solution provided by one or more embodiments of this application includes a first cyclic shift network module, a timing buffer module, a core transformation module, and a second cyclic shift network module, which can achieve optimization in multiple aspects. The first cyclic shift network module adapts to the operation mode of either matrix A or matrix C1 by switching the adaptation matrix, and the second cyclic shift network module adapts to matrix C2. The timing buffer module ensures matching processing delays between modules by aligning intermediate variables in time, thereby reducing latency and improving encoding real-time performance. At the same time, through parallel computation and XOR generation of extended parity bits, this architecture can improve encoding performance and throughput, meeting the requirements of efficient encoding. Overall, this encoder architecture balances the needs of low hardware overhead, low latency, high throughput, high performance, and extremely low power consumption in different application scenarios. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 A schematic diagram of a standards-oriented QC-LDPC encoder architecture provided for embodiments of this application; Figure 2 A schematic diagram illustrating four basic modules constituting an encoder, provided for embodiments of this application; Figure 3 A flowchart of the encoding method for a standard-oriented QC-LDPC encoder architecture provided in the embodiments of this application; Figure 4 A schematic diagram of timing scheduling relationships provided for embodiments of this application; Figure 5 A three-dimensional design space distribution diagram provided for embodiments of this application; Figure 6 Micrograph of a QC-LDPC encoder chip implemented using TSMC 65-nm CMOS technology, provided for embodiments of this application; Figure 7 A schematic diagram of the actual test environment for the QC-LDPC encoder chip provided in the embodiments of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] With the continuous evolution of communication technologies, especially in B5G and 6G scenarios, the demands and challenges faced by communication systems are becoming increasingly complex. Low-density parity-check codes (LDPC), as a highly promising error-correcting code, have become one of the core technologies of modern communication systems due to their performance approaching the Shannon limit. Particularly in fifth-generation mobile communication (5G), quasi-cyclic low-density parity-check codes (QC-LDPC) have been standardized for use in data channels. This development signifies the important role of LDPC coding in applications requiring high throughput, high reliability, and low latency.

[0020] However, with the advancement of technology towards B5G and 6G, the diversity and extreme differentiation of application scenarios have brought significant challenges to the design of existing QC-LDPC encoders. Firstly, the inefficiency of the general-purpose architecture becomes a prominent issue. In practical applications, especially in resource-constrained scenarios such as the Internet of Things (A-IoT), existing general-purpose architectures must retain a large amount of redundant resources to accommodate a wide set of parameters (such as the 51 boost values ​​in 5G NR), leading to wasted hardware resources and low efficiency. This inefficiency is particularly severe in certain scenarios, especially for applications that need to operate in low-power environments with low transmission throughput, where the performance of existing architectures falls far short of expectations.

[0021] Secondly, the limitations of manual custom design have become one of the challenges faced by QC-LDPC encoders. For different application scenarios, such as high-throughput immersive communication (eMBB+) or ultra-low latency ultra-reliable low-latency communication (HRLLC), traditional hardware custom design typically requires a significant amount of manual intervention and customization. This not only leads to excessively long design cycles but also makes it difficult to flexibly cope with the ever-changing and expanding parameter space in future 6G standards. In other words, manual design cannot meet rapidly changing market demands.

[0022] Meanwhile, while existing automation tools (such as High-Level Synthesis (HLS) tools can generate hardware code, they often suffer from control redundancy and fail to fully exploit the unique parallelism of QC-LDPC encoders. For example, QC-LDPC encoders possess cyclic shift bit-level parallelism, one of their core advantages, but existing tools fail to fully utilize this feature, resulting in generated circuits that do not achieve ASIC-level optimization in terms of timing and energy efficiency. Therefore, in high-performance hardware design, existing automation tools often cannot meet the demands of complex applications.

[0023] More seriously, existing automated design solutions suffer from deficiencies in physical verification. While many automated generation solutions have demonstrated potential in simulation environments, most have not yet been verified on actual advanced process nodes. Therefore, their feasibility and performance advantages in actual physical implementation cannot be proven, leading to significant uncertainty in their practical application and compromising their reliability in real-world deployments.

[0024] Furthermore, existing technologies face an irreconcilable conflict between hardware resource overhead (area / power consumption) and coding performance (throughput / latency). Existing solutions often prioritize high throughput or full-mode compatibility, accepting high hardware costs and resulting in excessive increases in area and power consumption; conversely, saving hardware resources often comes at the expense of coding speed or flexibility. This inherent contradiction makes it difficult for existing technologies to meet the demands of diverse application scenarios, especially in future B5G / 6G environments where solutions requiring low hardware overhead, low latency, and flexible adaptation to different standards are particularly urgent.

[0025] In summary, how to provide an automated QC-LDPC encoder architecture that can achieve low hardware overhead, low latency, and high performance in different application scenarios, while taking into account the requirements of high throughput and extremely low power consumption, is a technical problem that urgently needs to be solved.

[0026] To address the aforementioned technical problems, this application provides a standards-oriented QC-LDPC encoder architecture. Figure 1This diagram illustrates a standards-oriented QC-LDPC encoder architecture as provided in an embodiment of this application. Before explaining the encoder architecture, to better understand the technical solution of this application, the four basic modules constituting the encoder are first described, such as... Figure 2 As shown: Figure 2 In the diagram, (a) represents the register bank, used for timing buffering and delay alignment of the input data u. The diagram contains N parallel registers D, each independently delaying one data path. The data enable signal de controls the sampling and updating of the register bank; the registers will latch new data only when de is valid.

[0027] Figure 2 (b) in the diagram represents an XOR reduction tree. An XOR reduction tree combines multiple data items through an XOR operation to form a tree structure, ultimately yielding one or more summary results. Specifically, it uses N parallel inputs u1,…,u N The bitwise XOR operation outputs a single reduced result x, corresponding to the XOR reduction tree within each cyclic shifter network (CSN). In matrix operations, the outputs of multiple parallel cyclic shifters (CS) are XORed and reduced to obtain the result of a single-row matrix operation.

[0028] Figure 2 In the diagram, (c) is a multiplexer (MUX), which selects one of the two inputs R1 and R2 as the output R based on the selection signal sel. This corresponds to the matrix mode switching logic in the first cyclic shift network module. Controlled by the matrix selection signal sel output by the finite state machine (FSM), it selects the operation of matrix A in stage one and switches to the operation of matrix C1 in stage two, thereby realizing the reuse of the same hardware for different matrix operations.

[0029] Figure 2 In the diagram, (d) represents the circular shift operation, used to perform a circular shift on the input data u. The shift can be selected via the select signal sel. The input data u passes through a k-stage register delay chain, generating a k-bit circularly shifted version. The original data u and the shifted data are fed into an XOR gate, and then the output is selected via the MUX: when sel=1, the XOR result of the shifted data is output; when sel=0, the original data is output.

[0030] Please see Figure 1 As shown, the encoder architecture includes a first cyclic shift network module, a timing buffer module, a core transform module, and a second cyclic shift network module connected in sequence; wherein, The first cyclic shift network module is configured to switch between the operation modes of matrix A or matrix C1, and is used to perform corresponding matrix operations on the bit sequence of information to be encoded, generating intermediate variables and the first operation value of the extended check bit, respectively.

[0031] Specifically, the first cyclic shift network module is configured with two flexibly switchable operation modes, corresponding to matrix A operation mode and matrix C1 operation mode, respectively. The switching between the two modes is controlled by a matrix selection signal (sel signal) at the hardware level. This control signal is output by a finite state machine (FSM) automatically generated by this application, ensuring that the timing of the mode switching is precisely matched with the overall operation cycle of the encoder. Among them, matrix A and matrix C1 are both core sub-matrices of the base map matrix under a preset standard (e.g., 5G NR standard).

[0032] It should be noted that the encoder architecture of this application embodiment revolves around the QC-LDPC code structure and its two-step encoding algorithm defined by the preset standard (5G NR standard). Specifically, the QC-LDPC code under the 5G NR standard adopts a graph lifting structure, and its parity-check matrix H is derived from the graph matrix H... b The base graph matrix H is generated by expanding the parity check matrix by dimension Z. b Structurally, it is divided into an information matrix A, a core check matrix B with a double diagonal structure, extended check matrices C1 and C2, and an identity matrix I, specifically satisfying the following block structure: Correspondingly, the QC-LDPC coding of the 5G NR standard adopts a two-step coding mechanism, which involves solving H·c T =0 yields the complete codeword, i.e., the encoding result c=[s,p,q], where s is the sequence of information bits to be encoded, p is the core parity bit, and q is the extended parity bit. It should be noted that, utilizing the double-diagonal structure of the core parity matrix B, the encoding process is divided into two steps. First, the core parity bit p is calculated by solving the upper part (A and B): Where λ = [λ1, λ2, λ3, λ4] represents intermediate variables. The extended check bit q is then explicitly calculated from the lower part (C1 and C2), while using s and the newly calculated core check bit p: In other words, the operation mode of the adaptation matrix A corresponds to the first step of the standard to solve for the core check bits, which requires As. T Operations; the operation mode of the adaptation matrix C1 corresponds to the second step of the standard to solve for the C1s required for the extended parity bits. T Operations; the operation mode of the adaptation matrix C2 corresponds to the second step of the standard to solve for the extended parity bit C2p. T Operations. The hardware functions of the first cyclic shift network module are adapted to the above matrix structure and two-step encoding algorithm. It is configured to switch between the operation modes of matrix A or matrix C1, corresponding to the matrix operation requirements of different stages in the two-step encoding.

[0033] Furthermore, the first cyclic shift network module contains multiple parallel-operating cyclic shift networks (CSNs). Each CSN is an independent parallel computing unit, and each CSN further contains multiple parallel-operated cyclic shifters (CSs). During the encoding operation, the multiple CSNs execute operations synchronously in parallel. Simultaneously, within a single CSN, multiple cyclic shifters (CSs) also execute shift operations synchronously in parallel, thus forming a two-layer parallel computing architecture: the outer layer consists of parallel operations of multiple CSNs, and the inner layer consists of parallel operations of multiple cyclic shifters (CSs) within each CSN.

[0034] In matrix A operation mode, the input of the first cyclic shift network module receives the bit sequence s of the information to be encoded from the external input. The bit sequence of the information to be encoded is distributed in parallel to each cyclic shift network (CSN). Multiple cyclic shifters (CS) within each cyclic shift network (CSN) synchronously complete the cyclic shift operation according to the corresponding matrix coefficients. The operation results of each cyclic shifter (CS) are summarized into the XOR reduction tree within its respective cyclic shift network (CSN) for XOR reduction processing. Then, the data is integrated by the output control accumulator T1 (pop). Finally, the corresponding intermediate variable λ is output in parallel by each cyclic shift network (CSN). i , ), where a i,j Let s be the element in the i-th row and j-th column of matrix A. j Let k be the j-th information bit in the sequence of information bits to be encoded. b The base graph matrix H b The number of information bits. The intermediate variables output by the first cyclic shift network module are directly transmitted to... Figure 1 The timing buffer module that interfaces with it first allows intermediate variables to enter a specific depth D within the timing buffer module. (mb-4) / P The timing matching buffer BUF1 is used to match the processing delay of the cyclic shift network to achieve timing alignment. The output of the timing matching buffer BUF1 is then sent to the 4Z-bit depth register D4, and after being timed by the register, it is input to the core transform module. .

[0035] Using the double diagonal matrix structure corresponding to the core transformation module to find The core check bit p is obtained. The hardware operation logic in the above stages can be uniformly expressed by the hardware generation formula as follows: in, For the corresponding dimension, , and → represent the unit operation structure; → represents the data flow description symbol. The term refers to the parallel operation operator, which corresponds to a two-level parallel architecture (outer layer multi-cyclic shift network CSN parallelism, inner layer multi-cyclic shifter CS parallelism). This represents the first cyclic shift network module, which uses the first-layer inner folding factor P1 to configure the number of parallel paths and the matrix selection signal sel to switch between matrix A and C1 modes. BUF1 is the timing matching buffer, D4(de) is a 4Z-bit register controlled by the data enable signal, and B(B) is the core transformation module for solving the double diagonal matrix.

[0036] In matrix C1 operation mode, the first cyclic shift network module reuses the same hardware structure (parallel cyclic shifter, XOR reduction tree, output control accumulator T1(pop)) to perform the parallel operation corresponding to matrix C1 on the bit sequence of information to be encoded. The operation process is consistent with the hardware logic of matrix A operation mode, only switching the matrix parameters corresponding to the operation through the matrix selection signal, and finally generating the first operation value of the extended parity bit. .

[0037] The timing buffer module is connected between the first cyclic shift network module and the core transformation module. It is used to perform timing alignment buffering of intermediate variables to match the processing delay of the first cyclic shift network module.

[0038] Specifically, because the first cyclic shift network module adopts a multi-layer parallel operation structure, it contains multiple parallel cyclic shift networks (CSNs). Each CSN is further composed of multiple parallel cyclic shifters (CS), an XOR reduction tree, and an output control accumulator (T1, pop). The information bit sequence undergoes multiple hardware processing steps, including parallel distribution, cyclic shifting operations, multi-level XOR reduction, and output integration, inevitably introducing a fixed computational delay. Furthermore, under the configurable architecture of the five-dimensional parameter model LDPC(Z,BG,P,P1,P2) in this invention, different folding factors P, P1, and P2 will change the number of parallel operations and the data bit width, causing the processing delay of the first cyclic shift network module to vary with parameter configuration. If the intermediate variable λ is directly... i (i.e., the result of matrix operation As) T If the input is to the core transformation module, the timing of the data received by the core transformation module will be disordered due to the mismatch of the operation link delay, which will lead to encoding errors.

[0039] Therefore, the timing buffer module is used to process the intermediate variable λ output by the first cyclic shift network module. i The timing alignment buffer consists of a two-level buffer structure: the first level is the parameterized depth. The timing-matched buffer BUF1 has a depth determined by the interlayer folding factor P and the number of rows m in the base graph matrix. bThe first stage is jointly determined to match the overall processing delay of the cyclic shift network (CSN) under different configurations; the second stage is a fixed-depth 4Z-bit register D4, used to achieve bit-level synchronization and signal stabilization.

[0040] Through the above two-level buffering, the timing buffer module can dynamically absorb the computational delay generated by the first cyclic shift network module, ensuring the smooth operation of the intermediate variable λ. i Arriving precisely at the same moment the core transformation module is ready, this ensures that the core transformation module receives stable and valid input data under correct timing, and thus correctly executes equation Bp based on the double diagonal matrix structure. T =As T Solve for the core check bit p.

[0041] The core transformation module adapts to a pre-defined standard double diagonal matrix structure, receives intermediate variables after time-aligned buffering, solves for the core check bit, and transmits the core check bit to the second cyclic shift network module.

[0042] Specifically, the hardware architecture of the core transformation module is specifically adapted to the preset standard, namely the 5G NR standard adopted in this embodiment. The base map matrix H in the 5G NR standard... b It possesses a dual-diagonal structure, specifically the dual-diagonal matrix structure presented by the core parity check matrix B. The input of this module is connected to the output of the timing buffer module, used to receive the intermediate variable λ after timing alignment and buffer stabilization. i The intermediate variable corresponds to the matrix operation As. T The result is the only input for solving the core check bit p.

[0043] During the operation, the core transformation module uses the two-step coding algorithm defined by the 5G NR standard and leverages the inherent double-diagonal structure of the parity-check matrix B to directly perform equation solving. It can be solved quickly through simple recursion or shift operations, without the need for complex matrix inversion or iterative operations, thus greatly reducing hardware overhead and improving solution speed.

[0044] After solving the equation, the core transformation module outputs the core parity bit p, which is used as part of the encoding result for the final codeword synthesis. On the other hand, it is directly transmitted to the second cyclic shift network module as input data for the calculation of the second-stage extended parity bit q, enabling the second cyclic shift network module to perform matrix C2 operations based on the core parity bit p, thus achieving parallel operation collaboration with the first cyclic shift network module.

[0045] The second cyclic shift network module adapts to the operation mode of matrix C2 and is used to perform matrix operations on the core parity bits to generate the second operation value of the extended parity bits.

[0046] The extended check bit is obtained by XORing the first and second operation values ​​of the extended check bit, so as to obtain the encoding result that conforms to the preset standard and includes the sequence of information bits to be encoded, the core check bit, and the extended check bit.

[0047] Specifically, after solving for the core parity bit p, the encoder enters the stage of calculating the extended parity bit q. In this stage, the core parity bit p obtained in the first stage and the original input bit sequence s to be encoded are used to start two sets of CSN networks in parallel for synchronous operation: the first set of CSN networks reuses the hardware structure of the first cyclic shift network module and continues to process the bit sequence s to be encoded by switching to matrix C1 operation mode; the second set of CSN networks is implemented by the second cyclic shift network module and operates in matrix C2 operation mode to perform parallel operation on the core parity bit p.

[0048] Consistent with the architecture of the first cyclic shift network module, the second cyclic shift network module also adopts a two-layer parallel operation structure, consisting of multiple parallel cyclic shift networks (CSNs), with each CSN further containing multiple parallel cyclic shifters (CSs). The core parity bit p is distributed in parallel to each cyclic shift network (CSN) and its corresponding cyclic shifter (CS). Each CS synchronously performs cyclic shift operations according to the coefficients in matrix C2. The multiple shift results are aggregated into the XOR reduction tree within the respective cyclic shift network (CSN) for bit-by-bit XOR reduction. Then, the output control accumulator completes data integration and timing calibration, ultimately generating the second operation value C2p of the extended parity bit. T .

[0049] The computation results of the two CSN networks are calculated according to the extended parity bit calculation formula q. T =C1s T +C2p T Performing bitwise XOR synthesis yields the complete extended parity bit q. The hardware generation formula for this process is: in, This represents the second cyclic shift network module configured with parallel paths by the second-layer inner folding factor P2.

[0050] This embodiment provides a standards-oriented QC-LDPC encoder architecture, including a first cyclic shift network module, a timing buffer module, a core transform module, and a second cyclic shift network module, which can achieve optimization in multiple aspects. The first cyclic shift network module adapts to either matrix A or matrix C1 in a switchable operation mode, while the second cyclic shift network module adapts to matrix C2. The timing buffer module ensures matching processing delays between modules by aligning intermediate variables in time, thereby reducing latency and improving real-time encoding. Simultaneously, through parallel computation and XOR generation of extended parity bits, this architecture can improve encoding performance and throughput, meeting the requirements of efficient encoding. Overall, this encoder architecture balances the needs of low hardware overhead, low latency, high throughput, high performance, and extremely low power consumption in different application scenarios.

[0051] In some alternative implementations, the encoder architecture is configured based on a five-dimensional parameter model, which includes the verification matrix extension dimension, base graph type, inter-layer folding factor, first-layer inner folding factor, and second-layer inner folding factor.

[0052] Specifically, the parity check matrix extension dimension Z is used to characterize the boost value Z used by the QC-LDPC basemap matrix in the boosting operation in the 5G NR standard, determining the data bit width and processing granularity of the internal computational units of the encoder; the basemap type BG corresponds to different basemap matrix structures specified in the 5G NR standard, used to match encoding requirements of different lengths and bit rates; the inter-layer folding factor P is used to configure the number of cyclic shift networks (CSNs) running in parallel in the encoder, determining the overall computational parallelism of the encoder; the first-layer inner folding factor P1 and the second-layer inner folding factor P2 are used to configure the parallel number of cyclic shifters (CS) inside the first and second cyclic shift network modules, respectively, further refining the balance between hardware resources and computational efficiency. By combining and configuring the above five-dimensional parameters, the encoder's computational depth, parallel scale, data bit width, and matrix structure can be flexibly adjusted, enabling the encoder to adapt to encoding requirements under different scenarios, throughputs, and hardware resource constraints.

[0053] This embodiment, through configuration based on a five-dimensional parameter model, allows the encoder architecture to flexibly adapt to the needs of different application scenarios according to the parameters. Expanding the parity-check matrix dimension Z optimizes the data bit width and processing granularity, ensuring low hardware overhead and improved resource utilization efficiency when performing high-complexity operations, thereby effectively reducing latency. The selection of the base map type enables the encoder to support encoding requirements of different lengths and bit rates, providing high throughput and reliability. The inter-layer folding factor improves parallel processing capabilities, reduces encoding time, and further reduces latency. Simultaneously, the first and second layer inner folding factors optimize the parallelism of the internal cyclic shift network modules, achieving efficient allocation of hardware resources and improved computational efficiency.

[0054] In some optional implementations, the first cyclic shift network module and / or the second cyclic shift network module include multiple parallel cyclic shift networks, each cyclic shift network including multiple parallel cyclic shifters, XOR reduction trees, and output control accumulators; wherein, the number of parallel paths of the first cyclic shift network module and the second cyclic shift network module is determined by the inter-layer folding factor; the number of parallel paths of the cyclic shifters within each cyclic shift network in the first cyclic shift network module is determined by the first layer inner folding factor; and the number of parallel paths of the cyclic shifters within each cyclic shift network in the second cyclic shift network module is determined by the second layer inner folding factor.

[0055] Specifically, the number of parallel paths in the first cyclic shift network module, i.e., the number of cyclic shift networks (CSNs), is equal to (m... b -4) / P,m b Let P be the number of rows in the 5G NR standard basemap matrix, and P be the inter-layer folding factor. This can be understood as follows: the smaller P is, the more parallel cyclic shift networks (CSNs) there are, the larger the data block that can be processed in a single operation, and the higher the coding throughput; conversely, the larger P is, the fewer cyclic shift networks (CSNs) there are, and the less hardware resources are consumed. Similarly, the first-layer folding factor P1 and the second-layer folding factor P2 are used to determine the number of cyclic shifters (CSs) in each cyclic shift network (CSN) within the first and second cyclic shift network modules, respectively, and also satisfy an inverse proportional relationship: For the first cyclic shift network module, the number of parallel paths of the cyclic shifters CS within each cyclic shift network CSN is determined by P1, and can be expressed as: number of parallel paths of cyclic shifter CS = k b / P1,k b The base graph matrix H b Number of information positions; For the second cyclic shift network module, the number of parallel paths of the cyclic shifter CS inside each cyclic shift network CSN is determined by P2, which can be expressed as the number of parallel paths of the cyclic shifter CS = 4 / P2.

[0056] The smaller P1 and P2 are, the more parallel paths (i.e., the number of cyclic shifters CS) are inside the corresponding cyclic shift network CSN, and the higher the computational parallelism of a single cyclic shift network CSN; the larger P1 and P2 are, the fewer parallel paths (i.e., the number of cyclic shifters CS) are inside, and the more resources are saved.

[0057] This embodiment sets the first and second cyclic shift network modules as multiple parallel cyclic shift networks. Each cyclic shift network contains multiple parallel cyclic shifters, an XOR reduction tree, and an output control accumulator, achieving a highly parallelized processing architecture. This significantly improves encoder throughput and reduces latency. The number of parallel paths in the first and second cyclic shift network modules is determined by the inter-layer folding factor. The first and second inner-layer folding factors determine the number of parallel paths of the cyclic shifters within each cyclic shift network. By adjusting the folding factor, the parallelism and hardware resource consumption of the encoder can be flexibly controlled: a smaller folding factor increases the number of parallel processing units, improving computational parallelism and throughput to meet high-performance requirements; a larger folding factor reduces hardware resource consumption and power consumption, achieving a simplified design. This enables the implementation of a low-hardware-overhead, low-latency, and high-performance automatic QC-LDPC encoder architecture in different application scenarios, effectively balancing high throughput and extremely low power consumption.

[0058] In some optional implementations, the timing buffer module includes a timing matching buffer, the depth of which is determined based on the number of rows in the basemap matrix corresponding to a preset standard and the interlayer folding factor.

[0059] Specifically, the timing buffer module is used to perform timing alignment and delay buffering on the intermediate variables output by the first cyclic shift network module to ensure that data is sent to the core transformation module at the correct time. In this embodiment, the depth of the timing matching buffer used by the timing buffer module is not a fixed value, but is determined by the number of rows m of the basemap matrix corresponding to the 5G NR standard. b Together with the interlayer folding factor P, its depth is determined by the following relationship: (m b -4) / P.

[0060] In some optional implementations, the data bit width of the first cyclic shift network module and the second cyclic shift network module is consistent with the value of the extended dimension of the parity check matrix.

[0061] Specifically, the parity check matrix extension dimension corresponds to the boost value Z used in the boosting operation of the QC-LDPC code base map matrix in the 5G NR standard, which directly determines the bit width of the encoder's internal modules. The first and second cyclic shift network modules, as the core hardware units for matrix operations in the encoder, have data path bit widths in their internal cyclic shifters, XOR reduction trees, output control accumulators, and other operational circuits that are identical to the boost value Z. By strictly aligning the data bit width of the cyclic shift network modules with the parity check matrix extension dimension Z, it is ensured that information bits and parity bits are processed in parallel with a full Z bit width during encoding operations. This ensures that data transmission, shift operations, and XOR reduction operations are fully matched with the matrix boosting rules of the 5G NR standard, avoiding data truncation, timing errors, or hardware redundancy caused by bit width mismatch. This guarantees that the encoder achieves stable, efficient, and standard-compatible encoding operations under different parameter configurations.

[0062] In some optional implementations, the basemap type corresponds to the basemap matrix parameters of a preset standard, and the dimensions of the double diagonal matrix structure of the core transformation module match the basemap matrix parameters corresponding to the basemap type.

[0063] Specifically, the core transform module is designed specifically for solving double-diagonal matrix equations. Its internal operational circuitry, including bit width, number of levels, and number of registers, maintains the same hardware structure as the basemap matrix parameters corresponding to the current basemap type. This ensures complete compatibility between the core parity bit calculation process and the coding algorithm defined in the 5G NR standard. The basemap type is passed to the encoder via a five-dimensional parameter model, enabling the core transform module to adaptively match the double-diagonal matrix size under the corresponding basemap. This ensures correct equation solving under different basemap configurations, outputting stable and valid core parity bits, thereby improving the versatility and configurability of the encoder architecture.

[0064] This embodiment aligns the basemap type with the basemap matrix parameters of a preset standard and ensures that the dimensions of the core transform module's double-diagonal matrix structure match the basemap type. This allows the encoder to accurately execute the core parity bit solving process compatible with the 5G NR standard. The core transform module can adaptively adjust parameters such as bit width, number of levels, and number of registers in its hardware structure according to different basemap types, ensuring effective equation solving under various basemap configurations. This design enhances the encoder's versatility and adaptability while achieving low hardware overhead and low latency.

[0065] This application also provides an embodiment of an encoding method for a standard-oriented QC-LDPC encoder architecture. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here. Figure 3 A flowchart of an encoding method for a standards-oriented QC-LDPC encoder architecture provided in an embodiment of this application is shown. The process may include the following steps: Step S1: Input the information bit sequence to be encoded into the first cyclic shift network module, control the first cyclic shift network module to switch to the operation mode of the adaptation matrix A, perform matrix A operation on the information bit sequence to be encoded, generate intermediate variables and output them to the timing buffer module.

[0066] Specifically, in matrix A operation mode, the first cyclic shift network module is configured into a two-level parallel operation structure according to the five-dimensional parameter model. Multiple cyclic shift networks (CSNs) within the module work synchronously. Multiple cyclic shifters (CSs) in each cyclic shift network (CSN) perform parallel cyclic shift operations on the bit sequence of information to be encoded. The operation results are then integrated with the XOR reduction tree and the output control accumulator to generate the corresponding matrix operation. intermediate variable λ i The intermediate variables are then output to the subsequent timing buffer module to complete the matrix operation in the first stage of encoding.

[0067] In step S3, the timing buffer module performs timing alignment buffering on the received intermediate variables, matches the processing delay of the first cyclic shift network module, and then transmits the intermediate variables to the core transformation module. The core transformation module performs calculations on the intermediate variables based on a preset standard double diagonal matrix structure to obtain the core check bit, and then transmits the core check bit to the second cyclic shift network module.

[0068] Specifically, the timing buffer module determines the buffer depth based on the number of rows in the base map matrix and the inter-layer folding factor, performs delay alignment processing on intermediate variables, eliminates timing deviations caused by parallel operations in the preceding stages, and ensures that intermediate variables enter the core transformation module under the correct timing. The core transformation module adapts to the double diagonal matrix structure specified in the 5G NR standard, based on equations... The intermediate variables are solved to obtain the core check bit p, and the core check bit is used as the input data for the second stage operation and transmitted to the second cyclic shift network module.

[0069] Step S5: Control the first cyclic shift network module to switch to the operation mode of adaptation matrix C1, perform matrix C1 operation on the bit sequence of information to be encoded, and generate the first operation value of extended parity bit; and the second cyclic shift network module performs matrix C2 operation on the core parity bit based on the operation mode of adaptation matrix C2, and generates the second operation value of extended parity bit.

[0070] Specifically, in the extended parity bit calculation stage, the first cyclic shift network module reuses the original computational hardware structure and, by switching to matrix C1 operation mode, performs another operation on the original sequence of bits to be encoded to obtain the first value of the extended parity bit. Simultaneously, the second cyclic shift network module operates in parallel, performing parallel matrix operations on the core parity bit p according to the operation rules of matrix C2 to obtain the second operation value of the extended parity bit. The two sets of operations are executed simultaneously, improving the overall encoding throughput.

[0071] Step S7: Perform an XOR operation on the first and second values ​​of the extended parity bit to obtain the extended parity bit, and integrate the preset standard sequence of information bits to be encoded, the core parity bit, and the extended parity bit to output the encoding result that conforms to the preset standard.

[0072] Specifically, according to the extended parity bit operation relationship The two sets of operation values ​​are XORed bitwise to synthesize the extended parity bit q; finally, the information bit to be encoded s, the core parity bit p, and the extended parity bit q are combined into a complete codeword c=[s,p,q], and the QC-LDPC encoding result that meets the preset standards such as 5G NR is output, thus completing the entire encoding process.

[0073] This embodiment provides an encoding method for a standard-oriented QC-LDPC encoder architecture. By reusing hardware resources in different operation modes through the first cyclic shift network module, it achieves efficient processing of the bit sequence of information to be encoded and the extended parity bit, thereby reducing hardware overhead. The timing buffer module performs precise alignment of intermediate variables, eliminating timing deviations caused by parallel operations and effectively reducing processing latency. The parallel operation of the core transform module, the first cyclic shift network module, and the second cyclic shift network module, as well as the XOR processing of the extended parity bit, significantly improves the encoding throughput and overall operation performance.

[0074] In some alternative implementations, the encoding method further includes: The encoding operation cycle is determined based on the received base map type, inter-layer folding factor, first-layer inner folding factor, and second-layer inner folding factor. The hardware control logic and corresponding timing scheduling strategy for generating a finite state machine based on the encoding operation cycle are used by the finite state machine to perform coordinated operations on the encoder architecture by outputting matrix selection signals, data enable signals, and output enable signals according to the hardware control logic and timing scheduling strategy; wherein, The matrix selection signal is output from the finite state machine to the mode control terminal of the first cyclic shift network module. When the finite state machine is in matrix A operation, it triggers the first cyclic shift network module to adapt to matrix A operation; when the finite state machine is in matrix C1 operation, it triggers the first cyclic shift network module to switch to adapt to matrix C1 operation. The data enable signal is output by the finite state machine in a time-division manner to the first cyclic shift network module, the timing buffer module, the core transformation module and the second cyclic shift network module. The input data is received and the operation is executed only when the corresponding module enters the operation. The output enable signal is output by the finite state machine in a time-division manner to the first cyclic shift network module, the timing buffer module, the core transformation module, and the second cyclic shift network module. Data transmission is triggered only when the corresponding module completes the corresponding operation and generates valid data.

[0075] Specifically, based on the received base map type, inter-layer folding factor P, first-layer inner folding factor P1, and second-layer inner folding factor P2, the encoding operation cycle is dynamically determined, and the hardware control logic and timing scheduling strategy of the finite state machine (FSM) are generated accordingly. The FSM outputs a matrix selection signal sel, a data enable signal de, and an output enable signal pop based on the aforementioned hardware control logic and timing scheduling strategy, thereby realizing the collaborative operation and deterministic timing control of each module in the encoder architecture.

[0076] The matrix selection signal sel is output from the finite state machine (FSM) to the mode control terminal of the first cyclic shift network module: when the finite state machine is in the matrix A operation state, the matrix selection signal sel triggers the first cyclic shift network module to work in the matrix A operation mode; when the finite state machine switches to the matrix C1 operation state, the matrix selection signal sel triggers the first cyclic shift network module to reuse the hardware structure and switch to the matrix C1 operation mode, realizing the multi-mode reuse of the same set of operation modules.

[0077] The data enable signal `de` is output by the finite state machine (FSM) in a time-division multiplexing manner to the first cyclic shift network module, the timing buffer module, the core transformation module, and the second cyclic shift network module. The data enable signal `de` is valid only when the corresponding module enters the computation phase, triggering the module to receive input data and initiate the corresponding computation. Similarly, the output enable signal `pop` is also output by the finite state machine (FSM) to the above modules in a time-division multiplexing manner. The output enable signal `pop` is valid only after the module completes the computation and generates valid data, triggering the module to transmit the computation result to the next stage.

[0078] Through the aforementioned parameterized automatic generation method, the encoder can automatically generate the state transitions and control signal logic of the finite state machine (FSM) based on the configuration parameters, achieving precise scheduling of signals such as sel, de, and pop. This deterministic scheduling mechanism based on periodic calculation ensures the predictability and compatibility of hardware behavior. The specific timing scheduling relationship is as follows: Figure 4 As shown, the entire encoding process can be divided into two stages: Stage 1 and Stage 2. The time consumed in each stage and the total delay satisfy the following relationship: Phase 1 is the matrix A operation and core check bit p solution stage, which mainly completes the matrix A operation of information bit s by the first cyclic shift network module, the timing buffer, and the double diagonal solution of the core transformation module. The time consumption of this stage is: Phase two involves parallel computation of matrices C1 and C2, and the XOR synthesis of extended parity bits q. The first and second cyclic shift network modules operate in parallel, and the overall time consumption is determined by the longer of the two computation paths. The time consumption for this phase is: Between stage one and stage two, there is a fixed timing buffer and state transition overhead, with a fixed overhead of 5 cycles. Therefore, the total encoding latency of the encoder is: Combination Figure 4 As shown in the timing scheduling diagram, the finite state machine (FSM) outputs the matrix selection signal sel, the data enable signal de, and the output enable signal pop in sequence according to the above periodic division, realizing pipelined collaborative operation between modules: after the completion of the first stage, it switches to the second stage for parallel operation, and after the operation is completed, it outputs the complete codeword. There is no data conflict, no idle waiting, and no control ambiguity throughout the process, ensuring that the encoder has deterministic, predictable, and highly robust hardware behavior under different five-dimensional parameter configurations.

[0079] To fully verify the technical effectiveness and design flexibility of the five-dimensional parameterized configurable encoder architecture of this application, this embodiment performs a full traversal of the five-dimensional parameter space through the encoder, completes simulation and performance evaluation under different folding factor configurations, and generates a three-dimensional performance view including throughput, coding latency, and energy efficiency, such as... Figure 5The encoder provided in this application provides a three-dimensional design space distribution map of throughput-latency-energy efficiency in a 6G scenario, obtained by traversing the five-dimensional parameter space. Data points are distinguished by shape according to the inter-layer folding factor P and by color according to the intra-layer parallelism (P1, P2), visually representing the performance range corresponding to different folding architectures. This design space exploration, by traversing key parameters such as the inter-layer folding factor P, the first-layer intra-layer folding factor P1, and the second-layer intra-layer folding factor P2, obtains the performance distribution corresponding to different parameter combinations, thereby selecting the optimal hardware implementation configuration between theoretical expectations and actual simulation results.

[0080] Since there are certain deviations between the theoretically calculated clock cycle and operation delay and the highest operating frequency and dynamic power consumption obtained after actual circuit synthesis, placement and routing, this application does not directly use theoretically derived values ​​to determine the final configuration. Instead, it first determines the reasonable range of values ​​for the folding factor based on architectural constraints, and then performs small-scale traversal and simulation on the parameter combinations within the range. Finally, it uses the comprehensive performance of throughput, delay, area and energy efficiency as the selection basis.

[0081] Depend on Figure 5 The three-dimensional performance distribution shown can be used to derive: In eMBB+ application scenarios with high speed and large bandwidth, the ultimate throughput can be achieved by increasing the scale of parallel computing. The parameter combination (P=2, P1=2, P2=1) is the optimal configuration point, and its simulated peak throughput can reach 597Gbps, which is suitable for communication systems with extremely high throughput requirements.

[0082] In A-IoT application scenarios that are geared towards the Internet of Things and require low power consumption and wide coverage, hardware area and power consumption are the primary optimization goals. By increasing the interlayer folding factor P, the number of parallel computing units can be reduced. P=42 is the optimal configuration point for pursuing the ultimate area efficiency, which can achieve standard-compatible coding functions with minimal resource consumption.

[0083] To balance throughput, power consumption, area, and physical feasibility, and to adapt to the actual engineering needs in resource-constrained scenarios, a balanced optimal configuration LDPC (Z=64, BG1, P=21, P1=1, P2=1) was selected from the aforementioned three-dimensional design space for tape-out verification. This configuration is located in the optimal trade-off region between energy efficiency and routing density in the five-dimensional parameter space. It avoids the power consumption and area overhead caused by high parallelism, while ensuring the throughput and latency performance required by the actual system. At the same time, it significantly reduces dynamic power consumption, improves routing feasibility, and has excellent physical implementation characteristics.

[0084] The simulation and design space exploration results above fully demonstrate that the encoder architecture proposed in this application can be flexibly configured to adapt to multiple application scenarios, and can maintain a better wiring density while significantly reducing dynamic power consumption.

[0085] This application also provides a chip implemented using the aforementioned standards-oriented QC-LDPC encoder architecture. The chip operates at a maximum frequency of 1.28 GHz at 1.0V, with a peak throughput of 144.94 Gbps, an energy efficiency of 1.4 Tbit / J, and a normalized area efficiency of 426.3 Gbps / mm². 2 .

[0086] Specifically, the chip is fabricated using TSMC's 65-nm CMOS process, with a total die area of ​​1.44 mm². 2 It measures 1.20mm × 1.20mm, with the core logic area for encoding alone being 0.34mm². 2 While ensuring high throughput performance, it significantly saves hardware resources. The chip integrates a complete encoding core, SPI configuration interface, output buffer, and on-chip phase-locked loop (PLL), and has complete functions such as independent clock generation, parameter configuration, and data input / output. It can be directly integrated as an IP core into various baseband processing and communication systems, and has complete SoC integration capabilities and engineering practicality.

[0087] Through packaging and actual testing, the chip has been verified to achieve a maximum operating frequency of 1.28GHz under a core voltage of 1.0V, with a stable clock provided by an on-chip PLL. The peak encoding throughput reaches 144.94Gbps, while the core power consumption is only 103mW, corresponding to an energy efficiency of 1.4Tbit / J and a single-bit encoding energy consumption as low as 0.71pJ. Simultaneously, the normalized area efficiency reaches 426.3Gbps / mm². 2 .

[0088] Please see Figure 6 The micrographs of the QC-LDPC encoder chip implemented using TSMC 65-nm CMOS technology provided for the embodiments of this application visually demonstrate the physical layout and core module distribution of the chip; Figure 7 This is a schematic diagram of the actual test environment of the QC-LDPC encoder chip provided in the embodiments of this application, illustrating the functional verification and performance acquisition scenarios of the chip in an actual test platform.

[0089] Experimental results demonstrate that the QC-LDPC encoder chip automatically generated through parameterization in this application achieves excellent performance in key indicators such as standard compatibility, area, power consumption, and throughput. Compared with traditional manually designed FPGA or ASIC implementations, it achieves orders-of-magnitude improvements in energy efficiency and area efficiency while maintaining full compatibility with the 5G NR standard. The aforementioned physical implementation and experimental performance data fully validate the correctness, efficiency, and engineering feasibility of the proposed configurable encoder architecture based on five-dimensional parameters and folding factors, two-level parallel computing structure, and deterministic FSM timing scheduling strategy in silicon physical implementation.

[0090] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0091] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.

[0092] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A standards-oriented QC-LDPC encoder architecture, characterized in that, It includes a first cyclic shift network module, a timing buffer module, a core transform module, and a second cyclic shift network module connected in sequence; wherein, The first cyclic shift network module is configured to switch between the operation modes of the adaptation matrix A or matrix C1, and is used to perform corresponding matrix operations on the bit sequence of information to be encoded, respectively generating intermediate variables and the first operation value of the extended check bit; The timing buffer module is connected between the first cyclic shift network module and the core transformation module, and is used to perform timing alignment buffering on the intermediate variables to match the processing delay of the first cyclic shift network module. The core transformation module is adapted to a pre-defined standard double diagonal matrix structure, receives intermediate variables after time alignment buffering, solves for the core check bit, and transmits the core check bit to the second cyclic shift network module. The second cyclic shift network module adapts to the operation mode of matrix C2 and is used to perform matrix operations on the core parity bit to generate the extended parity bit second operation value; The extended check bit is obtained by XORing the first and second operation values ​​of the extended check bit to obtain an encoding result that conforms to the preset standard, including the sequence of information bits to be encoded, the core check bit, and the extended check bit.

2. The encoder architecture of claim 1, wherein, The encoder architecture is configured based on a five-dimensional parameter model, which includes the verification matrix extension dimension, base graph type, inter-layer folding factor, first-layer inner folding factor, and second-layer inner folding factor.

3. The encoder architecture of claim 2, wherein, The first cyclic shift network module and / or the second cyclic shift network module include multiple parallel cyclic shift networks, each of which includes multiple parallel cyclic shifters, XOR reduction trees, and output control accumulators; wherein the number of parallel paths of the first cyclic shift network module and the second cyclic shift network module is determined by the inter-layer folding factor; the number of parallel paths of the cyclic shifters within each cyclic shift network in the first cyclic shift network module is determined by the first intra-layer folding factor; and the number of parallel paths of the cyclic shifters within each cyclic shift network in the second cyclic shift network module is determined by the second intra-layer folding factor.

4. The encoder architecture according to claim 2, characterized in that, The timing buffer module includes a timing matching buffer, the depth of which is determined based on the number of rows in the base map matrix corresponding to the preset standard and the interlayer folding factor.

5. The encoder architecture of claim 2, wherein, The data bit width of the first cyclic shift network module and the second cyclic shift network module is consistent with the value of the extended dimension of the parity check matrix.

6. The encoder architecture of claim 2, wherein, The base graph type corresponds to the base graph matrix parameters of the preset standard, and the double diagonal matrix structure dimension of the core transformation module matches the base graph matrix parameters corresponding to the base graph type.

7. An encoding method for a standards-oriented QC-LDPC encoder architecture as described in any one of claims 1-6, characterized in that, The encoding method includes: The information bit sequence to be encoded is input into the first cyclic shift network module, and the first cyclic shift network module is controlled to switch to the operation mode of the adaptation matrix A, and matrix A operation is performed on the information bit sequence to be encoded to generate intermediate variables and output them to the timing buffer module. The timing buffer module performs timing alignment buffering on the received intermediate variables, matches the processing delay of the first cyclic shift network module, and then transmits the intermediate variables to the core transformation module; wherein, the core transformation module performs operations on the intermediate variables based on a preset standard double diagonal matrix structure to obtain the core check bit, and transmits the core check bit to the second cyclic shift network module. The first cyclic shift network module is controlled to switch to the operation mode of adaptation matrix C1, and performs matrix C1 operation on the information bit sequence to be encoded to generate the first operation value of extended parity bit; and the second cyclic shift network module performs matrix C2 operation on the core parity bit based on the operation mode of adaptation matrix C2 to generate the second operation value of extended parity bit. The first and second values ​​of the extended check bit are XORed to obtain the extended check bit. The extended check bit is then integrated with the preset standard, which includes the sequence of information bits to be encoded, the core check bit, and the extended check bit, to output an encoding result that conforms to the preset standard.

8. The encoding method according to claim 7, characterized in that, The encoding method further includes: The encoding operation cycle is determined based on the received base map type, inter-layer folding factor, first-layer inner folding factor, and second-layer inner folding factor. The hardware control logic and corresponding timing scheduling strategy for generating a finite state machine based on the encoding operation cycle are used by the finite state machine to perform coordinated operations of the encoder architecture by outputting matrix selection signals, data enable signals, and output enable signals according to the hardware control logic and the timing scheduling strategy; wherein, The matrix selection signal is output by the finite state machine to the mode control terminal of the first cyclic shift network module. When the finite state machine is in matrix A operation, it triggers the first cyclic shift network module to adapt to matrix A operation; when the finite state machine is in matrix C1 operation, it triggers the first cyclic shift network module to switch to adapt to matrix C1 operation. The data enable signal is output by the finite state machine in a time-division manner to the first cyclic shift network module, the timing buffer module, the core transformation module and the second cyclic shift network module, and is triggered to receive input data and perform operations only when the corresponding module enters the operation; The output enable signal is output by the finite state machine in a time-division manner to the first cyclic shift network module, the timing buffer module, the core transformation module, and the second cyclic shift network module. Data transmission is triggered only when the corresponding module completes the corresponding operation and generates valid data.

9. A chip implemented using the standards-oriented QC-LDPC encoder architecture as described in any one of claims 1-6.

10. The chip according to claim 9, wherein the chip has a maximum operating frequency of 1.28 GHz, a peak throughput of 144.94 Gbps, an energy efficiency of 1.4 Tbit / J, and a normalized area efficiency of 426.3 Gbps / mm² at 1.0 V. 2 .