LDPC decoder with multi-core architecture

The multi-core LDPC decoder architecture addresses the inefficiencies of conventional 5G decoders by using a limited set of predefined configurations, enabling parallel decoding and reducing hardware complexity, thus enhancing performance for satellite communications.

FR3155991A1Pending Publication Date: 2025-05-30AIRBUS DEFENCE & SPACE SAS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
FR2023012667
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Conventional 5G LDPC decoders face significant complexity and inefficiency due to their ultra-flexibility, leading to increased hardware complexity, power consumption, and reduced throughput, making them unsuitable for satellite communications.

Method used

A multi-core LDPC decoder architecture is proposed, which includes a group of cores with shared memory storing a limited number of predefined LDPC configurations. Each core can be dynamically configured to decode LDPC code words using these predefined configurations, optimizing memory usage and hardware complexity.

Benefits of technology

The multi-core approach enables parallel decoding of multiple codewords, optimizes memory resources, and reduces hardware complexity and power consumption, making the decoder more efficient and suitable for satellite communications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a multi-core decoder (31) for decoding low density parity check code words (LDPC code). The multi-core decoder comprises at least one group of several cores (10), several of these cores (10) being each capable of simultaneously decoding a different LDPC code word. Said at least one group comprises a memory (20) shared between the different cores (10) in which a number NC of predefined configurations is stored. Each predefined configuration comprises a binary parity matrix corresponding to an LDPC code. Each core is adapted to be dynamically configured with any one of the NC predefined configurations stored in the shared memory (20) in order to decode an LDPC code word using this configuration. Figure for abstract: Fig. 20
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: LDPC decoder with multi-core architecture Field of invention

[0001] The present invention belongs to the field of low density parity check (LDPC) codes. In particular, the invention relates to a multi-core architecture for an LDPC decoder. State of the art

[0002] LDPC codes are currently used in several communication technologies, notably for the IEEE 802.16 (WiMAX), IEE 802.lin (Wi-Fi) standards, the 5G standard of the 3GPP organization (“3rd Generation Partnership Project”), the DVB-S2 standard (“Digital Video Broadcasting, 2nd Generation”), or the CCSDS C2 (“Consultative Committee for Space Data Systems, C2”) space communications standard•

[0003] A binary LDPC code is a linear error correcting code defined by a low-density binary parity matrix (i.e., the number of undamaged elements of the matrix is ​​relatively small compared to the size of the matrix).

[0004] To reduce the hardware implementation complexity of an LDPC decoder, it is known to use particular structures of the parity matrix. In particular, quasi-cyclic LDPC codes (QC-LDPC for "Quasi-Cyclic Low Density Parity Check") are defined by parity matrices composed of sub-matrices of size Z x Z. The term Z is generally called "expansion factor". The sub-matrices of size Z x Z are generally called "circulant matrices". A parity matrix of a QC-LDPC code has a layered structure which makes it possible to parallelize the calculations of the parity check messages within a layer.

[0005] The 3GPP TS 38.212 technical specification document describes the LDPC configurations that must be supported by a device compatible with the 5G standard.

[0006] These 5G LDPC codes are compatible in terms of coding rate for Incremental Redundancy Hybrid Automatic ReQuest (IR-HARQ). HARQ is a technique for retransmitting data that has been corrupted by the communication channel. In the Chase Combining (CC-HARQ) variant, the same coding rate is used for each retransmission. In the IR-HARQ variant, a different coding rate is used for each retransmission.

[0007] 5G LDPC codes are based on quasi-cyclic parity matrices. Each 5G LDPC configuration corresponds to a triplet of parameters including the size K of the encoded word (this is the number of useful bits in the LDPC code word), the coding rate R (this roughly corresponds to the ratio K / N between the size K of the encoded word and the size N of the LDPC code word) and the expansion factor Z. By defining the parameters K and R, we can deduce the parameter Z to use, then dynamically construct the parity matrix to use.

[0008] A very large number (several thousand) of different configurations are supported by the specification (K can vary from a few bits to several thousand bits, Z can take about fifty possible values, and a large number of coding rates R are supported). This makes 5G LDPC codes very flexible and adaptable to a very wide range of transmission conditions.

[0009] This ultra-flexibility of LDPC codes, however, leads to a very significant complexity of the decoders implementing them. In addition, conventional 5G LDPC decoders are generally not very efficient for transmissions using a fixed spectral band. The high flexibility on the expansion factor Z leads to a complexity of the permutation network of the layered architecture. This complexity increases the footprint of the registers and logic gates of the circuit. This complexity also leads to a reduction in the overall throughput of the decoder. In addition, the parallel architecture of a conventional 5G LDPC decoder is often suboptimal for low values ​​of the expansion factor Z.

[0010] The constraints relating to the space domain (constraints in terms of size, weight and energy consumption in particular) mean that conventional 5G LDPC decoders are not suitable for being carried on board a satellite. Presentation of the invention

[0011] The present invention aims to remedy all or part of the drawbacks of the prior art.

[0012] To this end, and according to a first aspect, the present invention proposes a multi-core LDPC decoder for decoding low density parity check code words, or LDPC code. The multi-core decoder comprises at least one group of several cores, several of these cores being able to simultaneously decode each a different LDPC code word. Said at least one group comprises a memory shared between the different cores in which a number Nc of predefined configurations is stored, each predefined configuration being associated with a parity matrix corresponding to an LDPC code. Each core is adapted to be dynamically configured with any one of the Nc predefined configurations stored in the shared memory in order to decode an LDPC code word using this configuration.

[0013] This multi-core approach is particularly advantageous since it allows the decoding multiple codewords in parallel. Thus, at a given time, different codewords can be decoded in parallel by different cores according to different configurations. The multi-core approach also makes it possible to optimize memory resource requirements, by sharing memory between all the cores in the same group.

[0014] In particular embodiments, the invention may further comprise one or more of the following features, taken individually or in any technically possible combination.

[0015] In particular embodiments, the predefined Ncconfigurations form a subgroup, in the strict sense, of the 5G LDPC configurations defined in the 3GPP TS 38.212 standard.

[0016] In particular embodiments, each core comprises a local configuration memory whose size is at least equal to the largest of the sizes of the Nc predefined configurations stored in the shared memory, and the size of the local memory is at least four times smaller than the size of the shared memory.

[0017] The smaller the ratio between the size of the local memory of a core and the size of the memory shared between the different cores of the same group, the more efficient the sharing of memory between the different cores is.

[0018] In particular embodiments, the number Nc of predefined configurations stored in the shared memory is between 8 and 512, or even between 8 and 256, or even between 8 and 128.

[0019] By selecting a limited number of LDPC configurations, it is possible to precalculate and store the parity matrices corresponding to the selected configurations. It is then no longer necessary to dynamically construct the parity matrix each time a codeword must be decoded with a new LDPC configuration. This also makes it possible to optimize the hardware resources of the decoder, particularly in terms of memory requirements, but also in terms of hardware complexity at the level of the multiplexers and the permutation network, and in terms of energy consumption. This can also make it possible to optimize the scheduling of the parity check nodes for each configuration, in order to maximize the performance of the decoder.

[0020] In particular embodiments, each of the Nc configurations comprises an optimized scheduling of parity check nodes of the core. The scheduling is defined according to a degree of parity of each parity check node for the configuration considered.

[0021] Such arrangements make it possible to limit situations where certain control nodes cannot be executed because the output data of other control nodes are not available. In other words, this makes it possible to limit the loss of performance related to “data hazards”: the number of waiting cycles in the decoding process is lower, and the overall throughput of the decoder is improved.

[0022] In particular embodiments, decoding an LDPC codeword by a core comprises performing one or more iterations until a stopping criterion is satisfied. The stopping criterion comprises a check whether a maximum number of iterations is reached and, for a new codeword to be decoded by a core, the maximum number of iterations is defined as a function of a load rate of the multi-core decoder.

[0023] It is advantageous to have a lower maximum number of iterations when the load rate of the multi-core decoder is high. Indeed, when the load rate of the decoder is high, it is advantageous to quickly release the cores to be able to decode new code words.

[0024] In particular embodiments, the load rate at a given instant is calculated as a function of a number of cores simultaneously occupied in decoding an LDPC code word at said instant.

[0025] In particular embodiments, the maximum number of iterations is also defined as a function of the size of the new code word.

[0026] In particular embodiments, the maximum number of iterations is also defined as a function of a coding rate used for the new code word.

[0027] In particular embodiments, the maximum number of iterations is also defined as a function of an estimated signal-to-noise ratio for the new codeword.

[0028] In particular embodiments, the multi-core decoder comprises a scheduling module adapted to schedule the cores by favoring, for the decoding of a new code word, the selection of a core for which the last configuration used by the core for the decoding of a previous code word is the same as the configuration to be used for the decoding of the new code word.

[0029] Such arrangements can make it possible to avoid having to reload a new configuration into the local memory of a core for the decoding of a new code word. This makes it possible to optimize the decoding time and to limit concurrent accesses to the shared memory.

[0030] According to a second aspect, the present invention relates to a reception device comprising a multi-core decoder according to any one of the embodiments described above.

[0031] According to a third aspect, the present invention relates to a satellite comprising such a reception device. Presentation of figures

[0032] The invention will be better understood on reading the following description, given by way of non-limiting example, and made with reference to Figures 1 to 22 which represent:

[0033] [Fig-1] a schematic representation of a parity matrix of an LDPC code,

[0034] [Fig.2] a schematic representation of a bipartite graph (Tanner graph) associated with a parity matrix,

[0035] [Fig.3] an illustration of a method used to obtain a parity matrix of a quasi-cyclic LDPC code,

[0036] [Fig.4] a schematic representation of an exemplary implementation of a method for decoding an LDPC codeword with layered ordering,

[0037] [Fig.5] a schematic representation of an exemplary embodiment of an LDPC decoder according to the invention,

[0038] [Fig.6] a schematic representation of an exemplary embodiment of a core of the decoder of [Fig.5], for implementing a decoding method such as that described in [Fig.4],

[0039] [Fig.7] a table showing the memory footprint for parity check messages (uncompressed) for different candidate configurations based on BG1,

[0040] [Fig.8] a graphical representation of the memory footprint for parity check messages (uncompressed) for different candidate configurations based on BG1,

[0041] [Fig.9] a graphical representation of the memory footprint for parity check messages (uncompressed) for different candidate configurations based on BG2,

[0042] [Fig. 10] a table showing the memory footprint for the a priori estimation variables for different candidate configurations based on BG1,

[0043] [Fig. 11] a graphical representation of the memory footprint for the a priori estimation variables for different candidate configurations based on BG1,

[0044] [Fig. 12] a graphical representation of the memory footprint for the a priori estimation variables for different candidate configurations based on BG2,

[0045] [Fig. 13] a table listing a set of selected configurations based on BG1,

[0046] [Fig. 14] a table listing a set of selected configurations based on BG2,

[0047] [Fig. 15] a graphical representation of the selected configurations listed in the tables of Figures 13 and 14,

[0048] [Fig. 16] a graphical representation of the degrees of the parity check nodes of the basic matrices BG1 and BG2,

[0049] [Fig. 17] a graphical representation of the degrees of the parity check nodes of the basic matrices BG1 and BG2, after scheduling, for the minimum coding rate,

[0050] [Fig. 18] a graphical representation of the degrees of the parity check nodes of the basic matrices BG1 and BG2, after scheduling, for a coding rate 1 / 2,

[0051] [Fig. 19] a graphical representation of the degrees of the parity check nodes of the basic matrices BG1 and BG2, after scheduling, for the maximum coding rate,

[0052] [Fig.20] a schematic representation of another example of an embodiment of a LDPC decoder, with a multi-core architecture comprising a group of four cores,

[0053] [Fig.21] a schematic representation of a second example of embodiment of a multi-core LDPC decoder comprising two groups of four cores,

[0054] [Fig.22] a schematic representation of a third example of embodiment of a multi-core LDPC decoder with a two-level hierarchy.

[0055] In these figures, identical references from one figure to another designate identical or similar elements. For reasons of clarity, the elements represented are not necessarily on the same scale, unless otherwise stated. Detailed description of the invention

[0056] In the remainder of the description, the case of an LDPC decoder for space communications is considered in a non-limiting manner. The decoder may in particular be part of a receiver device on board a satellite intended to be placed in orbit around the Earth.

[0057] An LDPC code is defined by a parity matrix. [Fig.l] schematically represents a parity matrix H of an LDPC code. We consider the case of a binary LDPC code. The parity matrix H is therefore a binary matrix, which means that each element of the matrix H is either a '0' or a '1'. We consider that the matrix H is of size M x N, with M and N positive integers. The matrix therefore has M rows and N columns. The parity matrix H is of low density, that is to say that the number of elements of the matrix equal to '1' is relatively small compared to the total number M x N of elements of the matrix. For example, the number of unnullified elements of the matrix is ​​less than 1% of the total number of elements of the matrix.

[0058] As illustrated in [Fig.2], an LDPC code can also be represented in the form of a bipartite graph G (Tanner graph) having connections between N nodes of variable VNn (n varying between 1 and N) and M parity check nodes CNm (m varying between 1 and M). Each non-zero element of the parity matrix H corresponds to a connection between a node of variable VNn and a parity check node CNm. Each row of the parity matrix H corresponds to a parity equation associated with a parity check node CNm. Each column of the parity matrix H corresponds to a variable associated with a variable node VNn. A code word to be decoded corresponds to a set of values ​​taken respectively by the variables associated with the N variable nodes (this is the set of estimated values ​​of the bits of the code word).

[0059] A code word may have a relatively large size, for example a size greater than or equal to one thousand bits (N > 1000), or even greater than ten thousand bits (N > 10000). We are considering the case of an LDPC decoder which supports different code word sizes, and different coding rates. The coding rate corresponds to the ratio between the number of useful bits in a code word and the total number of bits in a code word. The closer the coding rate is to '1', the lower the computational complexity and the higher the useful bit rate can be; in return, the error correction power is lower. Conversely, the closer the coding rate is to '0', the greater the error correction power; in return, the computational complexity is higher and the useful bit rate is lower. The number of lines M and the density of the parity matrix generally depend on the coding rate.

[0060] The decoding of an LDPC code word is based on an iterative exchange of information on the likelihood of the values ​​taken by the bits of the code word. The iterative decoding process is based on a belief propagation algorithm which relies on an exchange of messages between the variable nodes VNn and the parity check nodes CNm.

[0061] As illustrated in [Fig.2], a message sent by a parity check node CNm to a variable node VNn is denoted [3m>n. The value of a message [3m>n is calculated at the parity node CNm for each of the variable nodes VNn connected to the parity check node CNm on the graph G.

[0062] A message sent by a variable node VNn to a parity check node CNm is denoted an>m. The value of a message an>m is calculated at a variable node VNn for each of the parity nodes CNm connected to the variable node VNn on the graph G.

[0063] This is an iterative process: the messages an>m are calculated from the messages[3mn previously calculated, and the messages [3m>n are calculated from the messagesan>m previously calculated. This iterative process takes as input a priori estimation variables of the code word which correspond for example to log-likelihood ratio (LLR) logarithms. These are values ​​representative of the probability that the value of a bit of the code word is equal to '1' or '0' (logarithm of the ratio between the probability that the value of the bit is equal to '0' and the probability that the value of the bit is equal to '1').

[0064] A posteriori estimation variables yn (n varying from 1 to N) of the bits of the code word are also calculated iteratively from the messages [3m>n. These yn values ​​are also representative of the probability that the value of a bit in the codeword is equal to '1' or '0'. They allow a decision to be made on the value of each of the bits in the codeword. A syndrome can then be calculated from the estimated values ​​of the bits in the codeword and the parity equations defined by the parity matrix H. If we denote by c = (ci, c2, ..., cN) the set of estimated values ​​of the bits in the codeword, then the syndrome s is defined by the matrix equation s = H * cT. A zero syndrome means that the estimated values ​​of the bits in the codeword satisfy the parity equations.

[0065] Different conventional algorithms can be considered for the iterative LDPC decoding process (BP-SPA, Min-Sum, Offset Min-Sum). These algorithms are known to those skilled in the art. Sections 2.3.1 and 2.3.3 of the document “Efficient Hardware Implementations ofLDPC Decoders through Exploiting Impreciseness in Message-Passing Decoding Algorithms”, TT Nguyen Ly (document subsequently referenced as “Refl”), respectively describe an example of an iterative LDPC decoding process with the BP-SPA algorithm and with the Min-Sum algorithm in the case of flood scheduling.

[0066] To reduce the hardware implementation complexity of the LDPC decoder, it is possible to use particular structures of the parity matrix H which give the matrix an organization in horizontal or vertical layers. For example, a horizontal layer of the parity matrix H can be defined as a set of consecutive rows defined in such a way that, for a given variable (i.e. for a given column of the parity matrix H), the layer has only one non-zero element.

[0067] This layered structure makes it possible to parallelize the calculations of the parity check messages within a layer because the parity equations of a layer do not involve a variable of the code word more than once. Indeed, if a layer has only one non-zero element for a given variable, this means that the variable nodes VNn connected to a parity check node CNm of a layer are not connected to another parity check node of said layer.

[0068] There are different ways to obtain a parity matrix H having a layered structure. In particular, and as illustrated in [Fig.3], it is possible to obtain a parity matrix H from a base matrix B of size R x C by replacing each element of the base matrix B with a matrix of size Z x Z corresponding either to a zero matrix, or to the identity matrix, or to a shift of the identity matrix. The parity matrix then has R x Z rows (M = R x Z) and C x Z columns (N = C x Z). The term Z is generally called an “expansion factor”. The sub-matrices of size Z x Z are generally called “circulant matrices”. The terms R, C, and Z are positive integers. An LDPC code defined by such a parity matrix H is called a "quasi-cyclic" LDPC code (QC-LDPC).

[0069] For example, and as illustrated in [Fig.3], each element of the basis matrix B is an integer having the value '-1', '0', or a value less than Z. An element of the basis matrix B having the value '-1' is replaced by the zero matrix; an element of the basis matrix B having the value '0' is replaced by the identity matrix; an element of the basis matrix B having a value d between 1 and (Zl) is replaced by an offset of value d of the identity matrix.

[0070] A horizontal layer of the parity matrix H can then be defined as a set of ZP consecutive rows of the parity matrix H coming from a row of the base matrix B, with ZP < Z.

[0071] In a horizontal layered architecture, the calculations are mainly centered on the parity check nodes CNm. The number ZP corresponds to the number of functional units used to execute in parallel the calculations performed at the parity check nodes. When ZP = Z, the level of parallelization is maximum.

[0072] Section 2.5.2 of the Refl document describes an example of an iterative LDPC decoding process with the Min-Sum algorithm in the case of horizontal layer scheduling. The posterior estimation variables yn are initialized with the a priori estimation variables (LLRs). The messages [3m>n are initialized to zero. Then, at each iteration, the different layers are processed successively. For each layer: the messages an>m are calculated from the messages [3m>n and the posterior estimation variables yn; the messages [3m>n are calculated from the messages an>m; the posterior estimation variables yn are calculated from the messages [3m>n; a partial syndrome can then be calculated from the a posterior estimation variables yn.

[0073] [Fig.4] schematically illustrates an example of implementation of a method 100 for decoding an LDPC code word with a horizontal layered architecture.

[0074] As illustrated in [Fig.4], the method 100 comprises performing one or more iterations until a stopping criterion is satisfied. Each iteration comprises successively processing the layers of the parity matrix H. The processing 110 of a layer comprises: - a calculation 111 of messages of variable an>m, for the nodes of variables VNn involved in said layer, - a calculation 112 of parity check messages [3m>n, for the parity check nodes CNm involved in said layer, - a calculation 113 of the a posteriori estimation variables yn, - a calculation 114 of a partial syndrome for said layer.

[0075] The calculation 111 of a variable message an>m is performed for each variable node VNn involved in the layer being processed and for each of the parity check nodes CNm connected to said variable node VNn. The messages an>m are calculated from the current values ​​of the a posteriori estimation variables yn and from the current values ​​of the parity check messages [3m>n. These current values ​​correspond either to the initialization values ​​(for the first iteration) or to the values ​​calculated during the previous iteration. For example, a message an>m is calculated such that an>m = yn - Pm,n-

[0076] The calculation 112 of a parity check message [3m>n is carried out for each parity check node CNm involved in the layer being processed and for each of the variable nodes VNn connected to said parity check node CNm. The messages [3m>n are calculated from the current values ​​of the variable messages an>m. For example, a message [3m>n is calculated by considering all the messages an>m associated with the parity check node CNm while excluding the message an>m associated with the variable node VNn; the absolute value of a message [3m>n is equal to the smallest absolute value of the messages an>m considered; the sign of a message [3m>n is equal to the product of the signs of the messages an>m considered.

[0077] The calculation 113 of a value of the a posteriori estimation variable yn is carried out for each bit of the code word. For example, the yn are calculated from the current values ​​of the parity check messages [3m>n and the current values ​​of the variable messages an>m such that yn = an>m + [3m>n.

[0078] The calculation 114 of a partial syndrome for the layer being processed is carried out by applying the parity equations of said layer to the a posteriori estimation variables yn. The partial syndrome is then a vector of size L.

[0079] The method 100 comprises, at the end of the processing of each layer, a check 120 whether the iteration is finished or not. The iteration is finished when all the layers have been processed.

[0080] At the end of an iteration, the method 100 comprises an evaluation 130 of a stopping criterion. It is for example conceivable to consider that the stopping criterion is satisfied when all the partial syndromes calculated respectively for the different layers are null. Other stopping criteria can however be envisaged: it is for example also possible to consider that the stopping criterion is satisfied when the partial syndromes of the different layers are all null for a predetermined number of successive iterations. According to yet another example, the stopping criterion is satisfied if, for a plurality of successive iterations, the number of iterations for which all the partial syndromes are null from which is subtracted the number of iterations for which at least one of the partial syndromes is non-zero is greater than or equal to a predetermined stopping threshold.

[0081] [Fig. 5] schematically represents an exemplary embodiment of an LDPC decoder 30 comprising at least one core 10 and a non-volatile memory 20 in which a number Nc of predefined configurations is stored. The core 10 is a computing unit (a processor) adapted to be dynamically configured with any one of the Nc predefined configurations stored in the memory 20 in order to decode an LDPC code word using this configuration. As illustrated in [Fig. 5], the core 10 comprises a local configuration memory 13 whose size is at least equal to the largest of the sizes of the Nc predefined configurations stored in the memory 20. Thus, each of the Nc configurations can be temporarily loaded into the local memory 13 for the decoding of an LDPC code word.A new configuration is loaded into the local memory 13 each time a new LDPC code word must be decoded with a different configuration than that which is currently loaded into the local memory 13. The core 10 is adapted to carry out the iterative LDPC decoding method 100 described above with reference to [Fig. 4]. The core 10 is for example implemented in the form of a specific integrated circuit of the ASIC type (acronym for “Application-Specific Integrated Circuit”), or a reprogrammable integrated circuit of the FPGA type (acronym for “Field-Programmable Gate Array”).

[0082] [Fig.6] schematically illustrates an exemplary embodiment of the core 10 described above with reference to [Fig.5]. It should however be noted that there are many possible architectures in the literature for implementing an LDPC core. In the example illustrated in [Fig.5], the core 10 comprises: - a first-in-first-out (FIFO) input buffer 11 for storing a data frame while another data frame is being processed, - an input alignment unit 12 for forming blocks of data bits to be decoded of the size of the parallelization factor ZP, - a local configuration memory 13, for example a volatile memory of the RAM type (“Random Access Memory”), in which configuration information relating to the LDPC code is stored, in particular the parity matrix H to be used (which can be stored in any suitable form), - a volatile memory 14 in which the current values ​​of the a posteriori estimation variables yn are stored, - a volatile memory 15 in which the current values ​​of the parity check messages [3m>n] are stored, - a processing unit 16 configured to execute the iterations of the decoding process, that is to say in particular to implement the network of per mutation (shift operations of the identity matrix), to perform the calculations of the a posteriori estimation variables yn, of the messages an>m and [3m>n, of the partial syndromes, and to determine whether the stopping criterion is satisfied, - a multiplexer 19 for directing into the memory 14 the values ​​of the a priori estimation variables (for the first iteration) or the values ​​of the a posteriori estimation variables yn calculated by the processing unit 16 (for the following iterations), - a volatile memory 17 in which the hard decision values ​​of the bits of the code word are stored, - an output alignment unit 18 for adapting the size of the decoded data bit blocks to the expected size at the output of the decoder 10. 5G Compatibility:

[0083] The 3GPP standard (acronym for "3rd Generation Partnership Project", a cooperation between telecommunications standards organizations), and more specifically the technical specification document 3GPP TS 38.212 (version 15.0.0 and later) describes the LDPC configurations that must be supported by a device compatible with the 5G standard (fifth generation of the 3GPP mobile communications standard).

[0084] Each 5G LDPC configuration corresponds to a triplet of parameters {K, R, Z] and a parity matrix. The parameter K corresponds to the size of the encoded word (this is the size of the useful data in the LDPC code word). The parameter Z corresponds to an expansion factor (“lifting size Z”). The parameter R corresponds to the coding rate. This corresponds substantially to the ratio between the size of the encoded word and the size of the LDPC code word (K / N ratio). More precisely, the coding rate R corresponds to the ratio K / (N-NP), where NP is a number of punctured bits.

[0085] 5G LDPC codes are based on quasi-cyclic parity matrices that can be constructed from a BG1 or BG2 base matrix (“Base Graph 1 or Base Graph 2”).

[0086] The basic matrix BG1 has a maximum of 46 rows and 68 columns. The basic matrix BG2 has a maximum of 42 rows and 52 columns. Each entry of the basic matrix can be extended with the expansion factor Z ("lifting size Z" in 3GPP TS 38.212). The number of rows and columns to be used depends on the coding rate R (the lower the coding rate, the larger the size of the basic matrix).

[0087] The choice of the basic matrix BG1 or the basic matrix BG2 is defined by the specification according to the values ​​of K and R (see section 6.2.2 of the 3GPP TS 38.212 document): • if K < 3824 and R < 0.67, then BG2 is selected, • if K < 292, then BG2 is selected, • if R < 0.25, then BG2 is selected, • otherwise, BG1 is selected.

[0088] Once the basic matrix is ​​selected, the specification defines a parameter Kb which represents the number of columns of the basic matrix to be used to process the information bits of the word of size K: • for BG1, Kb = 22; • for BG2: • if K > 640, then Kb = 10, • if 560 < K < 640, then Kb = 9, • if 192 < K < 560, then Kb = 8, • if K < 192, then Kb = 6.

[0089] The expansion factor Z is determined as the smallest value of Z satisfying Kb x Z > K.

[0090] Knowing the expansion factor Z, we can define from table 5.3.2-1 of the 3GPP TS 38.212 specification to which index iLS it corresponds. Table 5.3.2-2 then allows the basic matrix to be constructed by filling it with values ​​between 0 and (Zl). As described previously with reference to [Fig.3], an element of the basic matrix having the value '-1' is replaced by the zero matrix; an element of the basic matrix having the value '0' is replaced by the identity matrix; an element of the basic matrix having an integer value 'd' between 1 and (Zl) is replaced by an offset of value 'd' of the identity matrix.

[0091] A very large number of different configurations are supported by the specification. For example, K can vary from a few bits or a few tens of bits to several thousand bits (up to 8448 bits); the specification supports more than fifty possible values ​​for Z (values ​​between 2 and 384); a large number of different values ​​are possible for the coding rate R (the coding rate is directly related to the number of rows of the basic matrix).

[0092] The need to support a very large number of different configurations for 5G LDPC codes is explained in particular by the need to have compatible codes in terms of coding rate to support IR-HARQ. IR-HARQ is used in particular to introduce frequency diversity in order to minimize the impacts of multipath which is very often present in terrestrial communications. IR-HARQ is also used to optimize spectral efficiency for communications characterized by a very low propagation time.

[0093] 5G LDPC codes are therefore designed to be very flexible and adaptable to a very wide range of transmission conditions. This means that they use parity matrices of highly variable length, which require a significant amount of logic to be implemented in hardware. Also, the number of different configurations to be supported is so large (several thousand different configurations in 5G) that it is not feasible to pre-calculate the parity matrices associated with these configurations in order to store them in memory. It is then necessary to dynamically calculate the parity matrix of the selected configuration.

[0094] In return for this ultra-flexibility, conventional 5G LDPC decoders are generally relatively inefficient in terms of hardware for fixed spectral band processing. In particular, the high flexibility on the expansion factor Z leads to complexity in the permutation network of the layered architecture. This complexity increases the footprint of the registers and logic gates of the circuit. This complexity also leads to an increase in the number of pipeline stages in order to maintain a high maximum frequency of the circuit.This increase in the number of pipeline stages, however, reduces the overall throughput of the decoder due to data hazards which are usually resolved by wait cycles (the term "data hazards" is used here to refer to a situation in which an instruction cannot be executed because it depends on the value of another instruction which is not yet available).

[0095] With a conventional 5G decoder, it is possible to achieve high data rates with high values ​​for the coding rate R and for the expansion factor Z. Conventional 5G decoders use a layered architecture with several processing units working in parallel, as mentioned earlier. The parallelization factor ZP is often chosen to be as large as possible. However, for smaller expansion factor values, the parallel architecture often becomes suboptimal.

[0096] For high-speed satellite communications applications (and possibly for some other specific applications), the ultra-flexibility offered by 5G in the choice of LDPC configurations is not useful, and it can significantly increase the hardware complexity and power consumption of the device on board the satellite.

[0097] In the field of space communications, where the transmission channel is essentially of the AWGN type (English acronym for "Additive White Gaussian Noise"), the interest in using 1TR-HARQ is clearly reduced. The usefulness of considering all 5G LDPC codes (to maintain compatibility in terms of coding rate) is then reduced.

[0098] To obtain a 5G compatible LDPC decoder having good performance in terms of hardware efficiency, the present invention proposes to select a limited number of configurations from the very large number of LDPC configurations supported by the 3GPP TS 38.212 standard.

[0099] It may indeed be interesting to design an LDPC decoder that is compatible with 5G, in the sense that it supports at least certain 5G LDPC configurations, without necessarily supporting all 5G LDPC configurations. In a satellite communications system, the resources to be used to establish communication between a ground station (for example a 5G mobile terminal) and a satellite are allocated by a ground gateway station. Among these resources is the LDPC configuration to be used to encode (on transmission) and decode (on reception) a code word included in a message exchanged between the ground station and the satellite. The gateway station may in particular comprise a radio resource control module configured to allocate only 5G LDPC configurations included in the Nc predefined configurations.

[0100] In the following, Nc is the limited number of configurations supported by the decoder 30. This number Nc is advantageously low, such that the parity matrices associated with these configurations can be precalculated and stored in the memory 20. It is then no longer necessary to completely construct the parity matrix each time a new code word must be decoded with a new LDPC code. This makes it possible to limit the complexity of the decoder: the decoder does not need to implement the logic necessary for the construction of the parity matrix, and the permutation network can be simplified and optimized for the Nc configurations retained.

[0101] The predefined Ncconfigurations therefore form a strict subgroup of the 5G LDPC configurations defined in the 3GPP TS 38.212 standard.

[0102] In particular embodiments, the number Nc of predefined configurations is between 8 and 512, or even between 8 and 256, or even between 8 and 128. Such a value of Nc offers an interesting compromise between a sufficiently large number of configurations to cover a sufficient number of scenarios in terms of SNR and desired throughput, and a sufficiently small number of configurations to limit memory requirements and implementation complexity.

[0103] Each configuration corresponds to a triplet {K, R, Z], and to a parity matrix constructed as a function of at least part of the parameters K, R and Z. As explained previously, by defining K and R, we can deduce the basic matrix to use (BG1 or BG2), then the value of Z, then the complete parity matrix.

[0104] According to another example, it is possible to start by defining a set of some preferred values ​​for the expansion factor Z to which the N c configurations will be limited, then deduce different possible values ​​of K to support certain coding rates R. As explained previously, a parity matrix can be constructed for each triplet {K, R, Z] retained.

[0105] Advantageously, the set of values ​​taken by the parameter Z among the Nc predefined configurations comprises at least two different values. This thus makes it possible to manage different message sizes (different values ​​of K).

[0106] In particular embodiments, the set of values ​​taken by the parameter Z among the Nc predefined configurations comprises Nz different Z values; ordered in ascending order with the index i (i is an index varying between 0 and (Nz-1)), Nz being an integer between 2 and 32, or even between 4 and 16, the Z values; being advantageously chosen such that, for i varying between 1 and (Nz-1), the ratio Z is equal to an integer value. Z)

[0107] It is then advantageous to choose a value of Zo equal to the parallelization factor ZP of the decoder 30 (ZP = Zo). As a reminder, the parallelization factor ZP corresponds to the maximum number of functional units of the core 10 that can be used in parallel to execute the calculations of the parity check nodes. Such arrangements in fact make it possible to optimize the use of the hardware resources of the decoder for parallelization and, as will be seen later, to offer more flexibility for resolving data randomness situations.

[0108] It should however be noted that the value Zo is not necessarily the smallest value among the set of values ​​taken by the parameter Z among the Nc predefined configurations (nothing would prevent having a configuration with a value of Z strictly lower than ZP, for example if this configuration makes it possible to specifically deal with relatively rare cases of transmission of small messages).

[0109] To optimize the hardware implementation of the decoder, it is also advantageous to choose a value of Zo equal to a power of two (in other words, a value of Zo that can be written in the form 2n, where n is a positive integer). This in fact makes it possible to maximize the hardware use of the multiplexers and the permutation network.

[0110] As a non-limiting example, we can choose Zo = 32, Nz = 12, and Z; = (i + 1) x Zo for i varying between 0 and (Nz-1).

[0111] For the Zi values ​​retained, it is advantageous to select configurations which limit the size of the addressing bus of the volatile memory 15 of the core 10 in which the current values ​​of the parity check messages [3m>n] are stored.

[0112] The table in [Fig.7] gives the size in bits of the address bus of the memory 15 for different configurations with the values ​​Z; retained when the basic matrix to be used is BG1. Each column of the table corresponds to a particular value Z;, each row of the table corresponds to a particular coding rate value R. For a pair of values ​​{Z; ; R], it is possible to determine the corresponding value of the parameter K. Each box in the table of [Fig.7] therefore corresponds to a configuration candidate.

[0113]

[0114]

[0115] The size of the memory address bus 15 for parity check messages (uncompressed) [3m>n is equal to j, where Nconn corresponds to the number of unharmed elements (number of connections) in the parity matrix of the configuration considered. Nconn also corresponds to Zj multiplied by the number of unharmed elements in the mB rows and nB columns of the basic matrix used for the configuration considered. As illustrated in Figure 7, it is for example advantageous to select the supported configurations such that the size of the memory addressing bus 15 for the parity check messages [3m>n is less than or equal to 16 - pog (Zp) j (i.e. less than or equal to 11 when ZP = 32). [Fig.8] is a graphical representation of the table in [Fig.7] (it is a graphical representation of the bit size of the memory address bus 15 for different candidate configurations with the Z values; retained when the basic matrix to be used is BG1). [Fig.9] is a graphical representation of the bit size of the memory address bus 15 for different candidate configurations with the Z values ​​retained when the basic matrix to be used is BG2. It can be observed that the configurations based on BG1 are more restrictive than those based on BG2.

[0116] As described in section 4.3.1 of the Refl document, it is possible to "compress" the parity check messages [3m>n processed by a functional unit. For example, rather than storing the values ​​of the messages [3m>n associated with a control node, it may be sufficient to store the signs of these messages, the absolute value of at least two minimum values ​​among these messages, and the indexes associated with these minimum values ​​except one.

[0117] Thus, in the uncompressed case, each memory entry corresponds to a message [3m>n, whereas in the compressed case, each memory entry corresponds to the compressed data (signs, min, index) of a layer.

[0118] In the case where the parity check messages are compressed, the size of the addressing bus of the memory 15 for the parity check messages is equal to J" [Og ( m^z< jj, where mB corresponds to the number of rows of the basic matrix used, and corresponds to the number of layers in the parity matrix for the configuration considered (Mp • Zz- is equal to the number M of rows of the parity matrix for the configuration considered).

[0119] In the case where the parity check messages are compressed, it is for example advantageous to select the supported configurations such that the size of the memory address bus 15 for parity check messages either in less than or equal to 14- ' is less than or equal to 9 when ZP = 32).

[0120] For the values ​​Z; retained, it may also be advantageous to select configurations which limit the size of the addressing bus of the volatile memory 14 of the core 10 in which current values ​​of a posteriori estimation variables yn are stored (the term APP is also used to designate the a posteriori estimation variables).

[0121] The table in [Fig. 10] gives the size in bits of the address bus of memory 14 for different configurations with the values ​​Z; retained when the basic matrix to be used is BG1.

[0122] The size of the memory addressing bus 14 for the a posteriori estimation variables yn is equal to [Og AL jj, where N corresponds to the size of a code word for the configuration considered. The size N is equal to • Z^ where nB corresponds to the number of columns of the basic matrix used for the configuration considered.

[0123] As illustrated in Figure 10, it is for example advantageous to select the supported configurations such that the size of the addressing bus of the memory 14 for the a posteriori estimation variables yn is less than or equal to 14 _ pog (Zp) j (i.e. less than or equal to 9 when ZP = 32).

[0124] [Fig. 11] is a graphical representation of the table of [Fig. 10] (it is a graphical representation of the bit size of the address bus of the memory 14 for different candidate configurations with the Z values; retained when the basic matrix to be used is BG1). [Fig. 12] is a graphical representation of the bit size of the address bus of the memory 14 for different candidate configurations with the Z values; retained when the basic matrix to be used is BG2. It can be observed that the configurations based on BG1 are more restrictive than those based on BG2.

[0125] In the example considered, the selected configurations are listed in the tables in [Fig. 13] and 14. The table in [Fig. 13] lists fifty-one selected configurations based on BG1, while the table in [Fig.14] lists twenty selected configurations based on BG2. In this example, there are therefore a total of seventy-one configurations (Nc = 71) to be stored in the non-volatile memory 20.

[0126] In the tables of Figures 13 and 14, the column "Connections" indicates the number of connections (number of unharmed elements) in the basic matrix. The number Nconn of connections in the parity matrix corresponds to the number of connections in the basis matrix multiplied by the expansion factor Z.

[0127] [Fig. 15] is a graphical representation of the selected configurations listed in the tables of Figures 13 and 14. In this figure, the selected configurations based on BG1 are represented by squares, the selected configurations based on BG2 are represented by circles. It is clear from this representation that, advantageously, for K > 3824, the number of predefined configurations having an expansion factor equal to Zj is a decreasing function with the value of Zj. Similarly, for K < 3824 and R < 0.67, the number of predefined configurations having an expansion factor equal to Zj is a decreasing function with the value of Zj.

[0128] 5G LDPC codes have a very irregular distribution of CNm check node parity degrees. This means that different check nodes can have different (and potentially very different) numbers of data bits connected to them. The parity degree of a CNm parity check node corresponds to the number of data bits connected to it; it also corresponds to the number of parity equations in which the check node is involved.

[0129] [Fig. 16] illustrates this irregularity of 5G LDPC codes. [Fig. 16] represents the degrees of the parity check nodes of the basic matrices BG1 and BG2. In [Fig. 16], the degrees of the parity check nodes for BG1 are represented by squares, the degrees of the parity check nodes for BG2 are represented by crosses.

[0130] In a layered architecture, CNm parity check nodes can be executed in parallel. However, the uneven distribution of the parity degrees of the check nodes leads to situations where some check nodes cannot be executed until the output data of the other check nodes is available. It then becomes necessary to insert additional wait cycles into the decoding process. This reduces the overall throughput of the decoder.

[0131] To solve this problem, it is possible to schedule the control nodes in a particular and optimized way according to their parity degree. In particular, the control nodes can be scheduled in ascending order of parity degrees (this means that the control nodes with the fewest connected data bits are executed first).

[0132] It is also possible to optimize, using heuristics, the order of control nodes having the same degree of parity, or the order of calculations within a control node.

[0133] Each configuration selected and stored in the memory 20 can therefore advantageously include this optimized ordering of the control nodes (i.e. an indication of the order in which the control nodes must be executed). This approach cannot be implemented for conventional 5G LDPC decoders due to the large number of different configurations they must support. However, this approach applies very well to the decoder 30 according to the invention because the number Nc of predefined configurations is limited.

[0134] Figures 17, 18 and 19 respectively represent the degrees of the parity check nodes of the basic matrices BG1 and BG2 after ordering in ascending order of the parity degrees, for different values ​​of the coding rate R. [Fig. 17] corresponds to the minimum coding rate (case where all the rows of the basic matrices BG1 and BG2 are used), [Fig. 18] corresponds to a coding rate of value / 2 (the basic matrix BG1 then has twenty-four rows, while the basic matrix BG2 has twelve rows), and [Fig. 19] corresponds to a maximum coding rate (case where the basic matrices BG1 and BG2 have only four rows).

[0135] It appears from Figures 17 to 19 that, for BG1-based configurations, parity check nodes having a parity degree equal to nineteen are always present, regardless of the coding rate. Similarly, for BG2-based configurations, parity check nodes having a parity degree equal to eight or ten are always present, regardless of the coding rate.

[0136] For a given value of Z, the ordering of the control nodes of the configuration corresponding to the maximum coding rate can therefore be common to all the configurations. In other words, for a given value of Z, for each configuration selected and stored in the memory 20, the ordering of the control nodes can comprise a part common to all the configurations and a specific part specific to the configuration considered. The common part corresponds to the ordering of the control nodes of high degree; the specific part corresponds to the ordering of the control nodes of lower degree.

[0137] This common part makes it possible to limit the memory size necessary to memorize the configurations (the common part is memorized only once for all the configurations with the same value of Z). This concept is illustrated by the “Config size” and “Opt. config size” columns of the tables in Figures 13 and 14. For each selected configuration, the number indicated in the “Config size” column is representative of the total size necessary to memorize the ordering of the control nodes for this configuration. This number corresponds to the number Nconn of connections of the parity matrix divided by the parallelization factor ZP (it also corresponds to the number of connections of the basic matrix, indicated in the “Connections” column, multiplied by the ratio ^-).For a given Z value, the number shown in the "Config Size" column for the maximum coding rate (R=0.92) corresponds to the size of the common part; the number shown in the . The "Config Size" column for a lower coding rate corresponds to the sum of the size of the common part and the size of the specific part. For a given value of Z, the number indicated in the "Opt. Config Size" column corresponds either to the size of the common part (for the maximum coding rate), or to the size of the specific part (for lower coding rates). For each value of Z, it is sufficient to memorize the common part corresponding to the maximum coding rate (R=0.92) and the specific parts corresponding to the lower coding rates selected. Doing so makes it possible to almost halve the memory size required to memorize the selected configurations listed in the tables of figures 13 and 14.

[0138] As mentioned previously, it is advantageous to choose a parallelization factor ZP equal to the smallest Zo value among the supported Z; values, because this allows to limit the loss of parallelism while offering greater flexibility to resolve data hazard situations for high values ​​of the expansion factor. This choice goes against an established idea according to which it is appropriate to use a parallelization factor as large as possible. Using a parallelization factor as large as possible certainly has an advantage in limiting the decoding latency, but it offers less flexibility in the possibility of resolving data hazards with a particular ordering of the control nodes. In the field of space communications, the constraints relating to the latency of LDPC decoding are significantly less demanding than for terrestrial communications. Multi-core architecture:

[0139] The invention also relates to an LDPC decoder with a multi-core architecture. This multi-core architecture is particularly well suited to a 5G-compatible LDPC decoder as described above. However, nothing would prevent the implementation of an LDPC decoder with a multi-core architecture and a set of Nc predefined configurations without these configurations corresponding to 5G LDPC configurations.

[0140] Figures 20 to 22 show three examples 31-a, 31-b, 31-c of embodiment of a multi-core decoder 31 (in the following, the reference number 31 is used to generically indicate a multi-core LDPC decoder according to the invention).

[0141] In the example illustrated in [Fig.20], the decoder 31-a comprises a group of four cores 10, each core being similar to the core 10 of the decoder 30 described with reference to [Fig.5]. The memory 20 is shared between these four cores 10.

[0142] Each core 10 of the group is adapted to be dynamically configured with any one of the Nc predefined configurations stored in the shared memory 20 in order to decode an LDPC codeword using this configuration. Each predefined configuration comprises a binary parity matrix corresponding to a LDPC code (this is advantageously the complete parity matrix, and not just a sub-matrix allowing the parity matrix to be calculated).

[0143] Advantageously, several cores of the group can simultaneously decode each a different LDPC code word. Each time a new code word must be decoded by the decoder 31, a scheduling module 21 is configured to select an available core 10 from among the different cores 10 of the group, to load into the local memory 13 of the selected core 10 the LDPC configuration to be used to decode the code word, and to launch the decoding of the code word by the selected core 10.

[0144] Each time a core 10 has completed its process of decoding a code word, it sends information to the scheduling module 21 to indicate to it that it is available to decode a new code word.

[0145] This multi-core approach is particularly advantageous since it allows the decoding of several code words in parallel, i.e. simultaneously. Thus, at a given time, different code words can be decoded in parallel by different cores 10 according to different configurations.

[0146] The multi-core approach also makes it possible to optimize memory resource requirements. The memory cost is reduced because the memory 20 is shared with all the cores in the group. The smaller the ratio between the size of the local memory 13 of a core 10 and the size of the memory 20 shared between the cores 10 of the same group, the more efficient the sharing of the memory 20 between the different cores 10 is.

[0147] As seen previously, the decoding of an LDPC code word by a core 10 comprises the execution of one or more iterations until a stopping criterion is satisfied, for example based on the partial syndromes calculated for the different layers. It is also appropriate to set a maximum number of iterations after which the decoding process must stop.

[0148] It is important to correctly define this maximum number of iterations because it has a significant impact on the performance of the decoder. In particular, when the LDPC decoder is embedded in a satellite, constraints in terms of computing power and energy consumption lead to choosing a relatively low maximum number of iterations, generally of the order of ten. Each additional iteration therefore has an impact of the order of ten percent on performance.

[0149] To optimize the performance of the multi-core decoder 31, the maximum number of iterations is defined as a function of a load rate of the decoder. Each time a new code word must be decoded by a core 10 of the multi-core decoder 31, the maximum number of iterations to be used for the decoding of this new code word is calculated.

[0150] At a given instant, the load rate of the multi-core decoder 31 is calculated in function of the number of cores simultaneously occupied in decoding an LDPC code word. For example, for the multi-core decoder 31-a described with reference to [Fig.20], the load rate is respectively 0%, 25%, 50%, 75% or 100% depending on whether the number of cores 10 simultaneously occupied in decoding a code word is equal to zero, one, two, three or four. The load rate is for example calculated by the scheduling module 21.

[0151] To define the maximum number of iterations to be used for decoding a new code word, the load rate of the decoder 31 is calculated just before the start of decoding the new code word. The load rate may or may not take into account the fact that the core 10 responsible for decoding the new code word is busy. In the remainder of the description, we consider the case where the load rate does not take into account the occupancy of the core 10 which will be responsible for decoding the new code word. In this case, if the load rate is equal to 100%, then the scheduling module 21 must wait for a core 10 to become available to assign it the decoding of the new code word.

[0152] By way of non-limiting example, when the load rate is less than or equal to 50%, the maximum number of iterations is set at ten. When the load rate is strictly greater than 50%, then the maximum number of iterations is set at eight.

[0153] However, nothing would prevent setting another threshold value for the load rate, or other values ​​for the maximum number of iterations depending on the load rate relative to the threshold.

[0154] It is advantageous to have a lower maximum number of iterations when the load rate of the multi-core decoder 31 is high. Indeed, when the load rate of the decoder 31 is high, it is appropriate to quickly release the cores 10 to be able to decode new code words. Using a lower maximum number of iterations makes it possible to limit the decoding time of a code word (while accepting the possible risk of having errors in the decoding).

[0155] In particular embodiments, the maximum number of iterations can also be defined as a function of the coding rate R used for the new code word. As a non-limiting example, when the load rate is less than or equal to 50%, the maximum number of iterations is set at ten. When the load rate is strictly greater than 50% and the coding rate is 2 / 5, then the maximum number of iterations is set at eight. When the load rate is strictly greater than 50% and the coding rate is 8 / 9, then the maximum number of iterations is set at six. Here again, in variants, other values ​​could be envisaged for the load rate threshold, the coding rates and / or the values ​​of the maximum number of iterations.

[0156] It is advantageous to have a lower maximum number of iterations when the coding rate of the multi-core decoder 31 is higher because a higher coding rate important generally corresponds to a faster convergence of the decoding process.

[0157] In particular embodiments, the maximum number of iterations may also be defined as a function of the size N of the new code word. For example, the larger the size N of the code word, the lower the maximum number of iterations may be, since the decoding process tends to converge faster for large code word sizes.

[0158] In particular embodiments, the maximum number of iterations may also be defined as a function of an estimated signal-to-noise ratio (SNR) for the new codeword. For example, the higher the estimated value of the SNR, the lower the maximum number of iterations may be. Indeed, the decoding process also tends to converge more quickly for high SNR values.

[0159] The maximum number of iterations to be used by a core 10 for decoding a new codeword may be defined based on a combination of conditions depending on the load rate of the decoder, the coding rate R to be used for decoding the new codeword, the size N of the new codeword, or the estimated SNR level for the new codeword.

[0160] The multi-core decoder 31-a described above with reference to [Fig.20] comprises a single group of four cores 10. However, nothing would prevent, in variants, the design of a multi-core decoder comprising several groups of several cores. In the example illustrated in [Fig.21], the multi-core decoder 31-b comprises two groups of four cores. Each group comprises a memory 20 shared between the four cores 10 of the group.

[0161] Nothing would prevent having a different number of cores 10 within each group (for example only two cores, or eight cores), or having different numbers of cores for the different groups.

[0162] When there are several groups of cores, the load threshold at a given instant can be calculated as a function of the total number of cores simultaneously occupied in decoding an LDPC code word at said instant, taking into account all the groups. However, nothing would prevent, in variants, taking into account only the load rate of the group to which the core responsible for decoding the new code word belongs, or else defining an average load rate per group. The different possible methods for calculating the load rate and for defining the maximum number of iterations as a function of the load rate are only variants of the invention.

[0163] In particular embodiments, the scheduling module 21 is configured to schedule the different cores 10 of the multi-core decoder 31 by favoring, for the decoding of a new code word, the selection of a core 10 for in which the last configuration used by the core for decoding a previous codeword is the same as the configuration to be used for decoding the new codeword.

[0164] Such arrangements make it possible to avoid having to reload a new configuration into the local memory 13 of the core 10 for decoding the new code word. This makes it possible to optimize the decoding time and to limit concurrent accesses to the shared memory 20.

[0165] For this purpose, a buffer memory can be used by the scheduling module 21 to store the configuration last used by each of the cores 10 of the multi-core decoder 31. When a new code word is to be decoded according to a particular configuration, the scheduling module 21 is configured to identify, among the different cores, whether a core is available for which the configuration last used is the same as the configuration to be used for decoding the new code word. If such a core is identified, then it is preferentially selected for decoding the new code word.

[0166] As illustrated by [Fig.22], the concept of multi-core architecture of the LDPC decoder can also be generalized with several levels of hierarchy.

[0167] [Fig.22] schematically represents a third example 31-c of embodiment of a multi-core LDPC decoder with a two-level hierarchy. The multi-core decoder 31-c comprises four first-level groups each comprising four LDPC cores 10, a first-level shared memory 20 (“Mem. lvl. 1”) and a first-level scheduling module 21 (“Ord. lvl. 1”) configured to schedule the cores 10 of the first-level group. The multi-core decoder 31-c also comprises a second-level group comprising the four first-level groups, a second-level shared memory 22 (“Mem. lvl. 2”) and a second-level scheduling module 23 (“Ord. lvl. 2”) configured to schedule the first-level scheduling modules 21.

[0168] The concept illustrated in [Fig.22] could be extended to a number of hierarchy levels greater than two. The highest level shared memory 22 is a non-volatile memory that stores the set of Nc predefined configurations. Different groups of configurations can be dynamically loaded into the lower level shared memories 20. The size of a shared memory can be all the larger as its hierarchy level is high. Test results:

[0169] Tests were carried out to compare a conventional 5G decoder produced in FPGA (Xilinx) with an LDPC decoder according to the invention, for the decoding of a code word according to certain particular configurations.

[0170] It has been observed that the coded flow values ​​are significantly more homogeneous with the decoder according to the invention for the different configurations tested (the coded rate can be measured in MLLR / s, one MLLR / s corresponds to a rate of 106 LLR values ​​per second). The ratio between the highest coded rate and the lowest coded rate is of the order of five for the conventional 5G decoder, compared to a ratio of the order of 2.5 with an LDPC decoder according to the invention.

[0171] For high coding rates (e.g. for a configuration {K=5632, R=8 / 9, Z=256}), it has been observed that the memory requirements are about 1.25 times higher and the logic requirements are about 2 times higher for the conventional 5G decoder compared to the LDPC decoder according to the invention. The memory requirements can be measured in BRAM per MLLR / s; a BRAM corresponds to a 36 kbit RAM block. The logic requirements can be measured in LUT6 per MLLR / s; a LUT6 is an FPGA resource corresponding to a six-entry Look-Up Table.

[0172] For low coding rates (e.g. for a configuration {K=5632, R=2 / 5, Z=256}), it has been observed that the memory requirements are approximately 2.0 times higher and the logic requirements are approximately 3.6 times higher for the conventional 5G decoder compared to the LDPC decoder according to the invention.

[0173] It should be noted that the above results were obtained for an optimal parallelization rate for the conventional 5G decoder. With a more heterogeneous distribution of the tested configurations, the decoder according to the invention would have outperformed the conventional decoder even more significantly.

[0174] The above description clearly illustrates that, through its various characteristics and their advantages, the present invention achieves the set objectives. In particular, the decoder according to the invention is compatible with 5G while remaining particularly well suited and very efficient for satellite communications.

[0175] It should be noted that the embodiments considered above have been described as non-limiting examples, and that other variants are consequently conceivable. In particular, the value of the number Nc or ​​the specific examples of LDPC configurations retained are in no way limiting. Other configurations could be selected to form a strict subgroup of the set of 5G LDPC configurations defined by the 3GPP TS 38.212 specification.

[0176] Similarly, different values ​​may be considered for the number of cores or the number of groups of cores forming the LDPC decoder. Also, different methods may be considered for defining the maximum number of iterations as a function of a load rate of the decoder. The different possible choices for these values ​​or these methods constitute only variants of the invention.

[0177] The invention has been described by considering an LDPC decoder intended to be embarked in a satellite in orbit around the Earth. Nothing excludes however, according to other examples, to consider other specific areas in which it might be interesting to have a 5G-compatible LDPC decoder without necessarily supporting all 5G LDPC configurations.

[0178] The present application relates on the one hand to a multi-core architecture for an LDPC decoder, and on the other hand to obtaining a 5G-compatible LDPC decoder (i.e. compatible with at least some 5G LDPC configurations) which remains suitable and efficient for satellite communications. These two aspects can however be applied independently of each other: the proposed multi-core architecture is an invention in its own right, whether or not it is used for a 5G-compatible LDPC decoder; also, obtaining an LDPC decoder compatible with at least some 5G configurations while remaining suitable for satellite communications is an invention in its own right, whether or not it is combined with the proposed multi-core architecture. The proposed multi-core architecture is however particularly well suited for designing a 5G-compatible LDPC decoder suitable for satellite communications.

Claims

Claims

1. Multi-core decoder (31) for decoding low density parity check code words, or LDPC code, said multi-core decoder (31) comprising at least one group of several cores (10), several of these cores (10) being able to simultaneously decode each a different LDPC code word, said at least one group comprising a memory (20) shared between the different cores (10) in which a number Nc of predefined configurations is stored, each predefined configuration being associated with a parity matrix corresponding to an LDPC code, each core (10) being adapted to be dynamically configured with any one of the Nc predefined configurations stored in the shared memory (20) in order to decode an LDPC code word using this configuration.

2. Multi-core decoder (31) according to claim 1 in which the Nc predefined configurations form a subgroup, in the strict sense, of the 5G LDPC configurations defined in the 3GPP TS 38.212 standard.

3. Multi-core decoder (31) according to any one of claims 1 to 2 in which each core (10) comprises a local configuration memory (13) whose size is at least equal to the largest of the sizes of the Nc predefined configurations stored in the shared memory (20), and the size of the local memory (13) is at least four times smaller than the size of the shared memory (20).

4. Multi-core decoder (31) according to any one of claims 1 to 3 in which the number Nc of predefined configurations stored in the shared memory (20) is between 8 and 512, or even between 8 and 256, or even between 8 and 128.

5. Decoder (31) according to any one of claims 1 to 4 in which each of the Nc configurations comprises an optimized ordering of parity check nodes of the core (10), said ordering being defined as a function of a degree of parity of each parity check node for the configuration considered.

6. A multi-core decoder (31) according to any one of claims 1 to 5 wherein the decoding of an LDPC codeword by a core (10) comprises the execution of one or more iterations until a stopping criterion is satisfied, said stopping criterion comprising a check whether a maximum number of iterations is reached and, for a new code word to be decoded by a core (10), the maximum number of iterations is defined according to a load rate of the multi-core decoder (31).

7. A multi-core decoder (31) according to claim 6 wherein the load rate at a given instant is calculated as a function of a number of cores (10) simultaneously occupied in decoding an LDPC code word at said instant.

8. A multi-core decoder (31) according to any one of claims 6 to 7 wherein the maximum number of iterations is also defined as a function of the size of the new code word.

9. A multi-core decoder (31) according to any one of claims 6 to 8 wherein the maximum number of iterations is also defined as a function of a coding rate used for the new codeword.

10. A multi-core decoder (31) according to any one of claims 6 to 9 wherein the maximum number of iterations is also defined as a function of an estimated signal-to-noise ratio for the new codeword.

11. Multi-core decoder (31) according to any one of claims 1 to 10, comprising a scheduling module (21) adapted to schedule the cores (10) by favoring, for the decoding of a new code word, the selection of a core (10) for which the last configuration used by the core (10) for the decoding of a previous code word is the same as the configuration to be used for the decoding of the new code word.

12. Receiving device comprising a multi-core decoder (31) according to any one of claims 1 to 11.

13. Satellite comprising a receiving device according to claim 12.

Citation Information

Patent Citations

  • Load Balanced Decoder Systems And Methods

    US20220182076A1