A GPU-based fine-grained polar code decoding system and method
By adopting a GPU-based fine-grained polarized coding system in the polarized coding system, using the joint storage matrix M and fine-grained message propagation scheme, the problem of insufficient flexibility and throughput in the prior art is solved, and higher flexibility and throughput are achieved, which is suitable for the needs of software-defined networks.
Patent Information
- Application Number
- CN202211143479.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-09-20
AI Technical Summary
Existing polarized codecs based on dedicated hardware have shortcomings in flexibility and throughput, making them difficult to adapt to the needs of software-defined networks.
The GPU-based fine-grained polarized coding decoding system is adopted to improve the flexibility and throughput of the system through the combined storage matrix M and the fine-grained message propagation scheme.
It achieves higher flexibility and throughput, can effectively support the needs of software-defined networks, and improves hardware utilization.
Smart Images

Figure CN115567159B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of digital signal processing, and particularly to a fine-grained polar code decoding system and method based on GPU. Background Art
[0002] An error correcting code is a code that can be self-detected or corrected at the receiving end after an error occurs during transmission. It is mainly used to control errors when transmitting data in a channel with strong noise interference, and is also called error control coding. The error correcting code aims to overcome the damage of noise to the signal in the channel, so its encoding process is also called channel coding. The error correcting code improves the transmission reliability at the cost of reducing the transmission efficiency. The more redundancy is introduced, the stronger the error correction and detection capabilities are.
[0003] The polar code is currently the only error correcting code that can be proven to reach the Shannon limit. The polar code decoding performance is excellent and is adopted by the control channel of the fifth generation mobile communication (5G). There are two algorithms for polar code decoding. The first is the successive cancellation algorithm, which was proposed by Professor Arikan in 2009 (E. Arikan, "Channel Polarization: A Method for Constructing Capacity-Achieving Codes for Symmetric Binary-Input Memoryless Channels," in IEEE Transactions on Information Theory, vol. 55, no. 7, pp. 3051-3073, July 2009, doi: 10.1109 / TIT.2009.2021379.). However, due to its inherent sequential processing characteristics, the successive cancellation decoding algorithm has a low throughput.
[0004] To improve the throughput, the more common is the belief propagation algorithm, which was proposed by Gallager in 1962 (R. Gallager, "Low-density parity-check codes," in IRE Transactions on Information Theory, vol. 8, no. 1, pp. 21-28, January 1962, doi: 10.1109 / TIT.1962.1057683.). The belief propagation algorithm is based on message passing between the check nodes and variable nodes on the left and right sides of the factor graph. For example, for an (n, k) polar code, n = 2 mLet \(n\) denote the codeword length, \(k\) denote the number of information bits, and \(m\) be the number of stages in the factor graph. The iterative decoding process is a round-trip process of updating messages on the factor graph. The belief propagation algorithm can be parallelized, and decoders based on this algorithm usually have higher throughput.
[0005] Emerging communication paradigms, such as network function virtualization and software-defined networks, not only require communication modules to have high speed, but also need communication modules to have flexibility, scalability, and configurability. The prior art provides a decoder for polar codes implemented based on dedicated hardware (Syed Mohsin Abbas, YouZhe Fan, Ji Chen, and ChiYing Tsui, “High-throughput and energy-efficient belief propagation polar code decoder,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 25, no. 3, pp. 1098–1111, 2017.). Although the channel decoder module of this decoder has a very high throughput, since it is designed for a specific standard, it is very difficult to adapt to various modes. Moreover, the functions of the dedicated chip can hardly be changed, so its flexibility is poor and it cannot well meet the requirements for flexibility in software-defined networks. The prior art also provides a polar code decoding system based on GPU (graphics processing unit) (Z. Liu, R. Liu, Z. Yan and L. Zhao, "GPU-based Implementation of Belief Propagation Decoding for Polar Codes," ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019, pp. 1513 - 1517, doi: 10.1109 / ICASSP.2019.8683248.). The GPU architecture is adopted to improve flexibility, but the throughput is lacking. Summary of the Invention
[0006] This application provides a GPU-based fine-grained polar code decoding system and method to improve the flexibility and throughput of the decoding system.
[0007] In a first aspect, the present application provides a fine-grained polar code decoding system based on a GPU. The fine-grained polar code decoding system includes a host side and a GPU device side. The GPU device side includes a plurality of stream processors and a global memory connected to the host side. The global memory is used to receive the initialization message transmitted by the host side. The stream multiprocessors are used to process the iterative decoding of polar codes according to the initialization message. It is characterized in that the GPU device side further includes:
[0008] A combined storage matrix M, which is used to receive and store the initialization message of a single codeword transmitted by the global memory, and store the update message from the source end to the channel end and the update message from the channel end to the source end during the iterative decoding of the single codeword. The size of the combined storage matrix M is configured as n(m + 1), where n = 2 m is the codeword length, and m is the number of stages in the factor graph;
[0009] wherein, the stream multiprocessor is further configured as:
[0010] In each iterative decoding, update the messages of the first m - 1 stages from the source end to the channel end and the messages of m stages from the channel end to the source end.
[0011] In one implementation, the stream multiprocessor includes a plurality of stream processors. The stream multiprocessor is used to execute thread blocks, and the stream processor is used to execute threads;
[0012] Each stream multiprocessor accommodates at least one thread block. The number of threads in all thread blocks is the same. One thread block is used to complete the iterative decoding of one codeword, and one thread is used to complete the calculation of one processing unit in the factor graph. Different threads in each stage of the factor graph are processed in parallel. The number of threads in each thread block is set to n / 2;
[0013] The host side includes a configuration unit, which is used to configure the number of thread blocks in each stream multiprocessor and the storage address of the combined storage matrix M when the kernel of the GPU device side is called.
[0014] In one implementation, the stream multiprocessor further includes a shared memory. The configuration unit is further configured as:
[0015] Obtain the size S of the shared memory sm and the storage space size n(m + 1)S required by the combined storage matrix M e , where S e is the space required for each storage unit in the combined storage matrix M;
[0016] If S sm ≥n(m + 1)S e, the number of thread blocks in each streaming multiprocessor is set to and the combined storage matrix M is stored in the shared memory;
[0017] If S sm <n(m + 1)S e , the number of thread blocks in each streaming multiprocessor is set to 1, and the combined storage matrix M is stored in the global memory.
[0018] In one implementation, the global memory is used to transfer the decoding result to the host side. Among them, when the combined storage matrix M is stored in the shared memory, the global memory is used to receive the decoding result transmitted from the shared memory after decoding stops.
[0019] In one implementation, the GPU device side further includes a constant memory and a texture memory.
[0020] In one implementation, multiple kernels are started simultaneously by multiple streaming processors of the fine-grained polar code decoding system.
[0021] In one implementation, the data transfer between the host side and the GPU device side adopts an asynchronous data transfer method to overlap kernel execution and data transfer.
[0022] In a second aspect, the present application further provides a fine-grained polar code decoding method based on GPU, including:
[0023] Perform message initialization on the host side and transfer the initialization message to the GPU device side. Among them, the initialization message includes the first column message from the source end to the channel end and the last column message from the channel end to the source end in the factor graph;
[0024] Call the kernel on the GPU device side, configure the number of thread blocks in each streaming multiprocessor and the storage address of the combined storage matrix M, and store the initialization message in the combined storage matrix M;
[0025] Perform iterative decoding on the GPU side. During each iterative decoding process, read the initialization message or the message updated in the previous stage from the combined storage matrix M to update the messages in the first m - 1 stages from the source end to the channel end and the messages in m stages from the channel end to the source end. After each stage ends, synchronization is performed inside the thread block. Among them, the updated message is stored in the combined storage matrix M;
[0026] After decoding ends, write the decoding result into the global memory and then transfer it from the global memory to the host side.
[0027] In one implementation, it includes: performing iterative decoding on the GPU side. During each iterative decoding process, the initialization message or the message updated in the previous stage is read from the joint storage matrix M to update the messages of the first m - 1 stages from the source end to the channel end and the messages of m stages from the channel end to the source end. After each stage ends, synchronization is performed inside the thread block, including:
[0028] The iterative decoding process is a round-trip process from the source end to the channel end on the factor graph. First, the R i,j messages are updated stage by stage from stage 1 to m - 1, and then the L i,j messages are updated stage by stage from stage m to 1. The criteria for message update are as follows:
[0029]
[0030]
[0031]
[0032] L i,j+n / 2 = g(R i,j , L i+1,2j-1 ) + L i+1,2j ;
[0033] where R i,j is the message from the source end to the channel end, L i,j is the message from the channel end to the source end, i and j represent positions or coordinates in the factor graph, i represents the column number, j represents the row number, 0 < i ≤ m, 0 < j ≤ n, and g(a, b) = 2tanh -1 (tanh(a / 2)tanh(b / 2));
[0034] At the end of each iteration, the estimated value of the source vector is updated using the following model:
[0035]
[0036] until the decoding reaches the preset maximum number of iterations or meets the early termination criteria.
[0037] In one implementation, it further includes:
[0038] If the joint storage matrix M is stored in the shared memory, the initial value message is read from the global memory into the shared memory, and during the decoding process, information is read and written from the shared memory.
[0039] As can be seen from the above technical solutions, the present application provides a fine-grained polar code decoding system and method based on a GPU, which adopts two fine-grained strategies to improve the system throughput. First, fine-grained storage is adopted. The joint storage matrix M is used to store the initialization messages of a single codeword, as well as the update messages from the source end to the channel end and from the channel end to the source end during the iterative decoding process of a single codeword. The size of the joint storage matrix M is configured as n(m + 1), where n = 2 m is the codeword length, and m is the number of stages in the factor graph. The storage consumption of each codeword is very low, enabling the system to load more polar codewords.
[0040] Second, a fine-grained message propagation scheme is adopted to reduce the number of decoding steps. In each iterative decoding, only the messages of the first m - 1 stages from the source end to the channel end and the messages of m stages from the channel end to the source end are updated, thereby reducing the decoding time of a single polar code. It not only saves the storage overhead of a single codeword but also reduces the number of iterative steps, thereby improving the system throughput. Thanks to the fine-grained strategy, this system can make full use of the on-chip shared memory to accelerate data access.
[0041] In addition, the internal architecture design of the kernel is reasonably mapped to maximize the use of shared memory and improve the speed. The architecture design between kernels adopts multi-stream and asynchronous data transmission to improve the throughput. By adopting an efficient architecture design and software-hardware mapping strategy, the hardware utilization rate can be maximized to achieve a better throughput. Brief Description of the Drawings
[0042] Figure 1 is a schematic diagram of an (8, 4) polar code factor graph and a processing unit provided by an embodiment of the present application;
[0043] Figure 2 is a flowchart of polar code belief propagation decoding provided by an embodiment of the present application;
[0044] Figure 3 is a schematic diagram of message storage and update provided by the prior art;
[0045] Figure 4 is a schematic diagram of a message storage and update provided by an embodiment of the present application;
[0046] Figure 5 is a schematic structural diagram of a fine-grained polar code decoding system provided by an embodiment of the present application;
[0047] Figure 6 is a schematic diagram of multi-stream and asynchronous data transmission between kernels provided by an embodiment of the present application;
[0048] Figure 7It is a flowchart of a fine-grained polar code decoding method provided by an embodiment of the present application. Detailed implementation manners
[0049] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0050] In the present application, the character " / " generally indicates an "or" relationship between the associated objects before and after. For example, A / B can be understood as A or B.
[0051] The terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise specified, the meaning of "a plurality" is two or more.
[0052] In addition, the terms "including" and "having" and any variations thereof mentioned in the description of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the listed steps or modules, but optionally further includes other unlisted steps or modules, or optionally further includes other steps or modules inherent to these processes, methods, products, or devices.
[0053] In addition, in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, the use of words such as "exemplary" or "for example" is intended to present concepts in a specific manner.
[0054] The polar code is currently the only error-correcting code that can be proven to reach the Shannon limit, and the polar code is adopted by the control channel of the fifth-generation mobile communication (5G). The belief propagation algorithm is a decoding algorithm for polar codes. The belief propagation algorithm has inherent parallelism and is suitable for high-speed scenarios. Currently, much work focuses on designing dedicated hardware to accelerate the decoding process of polar codes. However, the hardware solution lacks flexibility. Emerging communication paradigms, such as software-defined networks, not only require the communication module to have high speed but also require the communication module to have high flexibility.
[0055] The GPU (Graphics Processing Unit) is a parallel processing device that can be connected to a general-purpose computer. It has a large number of parallel computing cores, usually with independent physical memory and a high-speed memory reading mechanism, which can increase the addition and multiplication operation speeds of a large amount of data by dozens or even hundreds of times on a general-purpose computer. In a GPU, software is executed sequentially in units of program cores, and each program core processes data in parallel. The parallel scale of each program core can be set independently.
[0056] To meet the requirements of modern communication systems for throughput and flexibility, the embodiments of this application adopt the belief propagation algorithm and provide a fine-grained polar code decoding system and method based on a GPU. It adopts a fine-grained storage scheme with very low storage consumption for each codeword, enabling the system to load more polar codewords, and introduces a fine-grained message propagation scheme to reduce the number of decoding steps. In addition, the fine-grained polar code decoding system of the embodiments of this application makes full use of on-chip shared memory to accelerate data access. In terms of architecture design, multi-stream and asynchronous data transmission are also adopted to improve throughput. In terms of flexibility, the fine-grained polar code decoding system can be efficiently configured through software to support different code lengths and code rates. The advantage in flexibility enables the fine-grained polar code decoding system to quickly adapt to the evolution of communication standards.
[0057] The following will elaborate in detail on the basis of polar codes, the decoding algorithm, and the storage method of messages between the source end and the channel end in the embodiments of this application.
[0058] For an (n, k) polar code, n = 2 m represents the codeword length, and k represents the number of information bits. Polar codes are based on the concept of channel polarization, that is, some channels are more reliable than others. According to channel reliability, the positions of the source vector u = [u 1 , u 2 ,..., u n can be divided into two groups, namely the information set I and the frozen set F. The information set I consists of the k positions of u on the first k most reliable channels, and the other n - k positions belong to the frozen set F. Information bits are only placed in I, and constant zeros are assigned in F. The source vector u is encoded into the codeword x by multiplying with the generator matrix G, that is, x = uG. In addition, polar codes can be represented by a factor graph. See Figure 1 , Figure 1 where (a) in Figure 1 shows the factor graph of an (8, 4) polar code.
[0059] The belief propagation decoding algorithm performs message passing according to the factor graph. As shown in (a) in Figure 1 , the factor graph is divided into m stages by dashed lines, and each stage contains n / 2 processing units. In Figure 1In (a), the processing units are marked with dashed boxes, and each processing unit is connected to four nodes (the nodes are drawn as solid black circles in Figure 1 ). Figure 1 In (b), the left and right nodes of a processing unit are shown, d i,j , d i,j+N / 2 , d i+1,2j-1 , d i+1,2j , and each node is associated with two one-way log-likelihood ratio messages. Taking the node d i,j as an example, the message R i,j from the source end to the channel end and the message L i,j from the channel end to the source end both pass through the node d i,j . Among them, i and j represent the positions or coordinates in the factor graph, i represents the column number, j represents the row number, 0 < i ≤ m and 0 < j ≤ n.
[0060] The codeword x is modulated by binary phase shift keying and then transmitted through an additive white Gaussian noise channel. The noise of the additive white Gaussian noise channel is a Gaussian random variable with a mean of zero and a variance of σ. Let y = [y 1 , y 2 ,..., y n represent the channel output. The decoding process S101 to S301 is described below in combination with Figure 2 .
[0061] S101. The initialization of decoding has two tasks. Task 1: L m+1,j is set to 2y j / σ 2 , 0 < j ≤ n. Task 2: For j ∈ I, R 1,j is set to zero, otherwise R 1,j is set to infinity. That is, the initialization messages of decoding include the first column messages (R 1,j ) from the source end to the channel end and the last column messages (L m+1,j ) from the channel end to the source end in the factor graph.
[0062] S102. The iterative process of decoding is a round-trip process on the factor graph, that is, updating the messages from left (source end) to right (channel end), and then from right to left. The criteria for message update are as follows:
[0063]
[0064]
[0065]
[0066] L i,j+n / 2 = g(R i,j , L i+1,2j-1)+L i+1,2j ;
[0067] where g(a, b) = 2tanh -1 (tanh(a / 2)tanh(b / 2)), 0 < i ≤ m, and 0 < j ≤ n.
[0068] According to the above message update criterion, during the iterative update process, the message R m+1,j (0 < j ≤ n) will not be used. Therefore, the embodiment of the present application utilizes the above regularized iterative process, adopts a fine-grained update mechanism in the iteration, and in the update from left to right, only updates the R i,j message stage by stage from stage 1 to m - 1, that is, first updates the R i,j message stage by stage from stage 1 to m - 1, and then updates the L i,j message stage by stage from stage m to 1.
[0069] It should be noted that when updating the message from left to right, for the message L from the channel end to the source end that needs to be used but has not been initialized i,j is defaulted to 0. For example, for the (8, 4) polarization code in Figure 1 , when starting to update from left to right at the beginning of the iteration, the processing units connected to the nodes d 1,1 , d 1,5 , d 2,1 , d 2,2 will perform the following update: R 2,1 = g(R 1,1 , L 2,2 + R 1,5 ), R 2,2 = g(R 1,1 , L 2,1 + R 1,5 ), where L 2,j is defaulted to 0.
[0070] S103. Update the estimated value of the source vector at the end of each iteration Adopt the following model:
[0071]
[0072] Stop when the decoding reaches the preset maximum number of iterations or meets the early termination criterion.
[0073] See Figure 3 , in the prior art, all the messages R i,j from left to right of the factor graph are stored using the matrix R from left to right,
[0074] where R i,j is located in the i-th column and the j-th row, and the size of the matrix R is n × (m + 1). For example, for the n = 2 polarization code:
[0075]
[0076] Figure 3 For the (8, 4) polar code, the size of matrix R is 8×4. Similarly, all the right-to-left messages L in the factor graph i,j can be stored using a right-to-left matrix L. The prior art first updates R i,j messages stage by stage from stage 1 to m, and then updates L i,j messages stage by stage from stage m to 1. For example, in stage 1 from left to right, according to the first column of R and the second column of L, the second column of R is obtained through the processing unit and stored in the second column of matrix R; in stage 2 from left to right, according to the second column of R and the third column of L, the third column of R is obtained through the processing unit and stored in the third column of matrix R; in stage 3 from left to right, according to the third column of R and the fourth column of L, the fourth column of R is obtained through the processing unit and stored in the fourth column of matrix R.
[0077] A GPU-based fine-grained polar code decoding system provided by an embodiment of the present application adopts a fine-grained storage strategy and a fine-grained update mechanism. The fine-grained update mechanism is as described in S102 above. In the left-to-right update, only the R i,j messages are updated stage by stage from stage 1 to m-1. The fine-grained storage strategy adopted in the fine-grained polar code decoding system provided by the embodiment of the present application will be described below.
[0078] In the storage method of matrix R and L provided by the prior art, the R matrix has n(m + 1) elements, and the L matrix has n(m + 1) elements, requiring a storage space of 2n(m + 1) elements. According to the above message update criterion, in the i-th stage of left-to-right propagation, for matrix R, the read address is the i-th column, and the write address is the (i + 1)-th column. For matrix L, the read address is the (i + 1)-th column, and the write address is empty. Therefore, the storage method in the prior art has redundancy. The fine-grained polar code decoding system provided by the embodiment of the present application adopts a fine-grained storage strategy in storage, jointly storing matrix R and matrix L in a matrix, that is, the joint storage matrix M. The size of the joint storage matrix M is configured as n(m + 1). Adopting the fine-grained storage strategy only requires n(m + 1) storage units to store the messages of a single codeword.
[0079] During the update process, the space occupied by messages that will no longer be used can be recycled. Refer to Figure 4 , the joint storage matrix M, which is used to receive and store the initialization messages of a single codeword, and store the update messages from the source end to the channel end and the update messages from the channel end to the source end during the iterative decoding process of a single codeword. Exemplarily, taking the left-to-right update in the factor graph as an example, for the (8, 4) polar code, only stages 1 and 2 are updated from left to right.Figure 4 The left side of the vertical bar stores R i,j message, and the right side stores L i,j message. Before the stage 1 update, R 1,j is stored in the first column of the combined storage matrix M, and L 2,j and L 3,j and L 4,j are stored in the second, third, and fourth columns of the combined storage matrix M (L 2,j and L 3,j are all 0 at this time if there is no initialization). During the stage 1 update, the message R 1,j in the first column and the message L 2,j in the second column are read from the combined storage matrix M. After being processed by the processing unit, R 2,j is obtained and stored in the second column of the combined storage matrix M, overwriting L 2,j . The details of stage 2 from left to right and stages 3 to 1 from right to left are not elaborated here. The specific storage locations and methods using the combined storage matrix M are only an example of this application. Undoubtedly, a combined storage matrix M of size n(m + 1) can be used to store the messages during the update process at other storage locations.
[0080] The following will elaborate in detail on the GPU-based fine-grained polar code decoding system architecture provided by the embodiments of this application.
[0081] Refer to Figure 5 , the fine-grained polar code decoding system includes a host side (such as a CPU) and a GPU device side. The GPU device side includes multiple stream processors and a low-speed global memory, constant memory, and texture memory connected to the host side. The global memory is used to receive the initialization messages transmitted from the host side. The stream multiprocessors are used to perform iterative decoding of the polar code based on the initialization messages. In each iterative decoding, the messages in the first m - 1 stages from the source end to the channel end and the messages in m stages from the channel end to the source end are updated.
[0082] The GPU device side also includes a combined storage matrix M. The combined storage matrix M is used to receive and store the initialization messages of a single codeword transmitted by the global memory, and to store the update messages of a single codeword from the source end to the channel end and from the channel end to the source end during the iterative decoding process. The size of the combined storage matrix M is configured as n(m + 1), where n = 2 m is the codeword length, and m is the number of stages in the factor graph.
[0083] Each stream multiprocessor contains multiple stream processors and a high-speed shared memory. Although the shared memory is fast, its capacity is very limited. Denote the number of stream multiprocessors in the GPU as N sm , and denote the size of the shared memory of each stream multiprocessor as S sm。At the software level, the kernel function of the GPU usually generates a large number of threads to fully utilize data parallelism. All the threads generated during the kernel call are collectively referred to as a grid. Each grid consists of one or more thread blocks, and all the thread blocks in the grid have the same number of threads. Hardware and software are closely related. Thread blocks are executed by streaming multiprocessors, and threads are executed by stream processors. In addition, one or more thread blocks can be accommodated in each streaming multiprocessor.
[0084] Reasonable architecture design can maximize hardware utilization and thus improve the throughput of the system. The following will elaborate on the fine-grained polar code decoding system provided by the embodiments of the present application in detail from the internal and kernel levels of the kernel.
[0085] First, introduce the architecture design and software-hardware mapping inside the kernel. For a single polar code, the message propagation between stages is serial, while the messages inside each stage are updated in parallel. The fine-grained polar code decoding system provided by the embodiments of the present application configures one thread to complete the calculation of one processing unit, and one stage has n / 2 processing units. Therefore, the number of threads N in each thread block t / b is set to n / 2.
[0086] During the iterative decoding process, the messages from left to right and from right to left are frequently accessed. To shorten the memory access time, if the capacity of the shared memory is sufficient, the combined storage matrix M is stored in the shared memory (as Figure 5 shown), and the information is read and written from the combined storage matrix M in the shared memory each time; if the capacity of the shared memory is insufficient, the combined storage matrix M is stored in the global memory ( Figure 5 not shown).
[0087] Therefore, the host side of the embodiments of the present application further includes a configuration unit, which is used to configure the number of thread blocks in each streaming multiprocessor and the storage address of the combined storage matrix M when the kernel is called on the GPU device side. The fine-grained polar code decoding system provided by the embodiments of the present application adopts a fine-grained storage scheme for the combined storage matrix M. The messages from left to right and from right to left of a codeword are stored in the M matrix, and n(m + 1) storage units are required. Assuming that the space required for each storage unit is S e , then the messages from left to right and from right to left of a single codeword occupy n(m + 1)S e in memory size.
[0088] The configuration unit is further configured to: obtain the size S of the shared memory sm and the storage space size n(m + 1)S required for the combined storage matrix M e ; if S sm ≥n(m + 1)Se , in order to maximize the utilization of shared memory, the number of thread blocks N in each streaming multiprocessor b / s is set to and the combined storage matrix M is stored in the shared memory; if S sm <n(m + 1)S e , then the number of thread blocks N in each streaming multiprocessor b / s is set to 1, and the combined storage matrix M is stored in the global memory. The number of thread blocks N in each grid b / g is then N sm ×N b / s .
[0089] Exemplarily, for the GPU-based fine-grained polar code decoding system provided in the embodiments of the present application, NVIDIA Tesla T4 GPU is selected. This GPU has a total of 40 streaming multiprocessors, that is, N sm = 40. The shared memory of each multi-stream processor is limited and configured to 32KB, that is, S sm = 32KB. The data is quantized to 8 bits, and the space required for a single storage unit is S e = 1B. Decoding a (512, 256) polar code, that is, n = 512, k = 256. According to the above mapping strategy, the number of threads N in each thread block t / b is set to n / 2 = 256. The messages of a single codeword from left to right and from right to left are stored in the combined storage matrix M, which requires n(m + 1)S e = 5K of memory, so the combined storage matrix M is stored in the shared memory of the streaming multiprocessor, and the number of thread blocks N in each streaming multiprocessor b / s is The number of thread blocks N in each grid b / g is set to N sm ×N b / s = 240.
[0090] Secondly, the system architecture design at the kernel level is introduced. Since the early termination is adopted in the decoding process, the number of iterations of each polar code is different. Therefore, the execution time of the thread blocks in the kernel is different, resulting in an unbalanced workload among different streaming multiprocessors and reducing the resource utilization rate. To balance the workload, the fine-grained polar code decoding system provided in the embodiments of the present application starts multiple kernels with multiple streams simultaneously. In addition, the data transfer time between the host CPU and the GPU cannot be ignored. To solve this problem, an asynchronous data transfer method is adopted to overlap the kernel execution and data transfer. Figure 6 A schematic diagram showing the use of multiple streams and asynchronous data transfer. Figure 6In this context, C2G and G2C respectively represent data transfer from the CPU to the GPU and from the GPU to the CPU. Exemplarily, three streams are shown, and data transfer and decoding between streams can be overlapped, which helps improve the overall system speed.
[0091] To more clearly understand the working process of the GPU-based fine-grained polar code decoding system provided by the embodiments of this application, the embodiments of this application also provide a GPU-based fine-grained polar code decoding method, which is executed using the GPU-based fine-grained polar code decoding system. Refer to Figure 7 This method includes S1 to S4.
[0092] S1. Initialize messages at the host side and transfer the initialization messages to the global memory on the GPU device side. Among them, the initialization messages include the first column of messages (R 1,j ) from the source end to the channel end and the last column of messages (L m+1,j ) from the channel end to the source end in the factor graph.
[0093] S2. Invoke the kernel on the GPU device side, configure the number of thread blocks in each streaming multiprocessor and the storage address of the combined storage matrix M, and store the initialization messages in the combined storage matrix M. How to configure has been elaborated in detail in the fine-grained polar code decoding system and will not be repeated here.
[0094] If the combined storage matrix M is stored in the shared memory, it is necessary to read the initialization messages from the global memory into the shared memory in advance.
[0095] S3. Perform iterative decoding on the GPU side. During each iterative decoding process, read the initialization messages or the messages updated in the previous stage from the combined storage matrix M to update the messages in the first m - 1 stages from the source end to the channel end and the messages in the m stages from the channel end to the source end. After each stage ends, synchronization is performed inside the thread block. Among them, the updated messages are stored in the combined storage matrix M.
[0096] The iterative decoding process is a round-trip process from the source end to the channel end on the factor graph. First, update the R i,j messages stage by stage from stage 1 to m - 1, and then update the L i,j messages stage by stage from stage m to 1. The criteria for message update are as follows:
[0097]
[0098]
[0099]
[0100] L i,j+n / 2 = g(Ri,j , L i+1,2j-1 ) + L i+1,2j ;
[0101] Among them, R i,j is the message from the source end to the channel end, and L i,j is the message from the channel end to the source end. i and j represent positions or coordinates in the factor graph. i represents the column number, and j represents the row number. 0 < i ≤ m, 0 < j ≤ n, and g(a, b) = 2tanh -1 (tanh(a / 2)tanh(b / 2));
[0102] Update the estimated value of the source vector at the end of each iteration using the following model:
[0103]
[0104] Until the decoding reaches the preset maximum number of iterations or meets the early termination criterion.
[0105] During the above decoding process, if the joint storage matrix M is stored in the shared memory, all information is read and written from the shared memory.
[0106] S4. After the decoding is completed, write the decoding result into the global memory and then transfer it to the host from the global memory.
[0107] The embodiment of the present application provides a GPU-based fine-grained polar code decoding system and method, which adopts two fine-grained strategies to improve the system throughput. First, adopt fine-grained storage. Use the joint storage matrix M to store the initialization messages of a single codeword, as well as the update messages from the source end to the channel end and from the channel end to the source end during the iterative decoding process of a single codeword. The size of the joint storage matrix M is configured as n(m + 1), where n = 2 m is the codeword length, and m is the number of stages in the factor graph. The storage consumption of each codeword is very low, enabling the system to load more polar code codewords.
[0108] Second, adopt a fine-grained message propagation scheme to reduce the number of decoding steps. In each iterative decoding, only update the messages of the first m - 1 stages from the source end to the channel end and the messages of m stages from the channel end to the source end, thereby reducing the decoding time of a single polar code. It not only saves the storage overhead of a single codeword but also reduces the number of iterative steps, thereby improving the system throughput. Thanks to the fine-grained strategy, this system can make full use of the on-chip shared memory to accelerate data access.
[0109] In addition, the internal architecture design of the kernel maps reasonably to maximize the utilization of shared memory and improve speed. The architecture design between kernels uses multi-stream and asynchronous data transfer to increase throughput. By adopting an efficient architecture design and software-hardware mapping strategy, the hardware utilization rate can be maximized, achieving a good throughput rate.
[0110] For the same or similar parts among the various embodiments of this specification, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the description in the method embodiment section.
[0111] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A GPU-based fine-grained polar code decoding system. The fine-grained polar code decoding system includes a host side and a GPU device side. The GPU device side includes multiple streaming multiprocessors and a global memory connected to the host side. The global memory is used to receive the initialization message transmitted from the host side. The streaming multiprocessor is used to process the iterative decoding of the polar code according to the initialization message. Characterized in that, The GPU device side further includes: A joint storage matrix M is used to receive and store the initialization message of a single codeword transmitted by the global memory, and store the update message from the source end to the channel end and the update message from the channel end to the source end during the iterative decoding process of the single codeword. The size of the joint storage matrix M is configured as n(m + 1), where n = 2 m is the codeword length, and m is the number of stages in the factor graph; Wherein, the streaming multiprocessor is further configured to: In each iterative decoding, update the messages of the first m-1 stages from the source end to the channel end and the messages of the m stages from the channel end to the source end.
2. A GPU-based fine-grained polar code decoding system according to claim 1, Characterized in that, The streaming multiprocessor includes multiple stream processors. The streaming multiprocessor is used to execute thread blocks, and the stream processor is used to execute threads; Each streaming multiprocessor accommodates at least one thread block. The number of threads in all thread blocks is the same. One thread block is used to complete the iterative decoding of one codeword, and one thread is used to complete the calculation of one processing unit in the factor graph. Different threads in each stage of the factor graph are processed in parallel. The number of threads in each thread block is set to n / 2; The host side includes a configuration unit. The configuration unit is used to configure the number of thread blocks in each streaming multiprocessor and the storage address of the combined storage matrix M when the kernel is called on the GPU device side.
3. A GPU-based fine-grained polar code decoding system according to claim 2, Characterized in that, The streaming multiprocessor further includes a shared memory. The configuration unit is further configured to: Obtain the shared memory size S sm and the storage space size n(m + 1)S required by the combined storage matrix M e , where S e is the space required for each storage unit in the combined storage matrix M; If S sm ≥ n(m + 1)S e , then the number of thread blocks in each streaming multi - processor is set to and the combined storage matrix M is stored in the shared memory; If S sm <n(m + 1)S e , then the number of thread blocks in each streaming multi-processor is set to 1, and the combined storage matrix M is stored in the global memory.
4. A GPU-based fine-grained polar code decoding system according to claim 3, Characterized in that, The global memory is used to transmit the decoding result to the host side. Wherein, when the combined storage matrix M is stored in the shared memory, the global memory is used to receive the decoding result transmitted from the shared memory after the decoding stops.
5. A GPU-based fine-grained polar code decoding system according to claim 1, Characterized in that, The GPU device side further includes a constant memory and a texture memory.
6. A GPU-based fine-grained polar code decoding system according to claim 1, Characterized in that, Multiple stream processors of the fine-grained polar code decoding system start multiple kernels simultaneously.
7. A GPU-based fine-grained polar code decoding system according to claim 1, Characterized in that, The data transmission between the host side and the GPU device side adopts an asynchronous data transmission method to overlap the kernel execution and data transmission.
8. A GPU-based fine-grained polar code decoding method, Characterized in that, Includes: Perform message initialization on the host side and transmit the initialization message to the GPU device side. Wherein, the initialization message includes the first column message from the source end to the channel end and the last column message from the channel end to the source end in the factor graph. Call the GPU device - side kernel, configure the number of thread blocks in each streaming multiprocessor and the storage address of the combined storage matrix M, and store the initialization message in the combined storage matrix M; the size of the combined storage matrix M is configured as n(m + 1), where n = 2m is the codeword length and m is the number of stages in the factor graph; Perform iterative decoding on the GPU side. During each iterative decoding process, read the initialization message or the message updated in the previous stage from the combined storage matrix M to update the messages in the first m - 1 stages from the source end to the channel end and the messages in m stages from the channel end to the source end. After each stage, synchronize inside the thread block, where the updated message is stored in the combined storage matrix M; After the decoding is completed, write the decoding result into the global memory and then transfer it from the global memory to the host side.
9. A fine - grained polar code decoding method based on GPU according to claim 8, wherein, it includes: The iterative decoding on the GPU side, during each iterative decoding process, read the initialization message or the message updated in the previous stage from the combined storage matrix M to update the messages in the first m - 1 stages from the source end to the channel end and the messages in m stages from the channel end to the source end. After each stage, synchronize inside the thread block, including: The iterative decoding process is a round-trip process from the source end to the channel end on the factor graph. First, the R i,j messages are updated stage by stage from stage 1 to m - 1, and then the L i,j messages are updated stage by stage from stage m to 1. The criteria for message update are as follows: L i,j+n / 2 = g(R i,j , L i+1,2j-1 ) + L i+1,2j ; where R i,j is the message from the source end to the channel end, L i,j is the message from the channel end to the source end, i and j represent positions or coordinates in the factor graph, i represents the column number, j represents the row number, 0 < i ≤ m, 0 < j ≤ n, and g(a, b) = 2tanh -1 (tanh(a / 2)tanh(b / 2)); Update the estimate of the source vector at the end of each iteration Adopt the following model: Until the decoding reaches the preset maximum number of iterations or meets the early termination criterion.
10. A fine - grained polar code decoding method based on GPU according to claim 8, wherein, it includes: If the combined storage matrix M is stored in the shared memory, read the initial value message from the global memory into the shared memory, and during the decoding process, read and write information from the shared memory.