Decoder used for cyclic codes and having cyclic equivalence and decoding method
By introducing the multiplexed design of cyclic equivalence and cyclic displacement operations of Tanner graphs into the Transformer decoder, the problems of high computational complexity and high resource consumption of existing cyclic code decoding methods are solved, and more efficient decoding performance and resource utilization are achieved.
Patent Information
- Application Number
- CN202510277384.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
When facing complex channels and high-dimensional data, the existing cyclic code decoding methods have high computational complexity, high resource consumption, and limited bit error rate optimization. The traditional Transformer decoder fails to fully utilize the intrinsic relationship between the algebraic structure of the cyclic code and the Tanner graph.
By introducing the multiplexed design of cyclic equivalence of Tanner graphs and cyclic displacement operations, the parameter matrix and embedding matrix of the Transformer decoder are optimized, reducing the computational complexity and number of parameters, while maintaining efficient decoding performance.
Significantly reduces the computational complexity and number of parameters of the decoder, improves decoding efficiency, and especially shows superior performance when processing long codewords and complex networks, while maintaining or improving decoding performance.
Smart Images

Figure CN120223102A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of communication and machine learning, and particularly to a decoder and a decoding method for cyclic codes with cyclic equivalence. Background Art
[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Cyclic codes are a class of linear codes with a special algebraic structure. Their basic property is that any codeword of a cyclic code and all its cyclic shifts still belong to the same codeword set. This periodic structure enables cyclic codes to achieve high efficiency in the encoding and decoding processes. This property means that cyclic codes can effectively utilize their structural characteristics during encoding and decoding, resulting in a lower computational complexity.
[0004] In a communication system, the data transmission process generally involves three main stages: transmission, channel transmission, and reception. In the transmission stage, the transmitting end encodes the information into a specific codeword for transmission, usually using error-correcting codes (such as cyclic codes) to ensure the reliability of transmission. The channel transmission process is subject to various interferences, such as noise and signal attenuation, which can cause bit errors during transmission. The receiving end then needs to use a decoding method to recover the original information and correct the errors in the signal.
[0005] In practical applications, the choice of decoding algorithm is crucial for system performance. Common decoding algorithms include Maximum Likelihood Decoding (ML) and Belief Propagation (BP). The goal of these decoding algorithms is to minimize the error between the received signal and the true information as much as possible and recover the most likely codeword. However, the existing decoding methods still have the following problems when facing complex channels and high-dimensional data:
[0006] High computational complexity: Especially when facing long codewords, traditional decoding algorithms (such as BP) often require a large number of computational and iterative steps.
[0007] High resource consumption: Traditional methods usually require high memory and computational capabilities, making it difficult to meet the requirements of large-scale systems and real-time decoding.
[0008] Limited Bit Error Rate (BER) optimization: Although traditional decoding methods perform well under low signal-to-noise ratios, as the signal-to-noise ratio increases, the improvement of the bit error rate is limited.
[0009] Therefore, with the development of information theory and deep learning techniques, more and more research has begun to explore how to use modern deep learning methods to optimize the decoding process, especially the decoder based on the Transformer architecture, which has significant advantages in processing sequence data and long-distance dependencies.
[0010] The Transformer architecture is a deep learning model based on the self-attention mechanism, which initially achieved great success in the field of natural language processing (NLP). Compared with traditional convolutional neural networks (CNNs) and recurrent neural networks (RNNs), the Transformer architecture can establish global dependencies between different positions in the input sequence through the self-attention mechanism, thereby capturing long-distance context information.
[0011] In communication systems, the Transformer decoder has gradually been introduced into encoding and decoding tasks due to its excellent long-distance dependency modeling ability. In particular, methods such as Error Correction Code Transformer (ECCT) have successfully applied the Transformer to channel decoding tasks and achieved good decoding performance. The advantage of the Transformer decoder lies in its ability to flexibly handle complex channel noise and variable coding structures, capturing long-distance dependencies in the input data through the self-attention mechanism, and thus recovering information in complex channel environments.
[0012] Although the Transformer has demonstrated excellent performance in various tasks, when applied to cyclic code decoding, traditional Transformer decoders do not take into account the special algebraic structure of cyclic codes. Existing Transformer decoders usually rely on fixed attention matrices to capture information, but they ignore the periodicity in cyclic codes and the structural information in Tanner graphs. Due to these factors, traditional Transformer decoders often have high computational complexity and a large number of parameters, making them difficult to process efficiently. Summary of the Invention
[0013] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a decoder and decoding method for cyclic codes with cyclic equivalence, a cyclic code decoder based on the Transformer. By introducing the cyclic equivalence of the Tanner graph and making full use of the algebraic structure of cyclic codes, the computational complexity and the number of parameters in the decoder are significantly reduced, while maintaining high decoding performance. This method reduces unnecessary calculations by introducing a reuse design of cyclic shifts during the decoding process and reduces the computational overhead caused by over-embedding through optimized embedding design. The decoder proposed by the present invention not only improves the decoding efficiency but also demonstrates superior performance when dealing with long codewords and complex networks.
[0014] To achieve the above object, the present invention is implemented through the following technical solutions:
[0015] In a first aspect of the present invention, a decoding method for cyclic codes with cyclic equivalence is provided, including the following steps:
[0016] Obtain the cyclic codeword to be decoded and preprocess the codeword;
[0017] Based on the cyclic equivalence of the cyclic code, construct a decoding model, set the initial parameters of the decoding model, initialize the parameter matrix of the decoding model according to the representative rows determined in the Tanner graph, preprocess the input codeword according to the parameter matrix to obtain an embedding matrix, associate the embedding matrix, the parameter matrix with the periodicity of the cyclic code, train the decoding model to obtain an optimized decoding model, wherein, based on the input simulated noise, use the decoding model to perform multiplexing decoding on the parameter matrix and the embedding matrix by means of cyclic shift, and the embedding matrix calculates the parameter matrix based on the self-attention mechanism during the multiplexing decoding process to realize the adjustment of the self-attention matrix;
[0018] Use the optimized decoding model to obtain a preliminary decoding result for the codeword to be decoded;
[0019] The preliminary decoding result is post-processed to obtain the final decoding result.
[0020] Further, the specific steps of preprocessing the input codeword according to the parameter matrix to obtain the embedding matrix include: mapping each bit in the vector to a high-dimensional space by using a fixed-dimension embedding method to obtain the embedding matrix.
[0021] Furthermore, in the fixed-dimension embedding method, the dimension is fixed for each code, and different codes use different embedding dimensions.
[0022] Further, the decoding model is a Transformer network structure, including a self-attention layer and a feed-forward network layer.
[0023] Furthermore, the specific steps of associating the embedding matrix, the parameter matrix with the periodicity of the cyclic code are:
[0024] Extract the representative rows of the embedding matrix and the parameter matrix from the Tanner graph;
[0025] When the embedding matrix and the parameter matrix are needed, generate new matrices by cyclically shifting the representative rows.
[0026] Further, the Tanner graph is the graph structure corresponding to the parity-check matrix of the cyclic code. In the Tanner graph, the cyclic relationships between nodes include two cyclic relationships: strong Tanner cycle equivalence and weak strong Tanner cycle equivalence.
[0027] Further, the specific steps for the embedding matrix to calculate the parameter matrix based on the self-attention mechanism during the multiplexed decoding process are as follows:
[0028] Associate the embedding matrix with the periodicity of the cyclic code;
[0029] Extract the representative rows of the embedding matrix and the mask matrix through the cyclic relationships between the nodes of the Tanner graph. The representative rows of the embedding matrix form a representative matrix, and the representative rows of the mask matrix form a new mask matrix;
[0030] Use the representative matrix and the new mask matrix to perform operations on the attention matrix in the Transformer, and perform a cyclic shift after the calculation is completed to obtain the complete attention matrix.
[0031] The second aspect of the present invention provides a decoder for cyclic codes with cyclic equivalence, including:
[0032] A preprocessing module configured to obtain a cyclic codeword to be decoded and preprocess the codeword;
[0033] A cyclic multiplexing module configured to construct a decoding model based on the cyclic equivalence of the cyclic code, initialize the parameter matrix of the decoding model according to the representative rows determined in the Tanner graph, preprocess the input codeword according to the parameter matrix to obtain an embedding matrix, associate the embedding matrix, the parameter matrix with the periodicity of the cyclic code, train the decoding model to obtain an optimized decoding model. Among them, based on the input simulated noise, use the decoding model to perform multiplexed decoding on the parameter matrix and the embedding matrix in a cyclic shift manner. The embedding matrix calculates the parameter matrix based on the self-attention mechanism during the multiplexed decoding process to realize the construction of the self-attention matrix; use the optimized decoding model to obtain a preliminary decoding result for the codeword to be decoded;
[0034] A postprocessing module configured to perform postprocessing on the preliminary decoding result to obtain the final decoding result.
[0035] The third aspect of the present invention provides a medium on which a program is stored, and when the program is executed by a processor, it implements the steps in the decoding method for cyclic codes with cyclic equivalence as described in the first aspect of the present invention.
[0036] A fourth aspect of the present invention provides a device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, the steps in the decoding method for cyclic codes with cyclic equivalence as described in the first aspect of the present invention are implemented.
[0037] The above one or more technical solutions have the following beneficial effects:
[0038] The present invention discloses a decoder and a decoding method for cyclic codes with cyclic equivalence, aiming to optimize the decoding process of cyclic codes based on Transformer by introducing the algebraic structure of cyclic codes and the cyclic equivalence of Tanner graphs. This solution reduces the computational complexity and storage requirements by reasonably designing the reuse of parameter matrices, the fixed-dimensional design of embedding vectors, and the optimization of the self-attention mechanism, while maintaining the decoding performance.
[0039] The present invention aims to optimize the balance between decoding performance and computational resources by introducing the algebraic structure of cyclic codes, the cyclic equivalence of Tanner graphs, and the self-attention mechanism of Transformer. Through innovative design, the present invention can reduce the computational complexity and the number of parameters while maintaining or improving the decoding performance, thus providing a new solution for efficient and low-resource-consuming decoding.
[0040] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings constituting a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0042] Figure 1 It is the flowchart of Transformer decoding in the first embodiment of the present invention;
[0043] Figure 2 It is the flowchart of fixed-dimensional embedding data processing in the first embodiment of the present invention;
[0044] Figure 3 It is the flowchart of data processing in the decoding process in the first embodiment of the present invention;
[0045] Figure 4 It is the flowchart of self-attention mechanism data processing in the first embodiment of the present invention;
[0046] Figure 5 It is the Hamming code parity-check matrix diagram in the first embodiment of the present invention;
[0047] Figure 6 It is the Tanner graph of the Hamming code check matrix in the first embodiment of the present invention;
[0048] Figure 7 It is the schematic diagram of the cyclic shift of the mask matrix in the first embodiment of the present invention;
[0049] Figure 8 It is the comparison diagram of the computational complexity when decoding BCH in the first embodiment of the present invention;
[0050] Figure 9 It is the comparison diagram of the number of parameters when decoding BCH in the first embodiment of the present invention. Detailed implementation manners
[0051] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0052] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof;
[0053] Embodiment 1:
[0054] The core objective of this embodiment is to solve several key technical problems existing in the existing cyclic code decoder, which are specifically as follows:
[0055] 1. High computational complexity: Traditional cyclic code decoding methods, such as the Belief Propagation (BP) algorithm, have a problem of high computational complexity when dealing with long codewords and complex coding structures. The BP algorithm calculates the marginal probability of each variable node by gradually transmitting information. Although it can provide good decoding effects, when facing large-scale data, especially in the case of low signal-to-noise ratio (SNR), it usually requires a large number of iterative calculations, which leads to a reduction in the efficiency of the decoding process. In some applications, the decoding time is too long, affecting real-time performance.
[0056] In addition, the BP decoding algorithm needs to maintain a large number of intermediate variables and probability information in the context of long codewords, resulting in a large memory consumption, which further increases the demand for computing resources. Therefore, how to reduce the computational complexity, reduce the number of iterations, and significantly improve the decoding efficiency while maintaining a high decoding accuracy has become an urgent technical problem to be solved.
[0057] 2. Excessive number of parameters: Due to its complex structure, the Transformer decoder often requires a large number of parameters to handle long - distance dependencies. When applied to the decoding of cyclic codes, the existing Transformer decoders usually fail to fully consider the algebraic structure of cyclic codes and the internal relationship of the Tanner graph, resulting in an inflation of the number of parameters in the decoder. As the scale of the model increases, the memory consumption and computational complexity during training and inference also increase, which brings huge computational and storage pressures in some practical applications.
[0058] 3. Need to optimize the balance between decoding performance and computational resources: Existing decoding methods usually face a trade - off between performance and computational resources. Although some methods improve decoding performance by increasing the number of parameters and layers, this improvement is usually accompanied by a significant increase in computational resources. Another type of method, although consuming less computational resources, usually has a lower decoding accuracy and cannot meet the requirements of high - reliability communication.
[0059] In view of the above - mentioned defects of the existing technologies, Embodiment 1 of the present invention provides a decoding method for cyclic codes with cyclic equivalence, as Figure 1 shown, including the following steps:
[0060] Step 1: Obtain the cyclic codeword to be decoded and pre - process the codeword.
[0061] The task of the receiving end is to recover the original transmitted codeword x from the received signal y through the decoding process. In a communication system, the received signal y is usually interfered by noise, so the receiving end needs to perform pre - processing and post - processing operations to recover the original information. The pre - processing process is as follows:
[0062] First, the received signal y is converted into a vector where |y| represents the absolute value (amplitude) of the received signal, and s(y) ∈ {0, 1} n-k is a binary mapping calculated from the parity - check matrix H and the signal y. Specifically, s(y) is defined by the following formula:
[0063] s(y)=Hy T .
[0064] where y b ∈ {0, 1} n is the binary mapping of the signal y, and this method is widely used in soft - decision decoding.
[0065] In a specific embodiment, the cyclic code has Tanner cyclic equivalence. In the study of cyclic codes, periodic equivalence is a crucial property. Applying this property to the Tanner graph reveals the deep relationship between variable nodes and check nodes, and these relationships are used to optimize the design of the Transformer decoder. Among them, the Tanner graph is the graph structure corresponding to the parity-check matrix of the cyclic code.
[0066] First, define the relevant node sets in the Tanner graph. Let VNs = {v0, v1, …, v n-1} be the set of variable nodes of the cyclic code C, and CNs = {c0, c1, …, c n-k-1} be the set of check nodes. For any node x, let its set of neighbor nodes be Γ(x). If c ∈ CNs is a check node, its neighbor set is denoted as Then define the relation function:
[0067]
[0068] where f(c) is the relation function of the check node.
[0069] Similarly, for the variable node v ∈ VNs, its neighbor set is Then its relation function is defined as:
[0070]
[0071] where f(v) is the relation function of the variable node.
[0072] On this basis, an equivalence relation is further defined:
[0073] For any check node c i and positive integer j, assume Then define
[0074]
[0075] For any variable node v i and positive integer j, assume Then define
[0076]
[0077] Combined with the Tanner graph by the equivalence relation, it is found that there may be cyclic relationships between different nodes. If there is a cyclic relationship between two nodes, the cyclic relationships are divided into two categories:
[0078] Strong Tanner cycle equivalence: It refers to the existence of exactly the same relationship between two check nodes. That is, in the Tanner graph, if there is a strong cycle equivalence between two check nodes c i and c j , then their corresponding positional relationships are exactly the same structurally and can be cyclically shifted relative to each other through the corresponding positional distance relationship. Specifically, if there is a strong cycle equivalence between check nodes c i and c j , then:
[0079] For any two check nodes c i , c j ∈CNs, and i < j, these two nodes are considered to have strong Tanner cycle equivalence. If and only if these two nodes have strong Tanner cycle equivalence, f(c j ) = f(c i ) + (j - i).
[0080] Similarly, for any two variable nodes v i , v j ∈VNs, and i < j, these two nodes are considered to have strong Tanner cycle equivalence. If and only if these two nodes have strong Tanner cycle equivalence, f(v j ) = f(v i ) + (j - i).
[0081] Where (j - i) is the relative positional distance between the two nodes.
[0082] Weak Tanner cycle equivalence: It means that in the Tanner graph, the relationship between two check nodes is not exactly the same, but there is still a certain cyclic relationship, that is:
[0083] For any two check nodes c i , c j ∈CNs, these two nodes are considered to have weak Tanner cycle equivalence. If and only if these two nodes have weak Tanner cycle equivalence, there exists a positive integer d such that f(c j ) = f(c i ) + d.
[0084] Similarly, for any two variable nodes v i , v j ∈VNs, these two nodes are considered to have weak Tanner cycle equivalence. If and only if these two nodes have weak Tanner cycle equivalence, there exists a positive integer d such that f(v j ) = f(v i ) + d.
[0085] In a specific implementation, the equivalence relationship in the Tanner graph is mapped into the parity-check matrix, thereby obtaining the following Algorithm 1 for finding the cyclic relationship in the parity-check matrix, denoted as Algorithm 1:
[0086] Algorithm 1: Determine the strong Tanner cycle equivalence of variable nodes.
[0087] Input: Parity-check matrix (pc_matrix)
[0088] Output: shift_map and row_indices
[0089] Initialize empty shift_map and row_indices.
[0090] Repeat the following steps:
[0091] Set new_rep = None.
[0092] For each column in the parity-check matrix:
[0093] If the column is not in shift_map or row_indices, set new_rep to the current column and exit the loop.
[0094] Add the column index of new_rep to row_indices.
[0095] For each column column after new_rep in the parity-check matrix:
[0096] If the column is not in shift_map or row_indices, and the cyclically shifted new_rep is equal to the current column, and the shift count shift_count is equal to the value of the column index of the current column column minus the column index of new_rep, then set the column index of shift_map[column] to the column index of new_rep, shift_count.
[0097] Repeat the above steps until new_rep == None.
[0098] Return shift_map and row_indices of the check nodes.
[0099] Among them, row indices is used to record which nodes can be used as representative nodes, that is, cannot be omitted. shift_map records those nodes that can be omitted, the representative nodes required for recovery, and their shift counts.
[0100] This embodiment makes a judgment based on the cyclic relationship. After traversing all nodes, the cyclic relationship in the matrix can be obtained. It is possible to know which nodes are cyclically equivalent based on the cyclic relationship. When performing the attention matrix operation or constructing the parameter matrix, this cyclic relationship is used to reduce the computational complexity and the number of parameters of the entire model.
[0101] Step 2: Construct a decoding model based on the cyclic equivalence of cyclic codes. Set the initial parameters of the decoding model. According to the representative rows determined in the Tanner graph, initialize the parameter matrix of the decoding model. Preprocess the input codeword according to the parameter matrix to obtain an embedding matrix. Associate the embedding matrix, the parameter matrix with the periodicity of the cyclic code, and train the decoding model to obtain an optimized decoding model.
[0102] Step 2.1: Construct a decoding model based on the cyclic equivalence of cyclic codes. In this embodiment, the decoding model is a Transformer network structure, including a self-attention layer and a feed-forward network layer.
[0103] In a specific implementation, the embedded representation, that is, the embedding matrix, enters the Transformer network, which is mainly composed of a self-attention layer and a feed-forward network layer.
[0104] As Figure 1 shown, the basic building blocks of the Transformer mainly include a self-attention layer with a self-attention mechanism and a feed-forward neural network, and also include an embedding layer for embedding data, a normalization layer for normalizing data, and an output processing network for outputting results.
[0105] In this embodiment, a decoding model is constructed by using N Transformer networks. The output of the previous Transformer network is used as the input of the next Transformer network. The model parameters are continuously adjusted through multiple Transformer networks until the decoding model is optimized to meet the preset conditions.
[0106] Similar to a traditional communication system, in the signal preprocessing stage, the Transformer decoder needs to first combine the information into a vector and then perform dimensionality increase on each bit of this vector, that is
[0107]
[0108] where is a trainable embedding matrix, is the Kronecker product, and ⊙ is the Hadamard product.
[0109] The core of the Transformer decoder lies in the self-attention mechanism, and its calculation process can be expressed by the following formula:
[0110]
[0111] Among them, Q, K, and V are the Query, Key, and Value matrices respectively, and d is the dimension of each attention head. Through Multi-Head Attention, the Transformer can parallelly process multiple attention patterns in each layer, improving the model's expressive ability.
[0112] Step 2.2: Set the initial parameters of the decoding model. Initialize the parameter matrix of the decoding model according to the representative rows determined in the Tanner graph, and preprocess the input codeword according to the parameter matrix to obtain the embedding matrix.
[0113] Step 1.2.1: Initialize the required parameter matrix according to the representative rows determined in the Tanner cycle.
[0114] Specifically, initialize the required parameter matrix according to the representative rows determined in the Tanner cycle, and when it needs to be brought in for calculation, restore it through the PCR() algorithm.
[0115] Step 1.2.2: Map each bit in the vector to a high-dimensional space using the fixed-dimensional embedding method to obtain the embedding matrix. Among them, in the fixed-dimensional embedding method, the dimension is fixed for each code, and different codes use different embedding dimensions, which helps to balance the expressive ability and computational complexity.
[0116] In a specific implementation manner, in order to optimize the efficiency of the Transformer decoder, this embodiment introduces a fixed-dimensional embedding method. As Figure 2 shown, in the traditional Transformer decoder, the embedding of the input signal usually uses a trainable high-dimensional embedding space to improve the decoding performance, but the performance improvement brought by unrestrictedly increasing the embedding dimension is limited. In order to obtain a balance between the embedding dimension and the decoding performance, this embodiment adopts a fixed-dimensional embedding design, that is, calculates a specific embedding dimension according to the actual code length and information bits of each cyclic code. At this time, the trade-off between the performance and the embedding dimension is approximately optimal.
[0117] Specifically, for a cyclic code C of length n with information length k, in the embedding stage, consider each preprocessed merged vector of each bit Embed it into a high-dimensional space
[0118] Among them, is the trainable one-hot embedding parameter for each bit, and s represents the code length index.
[0119] Step 2.3: Associate the embedding matrix, the parameter matrix with the periodicity of the cyclic code, train the decoding model, and obtain the optimized decoding model.
[0120] Among them, based on the input simulated noise, the decoding model is used to perform multiplexing decoding on the parameter matrix and the embedding matrix in a cyclic shift manner. During the multiplexing decoding process, the embedding matrix calculates the parameter matrix based on the self-attention mechanism to realize the adjustment of the self-attention matrix.
[0121] Specifically, the cyclic shift method refers to extracting the representative rows of the embedding matrix and the parameter matrix from the Tanner graph; when the embedding matrix and the parameter matrix are needed, the representative rows are cyclically shifted, and then the corresponding recovery operation is performed using the shift_map cyclic spectrum to generate a new matrix.
[0122] Different from the traditional Transformer, as Figure 3 shown, this model makes full use of the cyclic equivalence property of the Tanner graph in the cyclic code, and introduces two key "plug-and-play" cyclic multiplexing strategies in the design, the parameter matrix cyclic multiplexing strategy and the embedding matrix cyclic multiplexing strategy.
[0123] Step 2.3.1: Parameter Cyclic Reuse.
[0124] In the self-attention layer and the feed-forward network, the parameter matrices of each linear transformation are regarded as "relationship matrices" expressing the relationships between code bits. The model only needs to learn the parameter matrices composed of the representative rows, and then construct the complete parameter matrices by cyclically shifting and reusing these matrices, thus greatly reducing the number of parameters.
[0125] Step 2.3.1.1: Associate the relationships within the parameter matrix with the periodicity of the cyclic code.
[0126] Step 2.3.1.2: Determine the representative rows of the parameter matrix through the cyclic relationships between the Tanner graph nodes and initialize them.
[0127] Step 2.3.1.3: When the parameter matrix is needed, generate a temporarily used parameter matrix by cyclically shifting the representative rows.
[0128] In a specific embodiment, based on the cyclic equivalence of the Tanner graph, the present invention proposes a new design of cyclic reuse of parameter matrices, that is, reducing the computational amount through the reuse of parameter matrices. Specifically, the parameter matrices in the decoder are associated with the periodic characteristics of cyclic codes and can be reused through cyclic shift.
[0129] In practical applications, first, representative nodes are extracted from the Tanner graph according to the cyclic relationship by Algorithm 1, and these nodes are mapped to the representative rows in the parameter matrix to form a new representative matrix W'. When the parameter matrix is needed, a new parameter matrix W = PCR(W') is generated by cyclically shifting the representative rows, thereby realizing the reuse of the parameter matrix. Through this design, while maintaining the decoding performance, the required computational resources and storage requirements can be significantly reduced.
[0130] The specific shift algorithm (PCR) is recorded as Algorithm 2 as follows:
[0131] Algorithm 2: Parameter Cyclic Reuse (PCR).
[0132] Input: W' ∈ R row_indices×(2n-k)
[0133] Output: New parameter matrix W
[0134] Initialize W as a (2n - k, 2n - k) zero matrix.
[0135] Fill W in the representative rows:
[0136] Fill W' into the representative rows in W: W[row_indices, :] = W".
[0137] Apply shift_map and cyclic reuse:
[0138] For each target row, independently apply cyclic shift to the variable nodes and check nodes:
[0139] For the variable node part, execute W[t_row, :n - k] = roll(W[r_row, :n - k], shift_count).
[0140] For the check node part, execute W[t_row, n - k:] = roll(W[r_row, n - k:], shift_count).
[0141] Return the new parameter matrix W.
[0142] Step 2.3.2: Embedding Cyclic Reuse.
[0143] During the calculation of self-attention, the decoding model is used to reuse and decode the embedding matrix through circular displacement, adjusting the traditional calculation formula of the attention matrix, reducing the computational complexity, and maintaining effective information transmission.
[0144] Step 2.3.2.1: Associate the embedding matrix with the periodicity of the cyclic code.
[0145] Step 2.3.2.2: Extract the representative rows of the embedding matrix and the mask matrix through the cyclic relationship between the nodes of the Tanner graph. The representative rows of the embedding matrix form a representative matrix, and the representative rows of the mask matrix form a new mask matrix.
[0146] Step 2.3.2.3: Use the representative matrix and the new mask matrix to perform the operation of the attention matrix in the Transformer, and perform a cyclic shift after the calculation is completed to obtain the complete attention matrix. As Figure 3 shown, the output self-attention matrix is used as the input and input into the feed-forward neural network.
[0147] In a specific implementation, as Figure 4 shown, the self-attention mechanism is the core part of the Transformer architecture, and its calculation process usually involves a large number of matrix multiplications and weighting operations. To reduce the computational complexity, in this method, first, representative rows with cyclic equivalence are extracted to construct a new embedding matrix Φ' and a mask matrix. Then, by multiplying Φ' with the query matrix W Q and multiplying it with the key matrix K T and then adding the mask matrix, the attention matrix is obtained. Finally, it is restored by cyclic shift through Algorithm 2. In the whole calculation process, the calculation process of the attention matrix is further optimized by cyclic reuse, avoiding unnecessary repeated calculations.
[0148] Specifically, the calculation formula of the attention matrix is modified to:
[0149]
[0150] where represents finding the representative row from , is the selected representative mask matrix, and PCR() is the cyclic shift method, that is, Algorithm 2 defined above.
[0151] Step 2.4: Use the optimized decoding model to obtain a preliminary decoding result for the codeword to be decoded.
[0152] Step 3: The preliminary decoding result undergoes post-processing to obtain the final decoding result.
[0153] During the post - processing, the minimum mean - square error (MMSE) decoding method is adopted to estimate the multiplicative noise The goal of this process is to estimate the noise from the received signal y And the predicted codeword is obtained through the following formula:
[0154]
[0155] where is the predicted codeword after decoding, bin represents binarizing the value, and sign represents taking the sign function. Through these pre - processing and post - processing steps, the receiving end can effectively recover the received signal, reduce the influence of noise, and further improve the decoding accuracy.
[0156] To better illustrate the superiority of the method of the present invention, the following verification is carried out:
[0157] Figure 5 is the parity - check matrix of the (7, 4) Hamming code, Figure 6 is the Tanner graph of the parity - check matrix. As Figure 5 and Figure 6 shown, according to the definition of the relation function, it is known that f(c0)=(v0, v2, v3, v4: c0), f(c1)=(v1, v3, v4, v5: c1), f(c2)=(v2, v4, v5, v6: c2). Then, from the definition of the equivalence relation, it can be known that f(c2)=f(c1)+1=f(c0)+2. So there is a strong Tanner cycle equivalence between these three nodes. Further, it can be proved that there is a strong Tanner cycle equivalence between all check nodes.
[0158] Then consider variable nodes v0 and v1. Similarly, it is found that f(v1)=f(v0)+1=(v1, v3, v4, v5: c1). So there is also a strong Tanner cycle equivalence between these two nodes. In addition, consider v0 and v6. It can be found that the relationship structure between f(v0)=(v0, v2, v3, v4: c0) and f(v6)=(v2, v4, v5, v6: c2) is basically the same at this time, that is, f(v6)=f(v0)+2. At this time, these two nodes satisfy weak Tanner cycle equivalence.
[0159] Finally, due to the mapping relationship between the parity - check matrix and the Tanner graph, the relationship between the nodes in the Tanner graph can be mapped to the parity - check matrix, that is, Algorithm 1. According to Algorithm 1, process Figure 1The parity-check matrix of the (7,4) Hamming code, under strong Tanner cycle equivalence, yields row_indices=(0,2,3,4,5,6,7) and shift_map={1:(0,1),8:(7,1),9:(7,2)}. This means that when decoding the (7,4) Hamming code using a Transformer, only the rows found by row_indices can be used as representative rows for operations, and then the ignored rows can be restored using the relationship (target row: (representative row, number of shifts)) in shift_map. Specifically, as Figure 7 shown, taking the mask matrix as an example, under the operation of Algorithm 2, the mask matrix can participate in the operation only with the rows in row_indices as representatives, and the remaining rows can be restored by cyclic shifting using shift_map.
[0160] The cyclic code decoding method based on the Transformer architecture proposed by the present invention solves multiple problems existing in the existing decoding methods by optimizing the computational complexity, reducing the number of parameters, and improving the decoding efficiency, and has the following remarkable beneficial effects:
[0161] 1. Significantly reduce the computational complexity:
[0162] By introducing the cycle equivalence of the Tanner graph and the reuse of cyclic shift operations, the present invention can effectively reduce unnecessary computations and improve the computational efficiency of the decoder. The algebraic structure of the cyclic code and the periodic characteristics of the Tanner graph enable the identification of equivalent nodes and computations during the decoding process, thus avoiding the repeated processing of the same information and reducing the computational amount. Especially in the case of long codewords and high-dimensional data, traditional decoding methods require a large amount of computations and iterations, while the present invention greatly reduces the redundant computations by utilizing cycle equivalence and significantly improves the computational efficiency.
[0163] Specifically, as Figure 8 shown, when decoding BCH(31,16), BCH(63,36), BCH(63,45), BCH(63,51), the present invention only requires 10.6%, 14.0%, 27.1%, and 31.9% of the computational amount respectively compared to the self-attention matrix calculation of the previous Transformer decoder (ECCT).
[0164] 2. Greatly reduce the number of parameters:
[0165] Existing Transformer decoders usually contain a large number of parameters, especially the query, key, and value matrices and the feed-forward neural network in the self-attention mechanism, resulting in a sharp increase in the memory consumption and computational complexity of the model. By introducing the algebraic structure of cyclic codes and the cyclic equivalence of Tanner graphs, the present invention significantly reduces the number of parameters in the decoder by reusing parameter matrices and reducing redundant computations.
[0166] Specifically, through the analysis of the relationship between variable nodes and check nodes in the Tanner graph, the present invention reuses some of the relationship matrices through cyclic shift, thus reducing the need for repeated computations. This design not only effectively reduces the storage requirements of the decoder but also makes the training and inference of the model more efficient. As Figure 9 shown, C-numbers represent in which linear layers the cyclic reuse of parameters is adopted. For example, C-4 means that the cyclic reuse is adopted in the first four linear layers. Experiments show that in BCH(31,16), BCH(63,36), BCH(63,45), and BCH(63,51), the number of parameters of the decoder of the present invention is reduced by 64.4%, 61.2%, 47.6%, and 41.0% respectively compared with the past, greatly reducing the burden of hardware deployment and making it possible to deploy in embedded systems and mobile devices with limited computing resources.
[0167] 3. Decoding performance is maintained or improved:
[0168] Although the present invention has made significant optimizations in reducing computational complexity and the number of parameters, its decoding performance has not decreased. Table 1 is a comparison table of the decoding performance of the method of this embodiment and other similar neural network methods.
[0169] Table 1 Comparison table of decoding performance.
[0170]
[0171] As shown in Table 1, "weak" represents the adoption of weak Tanner cyclic equivalence, and the rest default to strong Tanner cyclic equivalence. N = 10 represents the number of layers of the Transformer, and the rest default to 6. All experiments are carried out under AWGN. The experimental results show that through the optimization of the Transformer self-attention mechanism, the fixed-dimensional embedding design introduced by the present invention, and the effective application of Tanner graph cyclic equivalence, the decoding performance of the decoder remains at a high level under multiple signal-to-noise ratio (SNR) conditions, and in some cases, the performance is improved compared with traditional decoders.
[0172] 4. Good scalability:
[0173] The specific innovation of the decoder of the present invention lies in discovering the application of the cyclic equivalence of cyclic codes in Transformer, paying more attention to the cyclic characteristics, and making basically no modification to the model. Therefore, our solution has the characteristics of plug-and-play and still has the same effect on other Transformer decoders, that is, it has good scalability.
[0174] Embodiment 2:
[0175] Embodiment 2 of the present invention provides a decoder for cyclic codes with cyclic equivalence, including:
[0176] A preprocessing module, configured to obtain a cyclic codeword to be decoded, and preprocess the codeword to obtain a parameter matrix and an embedding matrix;
[0177] A model construction module, configured to construct a decoding model based on the cyclic equivalence of cyclic codes, associate the embedding matrix, the parameter matrix with the periodicity of the cyclic code, and use the decoding model to multiplex and decode the embedding matrix and the parameter matrix by means of cyclic shift to obtain a preliminary decoding result. Among them, the embedding matrix adjusts the self-attention matrix during the multiplexing decoding process, and the generated self-attention matrix is used to correct the generation of the preliminary decoding result during the multiplexing decoding process of the parameter matrix;
[0178] A postprocessing module, configured to perform postprocessing on the preliminary decoding result to obtain a final decoding result.
[0179] Embodiment 3:
[0180] Embodiment 3 of the present invention provides a medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the decoding method for cyclic codes with cyclic equivalence described in Embodiment 1 of the present invention.
[0181] Embodiment 4:
[0182] Embodiment 4 of the present invention provides a device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the decoding method for cyclic codes with cyclic equivalence described in Embodiment 1 of the present invention.
[0183] The steps involved in the above Embodiments 2, 3, and 4 correspond to those in Method Embodiment 1. For the specific implementation manners, reference may be made to the relevant description parts of Embodiment 1.
[0184] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0185] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts on the basis of the technical solution of the present invention are still within the protection scope of the present invention.
Claims
1. A decoding method for cyclic codes with cyclic equivalence, characterized in that: The following steps are involved: Obtaining a codeword of a cyclic code to be decoded, and preprocessing the codeword; A decoding model is constructed based on the cyclic equivalence of cyclic codes, a parameter matrix of the decoding model is initialized according to representative rows determined in the Tanner graph, an input codeword is preprocessed according to the parameter matrix to obtain an embedding matrix, the embedding matrix and the parameter matrix are associated with the periodicity of the cyclic code, and the decoding model is trained to obtain an optimized decoding model, wherein based on the input simulated noise, the parameter matrix and the embedding matrix are multiplexed and decoded by the decoding model through cyclic shift, and the embedding matrix calculates the parameter matrix based on the self-attention mechanism during the multiplexed decoding process to achieve the adjustment of the self-attention matrix; Using the optimized decoding model to obtain a preliminary decoding result for the codeword to be decoded; The preliminary decoding results are post-processed to obtain the final decoding results.
2. The decoding method for cyclic codes with cyclic equivalence according to claim 1, characterized in that: The specific steps of preprocessing the input codeword according to the parameter matrix to obtain the embedding matrix include: mapping each bit in the vector to the high-dimensional space by using a fixed-dimensional embedding method to obtain the embedding matrix.
3. The decoding method for cyclic codes with cyclic equivalence according to claim 2, characterized in that: In the fixed-dimensional embedding method, the dimension is fixed relative to each code, and different codes use different embedding dimensions.
4. The decoding method for cyclic codes with cyclic equivalence according to claim 1, characterized in that: The decoding model is a Transformer network structure, including a self-attention layer and a feedforward network layer.
5. The decoding method for cyclic codes with cyclic equivalence according to claim 4, characterized in that: The specific steps to associate the embedding matrix, parameter matrix and periodicity of cyclic code are: Extract representative rows of the embedding matrix and parameter matrix through the Tanner graph; When the embedding matrix and parameter matrix are needed, a new matrix is generated by cyclically shifting the representative rows.
6. The decoding method for cyclic codes with cyclic equivalence according to claim 5, characterized in that: The Tanner graph is a graph structure corresponding to the check matrix of a cyclic code. In the Tanner graph, the cyclic relationship between nodes includes two types of cyclic relationships: strong Tanner cyclic equivalence and weak strong Tanner cyclic equivalence.
7. The decoding method for cyclic codes with cyclic equivalence according to claim 6, characterized in that: The specific steps of calculating the parameter matrix based on the self-attention mechanism during the reuse decoding process of the embedding matrix are: Associating the embedding matrix with the periodicity of the cyclic code; The representative rows of the embedding matrix and the mask matrix are extracted through the cyclic relationship between the nodes of the Tanner graph. The representative rows of the embedding matrix form the representative matrix, and the representative rows of the mask matrix form the new mask matrix. The representative matrix and the new mask matrix are used to operate the attention matrix in the Transformer, and a circular shift is performed after the calculation is completed to obtain the complete attention matrix.
8. A decoder for cyclic codes and having cyclic equivalence, characterized in that include: A preprocessing module is configured to obtain a codeword of a cyclic code to be decoded and preprocess the codeword; The cyclic multiplexing module is configured to construct a decoding model based on the cyclic equivalence of the cyclic code, set the initial parameters of the decoding model, initialize the parameter matrix of the decoding model according to the representative row determined in the Tanner graph, pre-process the input codeword according to the parameter matrix to obtain an embedding matrix, associate the embedding matrix and the parameter matrix with the periodicity of the cyclic code, train the decoding model, and obtain an optimized decoding model, wherein based on the input simulated noise, the parameter matrix and the embedding matrix are multiplexed and decoded by the decoding model through cyclic shift, and the embedding matrix calculates the parameter matrix based on the self-attention mechanism during the multiplexing decoding process to realize the construction of the self-attention matrix; the optimized decoding model is used to obtain a preliminary decoding result for the codeword to be decoded; The post-processing module is configured to post-process the preliminary decoding result to obtain the final decoding result.
9. A computer-readable storage medium, characterized in that: A plurality of instructions are stored therein, and the instructions are suitable for being loaded by a processor of a terminal device and executing a decoding method for cyclic codes with cyclic equivalence as described in any one of claims 1 to 7.
10. A terminal device, characterized in that: The invention comprises a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; and the computer-readable storage medium is used to store a plurality of instructions, wherein the instructions are suitable for being loaded by the processor and executing the decoding method for cyclic codes with cyclic equivalence according to any one of claims 1 to 7.