RS code compiling method and system based on data stream

By converting the target matrix into a coefficient matrix and performing matrix vector multiplication, the problem that existing compilers are difficult to be compatible with various RS code algorithms and implementing real-time data stream processing is solved, and efficient compilation and decoding operations and system performance improvements are achieved.

CN120074546AActive Publication Date: 2025-05-30深圳倚宿科技有限公司
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510150604.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

Existing compilers are difficult to be compatible with various RS code algorithms, and it is difficult to achieve real-time or near-real-time processing of data streams during data stream transmission, which limits the compiler's applicability in different application scenarios and affects the system's response time and delay time.

Method used

By converting the target matrix into a coefficient matrix, converting the data packet into a byte vector, and performing matrix vector multiplication based on the coefficient matrix and byte vector, it supports any custom generation matrix and native polynomials, and is compatible with various RS code algorithms, and implements the path-to-path processing of compilation and decoding.

Benefits of technology

Compatibility with different RS code algorithms is achieved, the applicability of the compiler in different application scenarios is improved, and the system response time and delay time are improved through real-time or near-real-time data stream processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120074546A_ABST
    Figure CN120074546A_ABST
Patent Text Reader

Abstract

The invention discloses a data stream-based RS code compiling method and system. Relates to the technical field of coding and decoding. Comprising the following steps: acquiring k target data packets and a corresponding m * k target matrix; converting the target matrix into a coefficient matrix, and converting the data packet into a byte vector; matrix vector multiplication operation is carried out based on the preprocessed target matrix and the data packet to obtain target data; according to the scheme, the method and the structure are improved on the basis of the existing coding and decoding technology, the target matrix is converted into the coefficient matrix, the data packet is converted into the byte vector, the matrix vector multiplication operation is carried out on the basis of the coefficient matrix and the byte vector to obtain the target data, any self-defined generation matrix and primitive polynomials are supported, and the coding and decoding efficiency is improved. Various known RS code algorithms or RS code algorithms possibly appearing in the future are compatible, and channel associated processing of coding and decoding is achieved based on a data stream type data processing mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of encoding and decoding, and particularly relates to an RS code encoding and decoding method and system based on data streaming. Background Art

[0002] Due to the diversity of RS code algorithms, including RS codes of different lengths, different redundancies, and different primitive polynomials and generating polynomials that may be adopted, the compiler needs to be able to adapt to these different parameter configurations to ensure the correct implementation of the algorithm; existing compilers may not be compatible with various RS code algorithms, which limits the applicability of the compiler in different application scenarios; and existing compilers are difficult to process data streams in real time or near real time during the data stream transmission process, which is crucial for ensuring the response time of the system and reducing latency. Summary of the Invention

[0003] The technical problem to be solved by the present invention is that existing compilers may not be compatible with various RS code algorithms and are difficult to process data streams in real time or near real time during the data stream transmission process, which limits the applicability of the compiler in different application scenarios and also affects the response time and latency of the system; the purpose of the present invention is to provide an RS code encoding and decoding method and system based on data streaming, which makes improvements in method and structure on the basis of existing encoding and decoding technologies. This solution converts the target matrix into a coefficient matrix, converts the data packet into a byte vector, and performs matrix-vector multiplication operations based on the coefficient matrix and the byte vector to obtain the target data, supports any user-defined generating matrix and primitive polynomial, is compatible with various known or future possible RS code algorithms, and realizes the in-line processing of encoding and decoding based on the data streaming data processing mechanism.

[0004] The present invention is achieved through the following technical solutions:

[0005] This solution provides an RS code encoding and decoding method based on data streaming, including:

[0006] Obtain k target data packets and the corresponding target matrix; the dimension of the target matrix is m×k;

[0007] Preprocess the target matrix and the target data packet: convert the target matrix into a coefficient matrix, and convert the data packet into a byte vector;

[0008] Perform matrix-vector multiplication operations based on the preprocessed target matrix and data packet to obtain the target data, where the target data includes encoded data or decoded data;

[0009] Among them, when the target data packet is an information packet to be encoded, the target matrix is the generating matrix in the encoding process, and the target data is the encoded data; when the target data packet is an information packet to be decoded, the target matrix is the decoding matrix in the decoding process; the target data is the decoded data.

[0010] Working principle of this solution: This solution is provided based on a data-streaming RS code encoding and decoding method, which improves the method and structure on the basis of existing encoding and decoding technologies. This solution converts the target matrix into a coefficient matrix, converts the data packet into a byte vector, and performs matrix-vector multiplication operations based on the coefficient matrix and the byte vector to obtain the target data. It supports any user-defined generating matrix and primitive polynomial, is compatible with various known or future possible RS code algorithms, and realizes in-path processing of encoding and decoding based on a data-streaming data processing mechanism.

[0011] A further optimized solution is that the preprocessing of the target matrix and the target data packet includes the method:

[0012] Represent the m×k elements in the target matrix with 8×8 bit matrices respectively to obtain an 8m×8k target bit matrix, and transpose the target bit matrix to obtain a coefficient matrix;

[0013] Convert each target data packet into 8 byte elements to obtain a column vector composed of 8k byte elements, where the byte configuration of each element is stride.

[0014] A further optimized solution is that the method for obtaining the coefficient matrix includes:

[0015] Calculate the element c in the i-th row and j-th column of the coefficient matrix according to the following formula i,j :

[0016] c i,j =[a j,i / 8 <<(i%8)]%P;

[0017] 0≤i<8k;

[0018] 0≤j<m;

[0019] Among them, (#) / (*) represents performing integer division operations on (#) and (*); (#)%(*) represents performing exclusive OR operations on (#) and (*); P represents the primitive polynomial; (#)<<(*) represents shifting (#) to the left by (*) times; a j,i / 8 represents the element in the j-th row and i / 8 of the target matrix.

[0020] A further optimized solution is that the method of converting each target data packet into 8 byte elements to obtain a column vector composed of 8k byte elements includes the method:

[0021] Obtain the fixed bit width w of the matrix-vector multiplication operation process;

[0022] Configure the bit width stride of each element of the column vector;

[0023] When the bit width stride is not an integer multiple of the fixed bit width w, fill the elements corresponding to the last data packet in the column vector with 0.

[0024] A further optimization solution is that the method for obtaining the decoding matrix includes:

[0025] Obtain the encoding generation matrix A corresponding to the decoding matrix with dimensions m A ×k A , the k A -dimensional data packet column vector B, the encoded m A -dimensional redundant packet column vector C, and the number Q of invalid data packets in k A +m A data packets; where Q < k A +m A ;

[0026] Based on the encoding generation matrix A, the data packet column vector B, and the redundant packet vector C, construct a matrix-vector multiplication equation: A×B = H×C; where H represents the transformation matrix;

[0027] Reorganize the matrix-vector multiplication equation based on the invalid data packets;

[0028] Multiply both sides of the reorganized matrix-vector multiplication equation by the inverse matrix N, where the inverse matrix N is the inverse matrix of the Q×Q dimensional coefficient matrix V, to obtain the decoding matrix calculation model;

[0029] Solve the decoding matrix based on the Gaussian elimination method and the sum code matrix calculation model.

[0030] A further optimization solution is that the reorganization of the matrix-vector multiplication equation based on the invalid data packets includes the method:

[0031] Reorganize the encoding generation matrix A in the matrix-vector multiplication equation: delete the elements that are the coefficients of the invalid data packets in the encoding generation matrix A, and fill them with the elements in the transformation matrix H by shifting;

[0032] Reorganize the data packet column vector B in the matrix-vector multiplication equation: delete the invalid data packets in the data packet column vector B, and fill them with the elements in the redundant packet column vector C by shifting;

[0033] Reorganize the redundant packet column vector C in the matrix-vector multiplication equation: replace the elements in the redundant packet column vector C with the invalid data packets;

[0034] Recombine the coefficient matrix H in the matrix-vector multiplication equation: Replace the elements in the coefficient matrix H with the elements in the encoding generation matrix A that are the coefficients of the lost data packets;

[0035] The shift filling means that the empty positions after deleting elements in the matrix to be filled are filled by the elements in the subsequent column in sequence, and then the newly generated empty positions are filled.

[0036] A further optimized solution is that the coefficient matrix V is composed of the elements in the encoding generation matrix A that are the coefficients of the lost data packets.

[0037] This solution also provides a data-streaming-based RS code encoding and decoding system for implementing the above-mentioned data-streaming RS code encoding and decoding method; the system includes:

[0038] A data input module for obtaining k target data packets and corresponding target matrices; the dimension of the target matrix is m×k; the data input module is also used to convert the data packets into byte vectors;

[0039] A coefficient matrix module for converting the target matrix into a coefficient matrix;

[0040] A calculation array module for performing matrix-vector multiplication operations on the preprocessed target matrix and data packets to obtain target data;

[0041] Among them, when the target data packet is an information packet to be encoded, the target matrix is the generation matrix in the encoding process, and the target data is the encoded data; when the target data packet is an information packet to be decoded, the target matrix is the decoding matrix in the decoding process; the target data is the decoded data.

[0042] A further optimized solution is that it also includes a PCIe module. The data input module batch-transfers k target data packets from the host through the PCIe module in the manner of a DMA engine. The data input module is also used to configure the bit width stride of each element of the column vector, and when the bit width stride is not an integer multiple of the fixed bit width w in the matrix-vector multiplication operation process, fill the elements corresponding to the last data packet in the column vector with 0.

[0043] A further optimized solution is that it also includes a calculation buffer module for temporarily storing some data of the accumulation operation during the matrix-vector multiplication operation. The calculation buffer module and the calculation array module form a circular data stream channel.

[0044] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0045] The RS code encoding and decoding method and system based on data streaming provided by the present invention improve the method and structure on the basis of the existing encoding and decoding technology. This solution converts the target matrix into a coefficient matrix, converts the data packet into a byte vector, and performs matrix-vector multiplication operations based on the coefficient matrix and the byte vector to obtain the target data. It supports any user-defined generator matrix and primitive polynomial, is compatible with various known or future possible RS code algorithms, and realizes the in-line processing of encoding and decoding based on the data streaming data processing mechanism. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0047] Figure 1 It is a schematic flowchart of the RS code encoding and decoding method based on data streaming;

[0048] Figure 2 It is a schematic diagram of the inference process of element-wise multiplication from the vector multiplication mode to the matrix-vector multiplication mode;

[0049] Figure 3 It is a schematic diagram of the RS code encoding process represented by the generator matrix;

[0050] Figure 4 It is a schematic diagram of the encoding operation process represented by the bit generator matrix;

[0051] Figure 5 It is a schematic diagram of the generation process of the coefficient matrix;

[0052] Figure 6 It is a schematic diagram of the generation process of the decoding matrix D;

[0053] Figure 7 It is a schematic diagram of the structure of the RS code encoding and decoding system based on data streaming;

[0054] Figure 8 It is a schematic diagram of the working principle of the RS code encoding and decoding system based on data streaming;

[0055] Figure 9 It is a schematic diagram of the storage method of a single data packet in the memory;

[0056] Figure 10 It is a schematic diagram of the microarchitecture of the data input module;

[0057] Figure 11 It is a schematic diagram of the microarchitecture of the input buffer module;

[0058] Figure 12 Schematic diagram of the microarchitecture of the coefficient matrix module;

[0059] Figure 13 Schematic diagram of the microarchitecture of the computing buffer module;

[0060] Figure 14 Schematic diagram of the microarchitecture of the data output module. Detailed implementation manners

[0061] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the embodiments and the accompanying drawings. The illustrative embodiments and descriptions thereof of the present invention are only used to explain the present invention and are not intended to limit the present invention.

[0062] Existing compilers may not be compatible with various RS code algorithms, and it is difficult to process data streams in real time or near real time during the data stream transmission process, which limits the applicability of the compiler in different application scenarios and also affects the response time and latency of the system. In view of this, the present invention provides the following embodiments to solve the above technical problems.

[0063] Embodiment 1

[0064] This embodiment provides an RS code compilation method based on data streaming, as Figure 1 shown, including:

[0065] Step 1, obtaining k target data packets and corresponding target matrices; the dimension of the target matrix is m×k;

[0066] Step 2, preprocessing the target matrix and the target data packets: converting the target matrix into a coefficient matrix and converting the data packets into byte vectors; this step specifically includes the method:

[0067] Representing each of the m×k elements in the target matrix with an 8×8 bit matrix to obtain an 8m×8k target bit matrix, and transposing the target bit matrix to obtain a coefficient matrix;

[0068] Converting each target data packet into 8 byte elements to obtain a column vector composed of 8k byte elements, where the byte configuration of each element is stride. Specifically:

[0069] S01, obtaining the fixed bit width w of the matrix-vector multiplication process;

[0070] S02, configuring the bit width stride of each element of the column vector;

[0071] S03. When the bit width stride is not an integer multiple of the fixed bit width w, the elements corresponding to the last data packet in the column vector are padded with 0.

[0072] In application scenarios such as storage, RS codes need to process a large amount of data in batches. To improve efficiency, encoding cannot be performed only in units of bytes. In this solution, each data packet is regarded as 8 fixed elements, and the size of each element changes with the size of the data packet. The size of a single element is denoted as stride bytes.

[0073] The encoding and decoding process of RS codes is a matrix-vector multiplication operation process in which element operations are defined in the finite field GF(2 8 ). In the finite field, the multiplication operation requires feedback for modulo calculation. To eliminate the differences brought by different primitive polynomials in the finite field, this solution makes transformations on the representation and multiplication operation of finite field elements, improves the parallelism of operations, and facilitates hardware acceleration.

[0074] There are two representation methods for elements in the finite field, namely vector representation and matrix representation. For any two elements A and B in the finite field, the vector representations of element A and element B are {a 7 , a 6 , a 5 , a 4 , a 3 , a 2 , a 1 , a 0} and {b 7 , b 6 , b 5 , b 4 , b 3 , b 2 , b 1 , b 0}, where a i and b i are both single bits, taking values of 0 or 1. The operation of C = A × B can be expressed as vector multiplication, matrix-vector multiplication, and matrix multiplication, and the obtained results are the vector representation and matrix representation of C respectively. In this solution, the M × K elements in the generator matrix G in the encoding process of RS codes and the decoding matrix D in the decoding process are represented by 8 × 8 bit matrices, and are extended to an 8M × 8K bit matrix as the coefficient matrix. The data packets are directly represented by vectors, and the encoding and decoding operations are completed in the form of matrix-vector multiplication. The reasoning process of the multiplication between elements from the vector multiplication mode to the matrix-vector multiplication mode is as shown in Figure 2 . In the figure, the multiplication and addition operations between bits are defined in GF(2), that is, the addition is the exclusive OR operation and the multiplication is the AND operation:

[0075] In the figure, formula ① is vector multiplication: After aligning vector A by shifting it according to the weights of the coefficients in vector B, a bitwise exclusive OR operation is performed, and a polynomial P is generated based on the modulus of the final result.

[0076] In formula ②, for each row of the shaded part of vector A, the exclusive OR operation with the primitive polynomial P can be performed in advance, making the subsequent operations independent of P. That is, vector A is shifted one bit to the left each time, with 0 filled in the low bit. If the highest bit shifted out is 1, the result is exclusive ORed with the lower 8 bits of P.

[0077] In formula ③, vector A is shifted 7 times and the modulus with respect to P is taken to obtain 8 column vectors.

[0078] In formula ④, the multiplication can be represented as the multiplication of the bit matrix representation of A, which is a combination of 8 column vectors, and the column vector B.

[0079] During the encoding and decoding process, the coefficient matrix is related to the encoding and decoding mode and can be determined and converted into a bit matrix representation before the operation starts. That is, A corresponds to the elements in the generator matrix or the decoding matrix, and B corresponds to the elements in the data packet. The matrix representation of A is the responsibility of the software, and the encoding and decoding calculations of the hardware accelerator are independent of the primitive polynomial P.

[0080] As Figure 3 shown, the generator matrix G of the RS code in this embodiment has M rows and K columns, and the elements are defined in GF(2 8 ). The encoding process involving the bit matrix is shown in Figure 4 , compared with Figure 3 , each element in the generator matrix G is changed to an 8×8 bit matrix, and the M×K generator matrix G is thus transformed into a target bit matrix of 8M×8K. And the K data packets are transformed into a vector-like form of 8K elements, and each element has stride bytes. Since the hardware accelerator processes data on the fly and does not store the input data packets, the coefficient matrix participates in the calculation column by column. The elements of each data packet (i.e., stride bytes) are exclusive ORed and accumulated under the bit control of 8M column vectors. The pseudocode is as follows:

[0081] "for(j = 0; j < 8*K; j++)

[0082] for(i = 0; i < 8*M; i++)

[0083] Sum[i] += Data[j] * G[i][j]; / / Sum[i] and Data[j] are both stride-byte vectors

[0084] / / G[i][j] is a single bit".

[0085] It can be seen from the operation process represented by the pseudocode that the K data packets Data only need to be read once, and for every stride bytes fetched, the accumulation operation is completed according to the columns of the generating matrix without storing the input data. Since the generating matrix G is used column by bit, the extended bit generating matrix needs to be transposed into the coefficient matrix for subsequent accelerator calculations, as shown in Figure 5 as follows:

[0086] Figure 5 Equation ① is to expand each element a 8 in GF(2 i,j ) of G into an 8×8 bit matrix. For example, the 8×8 bit matrix of the shaded elements from g 0,0 to g 7,7 is the bit matrix representation of a 0,0 ;

[0087] Figure 5 In Equation ②, the matrix is derived column by column, and every 7 column bits form a byte. For example, each of the two dashed boxes in the figure has 8 bits, corresponding to the elements c 0,0 and the element c 1,0 respectively;

[0088] Figure 5 Equation ③ is the transformation relationship between the elements of the coefficient matrix and the elements of the generating matrix obtained by considering the above content:

[0089] The element c i,j in the i-th row and j-th column of the coefficient matrix is obtained according to the following formula:

[0090] c i,j = [a j,i / 8 << (i % 8)] % P;

[0091] 0 ≤ i < 8k;

[0092] 0 ≤ j < m;

[0093] where, (#) / (*) represents the integer division operation of (#) and (*); (#) % (*) represents the exclusive OR operation of (#) and (*); P represents the primitive polynomial; (#) << (*) represents shifting (#) to the left by (*) times; a j,i / 8 represents the element in the j-th row and i / 8-th of the target matrix.

[0094] Step 3: Perform matrix-vector multiplication on the preprocessed target matrix and data packets to obtain target data, where the target data includes encoded data or decoded data. When the target data packet is an information packet to be encoded, the target matrix is the generation matrix in the encoding process, and the target data is the encoded data. When the target data packet is an information packet to be decoded, the target matrix is the decoding matrix in the decoding process, and the target data is the decoded data.

[0095] The method for obtaining the decoding matrix includes:

[0096] Obtain the encoding generation matrix A of dimension m A ×k A for the decoding matrix, the data packet column vector B of dimension k A the redundant packet column vector C of dimension m obtained by encoding, and the number Q of failed data packets in k A +m A data packets; where Q < k A +m A +m A ;

[0097] S11: Construct a matrix-vector multiplication equation based on the encoding generation matrix A, the data packet column vector B, and the redundant packet vector C: A × B = H × C, where H represents the transformation matrix.

[0098] S12: Reorganize the matrix-vector multiplication equation based on the failed data packets. This step specifically includes the following methods:

[0099] Reorganize the encoding generation matrix A in the matrix-vector multiplication equation: Delete the elements that are coefficients of the failed data packets in the encoding generation matrix A, and fill the positions with elements shifted from the transformation matrix H.

[0100] Reorganize the data packet column vector B in the matrix-vector multiplication equation: Delete the failed data packets in the data packet column vector B, and fill the positions with elements shifted from the redundant packet column vector C.

[0101] Reorganize the redundant packet column vector C in the matrix-vector multiplication equation: Replace the elements in the redundant packet column vector C with the failed data packets.

[0102] Reorganize the coefficient matrix H in the matrix-vector multiplication equation: Replace the elements in the coefficient matrix H with the elements that are coefficients of the failed data packets in the encoding generation matrix A.

[0103] The shifted filling means that the empty positions after deleting elements in the matrix to be filled are filled by the elements in the subsequent column in sequence, and then the newly generated empty positions are filled.

[0104] S13. Multiply both sides of the recombined matrix-vector multiplication equation by the inverse matrix N, where the inverse matrix N is the inverse matrix of the Q×Q dimensional coefficient matrix V, to obtain a decoding matrix calculation model; the coefficient matrix V is composed of the elements in the encoding generation matrix A that are coefficients of the failed data packets.

[0105] S14. Solve the decoding matrix based on the sum code matrix calculation model using Gaussian elimination.

[0106] In this solution, the difference between RS decoding and encoding lies in the different coefficient matrices calculated by software, and matrix-vector multiplication does not distinguish between encoding and decoding; the calculation of the decoding matrix D depends on the encoding generation matrix A, the primitive polynomial P, and the pattern of packet loss or damage, and it is required that the number of failed data packets Q (in this embodiment, the failed data packets mainly include lost or damaged data packets;) is not greater than m A . The inverse matrix is solved with k A = 4, m A = 2, Q = 2 as an example, as Figure 6 shown:

[0107] Figure 6 Equation ① of is the constructed matrix-vector multiplication equation, where A is the encoding generation matrix, B is the data packet column vector composed of 4 information packets, C is the redundant packet column vector composed of 2 encoded redundant packets, and a total of b 0 , b 1 , b 2 , b 3 , c 0 , c 1 A total of 6 data packets, each data packet corresponding to a coefficient column. The columns of the information packets come from the encoding matrix, the columns of the redundant packets are one-hot encoded, and the i-th row of the coefficient column of the i-th redundant packet is 1;

[0108] Figure 6 Equation ② of is the matrix-vector multiplication equation recombined based on the failed data packets. Among them, data packets b 1 and data packet b 2 are damaged or lost. The addition in the finite field is an exclusive OR operation, and this equation can be obtained. The left matrix is recombined from the coefficient columns of the received or error-free data packets, and the right matrix is recombined from the coefficient columns of the lost or error-occurring data packets;

[0109] Figure 6 Equation ③ of is the calculation formula of the decoding matrix D obtained by multiplying both ends of Equation ② by the inverse matrix N, where multiplying the inverse matrix N by the recombined encoding generation matrix A can obtain the decoding matrix D.

[0110] Embodiment 2

[0111] This embodiment provides an RS code compilation system based on data streaming, which is used to implement a data streaming RS code compilation method described in Embodiment 1; as Figure 7 and Figure 8 shown, the system includes:

[0112] A data input module, which is used to obtain k target data packets and corresponding target matrices; the dimension of the target matrix is m×k; the data input module is also used to convert the data packets into byte vectors;

[0113] A coefficient matrix module, which is used to convert the target matrix into a coefficient matrix;

[0114] A calculation array module, which is used to perform matrix-vector multiplication operations based on the preprocessed target matrix and data packets to obtain target data;

[0115] Among them, when the target data packet is an information packet to be encoded, the target matrix is the generating matrix in the encoding process, and the target data is the encoded data; when the target data packet is an information packet to be decoded, the target matrix is the decoding matrix in the decoding process; the target data is the decoded data.

[0116] It also includes a PCIe module. The data input module batch-transports k target data packets from the host computer through the PCIe module in the manner of a DMA engine. The data input module is also used to configure the bit width stride of each element of the column vector, and when the bit width stride is not an integer multiple of the fixed bit width w in the matrix-vector multiplication operation process, fill the elements corresponding to the last data packet in the column vector with 0.

[0117] It also includes a calculation buffer module, which is used to temporarily store part of the data of the accumulation operation during the matrix-vector multiplication operation. The calculation buffer module and the calculation array module form a circular data flow channel.

[0118] In this embodiment, the hardware of the entire system is mainly an accelerator, and the accelerator is connected to the host computer in the form of a PCIe acceleration card to complete the main calculation tasks of encoding and decoding. As Figure 8Among them, the core of the acceleration card is the FPGA chip, which internally includes two major parts: the PCIe module and the computing engine. The PCIe module implements a PCIe Endpoint device, and the application layer implements a DMA engine. PIO read and write are achieved through the AXI Lite interface, and DMA read and write are achieved through the AXI Stream interface. The computing engine includes parts such as a data input module, a computing array module, a coefficient matrix module, a computing buffer module, and a data output module, which realizes the acceleration of data-streaming RS code encoding and decoding calculations. When the data input module, the coefficient matrix module, and the computing buffer module can all provide relevant data or coefficients for starting the calculation, the calculation is executed in a pipeline in the computing array module without the need for additional start control or synchronization signals. Generally, the calculation is carried out under the drive of the data provided by the data input module. After the calculation task is completed, the data output module takes away the result in the computing buffer module and writes it back to the host computer. The data input module and the data output module exchange data with the host computer through DMA, and the coefficient matrix module and the configuration registers of the computing engine accept the configuration from the host computer through the PIO method. The bit width of the computing engine is defined by the width w bytes of the exclusive OR operation in the computing array module.

[0119] The data input module receives the data packets transported by DMA from the host computer to the board through the AXI Stream interface, sorts out the data, and then hands it over to the computing array module to promote the progress of data-streaming calculations. A single encoding and decoding calculation task can contain multiple batches of calculation data. A single batch of calculation data consists of K data packets, each data packet consists of 8 calculation stripes, and each calculation stripe contains stride bytes. The storage method of a single data packet in the memory is as Figure 9 shown. DMA transports the data to the board in the form of a byte stream and hands it over to the data input module. Stride, as an encoding parameter, can be configured, while the bit width w of the computing engine is fixed. When Stride is not an integer multiple of w, the data input module needs to pad the data in the last cycle.

[0120] The microarchitecture of the data input module is as Figure 10As shown in the figure, the data input module receives data fetched from the host by the DMA read channel. There is a multiple relationship between the data width dw of the AXI Stream interface of the DMA and the bit width w of the computing engine. Since the stripes of the data packet are stored continuously in memory, the data obtained by the DMA per cycle may contain data of two adjacent strides. However, the data input to the computing engine per cycle must belong to the same stripe. The byte slipbuffer module is responsible for byte alignment operations, cutting and splicing data that does not belong to the same stripe, and padding with 0 for data less than w bits. The width of the input buffer is w bits, with the same frequency and width as the computing array module, belonging to the same clock domain. Since each stripe needs to perform an exclusive OR operation with 8*K cached stripes under the control of 8*K coefficients, the reusebuffer caches the multi-cycle data of a single stripe. When calculating the last coefficient of the previous stripe, the new stripe enters the reuse buffer from the data buffer, and then loads the data from the input to the input port in a circular buffer mode, reusing 8*K cycles. The Reuse counter state machine controls the data scheduling and usage of the data buffer and the reuse buffer.

[0121] The computing array completes the exclusive OR operation of the coding and decoding calculation core, performs an exclusive OR calculation on the input data provided by the input buffer module and the sum provided by the computing buffer module under the control of the coefficient bits provided by the coefficient matrix module, and writes it back to the computing buffer module, as Figure 11 shown. Only when the input data, calculation coefficients, and sum are all valid, and there is space in the computing buffer module, the output is valid, and the signal indicating the calculation is valid is fed back to the modules in four directions. Signals such as data_in, sum_out, and sum_in are w bits wide, and the rest of the signals are single-bit wide.

[0122] The coefficient matrix module is responsible for storing and reusing the coding and decoding coefficients, and controlling the execution of the exclusive OR operation in the computing array, as Figure 12 shown. The width of the coefficient buffer is one byte, and the depth is 128K, that is, the capacity is 128KB, which can support the coding and decoding of RS codes of any dimension within GF(28). The coefficient buffer is a circular buffer that reuses the coefficient matrix and can complete the calculation of multiple batches of data in a single task. Only when the configuration of the coding and decoding of the calculation task changes, the coefficient matrix needs to be reinitialized. The coe_rd signal is equivalent to the coe_in_rdy of the computing array module. Whenever a cycle of calculation is completed, the reuse counter is incremented by 1, and each coefficient is reused Next, when the reuse counter overflows, the rotate counter is incremented by 1, and the next bit in the selected byte is used as the coefficient output. When the rotate counter reaches 8, it is shifted out, the byte is removed from the coefficient buffer and rewritten to the end of the buffer. During the initialization phase of the computing task, the coefficient buffer receives the coefficient matrix configured by software from the PIO channel.

[0123] The computing buffer module is used to temporarily store part of the accumulated redundant packet data during encoding and decoding, with a total of M * 8 * stride bytes, as Figure 13 shown. During encoding and decoding calculations, the computing buffer and the computing array module form a circular data stream. The data in the computing buffer is XORed with the sum_in of the computing array and the data_in of the computing array under the control of coe_in, and the result is written back to the computing buffer module as the sum_out of the computing array. During the initialization phase, the computing buffer can load a partial sum of all 0s equal to the size of the redundant packet, or load the partial sum of the previous calculation or a specific initial vector from the data input module. At the end stage, the data in the computing buffer can be exported to the data output module.

[0124] It also includes, as Figure 14 shown, a data output module. The data output buffer retrieves the calculation result from the computing buffer module, processes it, and then sends it back to the host main memory through a write operation by the DMA engine. The byte alignment buffer (byte slipbuffer) module is responsible for removing the filled 0s in the result and merging the stride into a continuous byte stream for DMA transmission.

[0125] This solution supports encoding and decoding algorithms for various RS codes, including those based on Vandermonde matrices and Cauchy matrices. The system completes the computing task in a software-hardware co - operation manner. The software is responsible for configuring the working mode of the hardware accelerator and scheduling the execution of multiple computing tasks. The hardware reads and writes data through PCIe in DMA mode and completes a large number of data encoding and decoding operations within a certain time in the software - configured working mode, completing each computing task one by one. Compared with various software and hardware implementations of existing RS code encoders and decoders, it has the following advantages:

[0126] Compared with software encoders and decoders, the computing core of this solution is implemented in hardware, with strong computing power and high energy efficiency. The encoding and decoding overhead and performance of the hardware implementation are the same.

[0127] Compared with various dedicated RS code hardware accelerators, this solution supports any user - defined generator matrix and primitive polynomial, and is compatible with various known or future - possible RS code algorithms.

[0128] This solution is based on a data streaming data processing mechanism, which realizes in-channel processing of encoding and decoding. The system does not require additional storage overhead such as board-level memory, and the calculation meets the real-time index of hard delay. The calculation does not limit the data supply mode and bandwidth.

[0129] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A data stream-based RS code compilation method, characterized in that: include: Get k target data packets and corresponding target matrices; The dimension of the target matrix is ​​m×k; Preprocess the target matrix and target data packet: convert the target matrix into a coefficient matrix and convert the data packet into a byte vector; Performing a matrix-vector multiplication operation based on the preprocessed target matrix and the data packet to obtain target data, wherein the target data includes encoded data or decoded data; Wherein, when the target data packet is an information packet to be encoded, the target matrix is ​​a generation matrix in the encoding process, and the target data is the encoding data; when the target data packet is an information packet to be decoded, the target matrix is ​​a decoding matrix in the decoding process; The target data is decoded data.

2. The RS code compilation method based on data stream according to claim 1 is characterized in that: The preprocessing of the target matrix and the target data packet includes the following method: The m×k elements in the target matrix are represented by 8×8 bit matrices respectively to obtain a 8m×8k target bit matrix, and the target bit matrix is ​​transposed to obtain a coefficient matrix; Each target data packet is converted into 8 byte elements, and a byte column vector consisting of 8k byte elements is obtained, where the bytes of each element are configured as stride.

3. The RS code compilation method based on data stream according to claim 2 is characterized in that: The method for obtaining the coefficient matrix includes: The element c in the i-th row and j-th column of the coefficient matrix is ​​calculated according to the following formula: i,j : c i,j =[a j,i / 8 <<(i%8)]%P; 0≤i<8k; 0≤j <m; Among them, (#) / (*) means integer division of (#) and (*); (#)%(*) means exclusive-OR operation of (#) and (*); P means primitive polynomial; (#)<<(*) means shifting (#) to the left (#) times; a j,i / 8 Represents the element at row i / 8 in the target matrix.

4. The RS code compilation method based on data stream according to claim 2 is characterized in that: The method of converting each target data packet into 8 byte elements to obtain a column vector consisting of 8k byte elements includes: Get the fixed bit width w of the matrix-vector multiplication process; Configure the bit width stride of each element of the column vector; When the bit width stride is not an integer multiple of the fixed bit width w, the element corresponding to the last data packet in the column vector is filled with 0.

5. The RS code compilation method based on data stream according to claim 1 is characterized in that: The method for obtaining the decoding matrix includes: Get the decoding matrix corresponding to m A ×k A The encoding matrix A, k of dimension A dimensional data packet column vector B, the encoded m A dimensional redundant packet column vector C, and k A +m A The number of failed packets among the packets is Q; where Q < k A +m A ; Based on the coding matrix A, the data packet column vector B and the redundant packet vector C, a matrix-vector multiplication equation is constructed: A×B=H×C; where H represents the conversion matrix; reorganizing the matrix-vector multiplication equation based on the failed data packet; Multiply both sides of the reorganized matrix-vector multiplication equation by the inverse matrix N, which is the inverse matrix of the Q×Q dimension coefficient matrix V, to obtain a decoding matrix calculation model; The decoding matrix is ​​solved based on the Gaussian elimination method and the code matrix calculation model.

6. The data stream-based RS code compilation method according to claim 5, characterized in that: The method of reorganizing the matrix-vector multiplication equation based on the failed data packet includes: The coding matrix A in the matrix-vector multiplication equation is reorganized: the elements of the coding matrix A that are coefficients of invalid data packets are deleted, and the elements in the conversion matrix H are shifted and filled; The data packet column vector B in the matrix-vector multiplication equation is reorganized: the invalid data packets in the data packet column vector B are deleted, and the elements in the redundant packet column vector C are shifted and filled; The redundant packet column vector C in the matrix-vector multiplication equation is reorganized: the elements in the redundant packet column vector C are replaced with the failed data packets; The coefficient matrix H in the matrix-vector multiplication equation is reorganized by replacing the elements in the coefficient matrix H with the elements in the coding generation matrix A that are coefficients of the failed data packets; The shift filling means that the vacancies after deleting the elements in the matrix to be filled are filled with the elements of the following column in sequence, and then the newly generated vacancies are filled.

7. The RS code compilation method based on data stream according to claim 6 is characterized in that: The coefficient matrix V is composed of elements in the coding generation matrix A that are coefficients of failed data packets.

8. The RS code compilation system based on data stream is characterized by: A data stream RS code compilation method for implementing any one of claims 1 to 7; the system comprises: A data input module, used to obtain k target data packets and corresponding target matrices; the dimension of the target matrix is ​​m×k; the data input module is also used to convert the data packets into byte vectors; The coefficient matrix module is used to convert the target matrix into a coefficient matrix; A calculation array module is used to perform matrix-vector multiplication operation based on the preprocessed target matrix and the data packet to obtain the target data; Among them, when the target data packet is an information packet to be encoded, the target matrix is ​​a generation matrix in the encoding process, and the target data is the encoded data; when the target data packet is an information packet to be decoded, the target matrix is ​​a decoding matrix in the decoding process; the target data is the decoded data.

9. The RS code compilation system based on data stream according to claim 8, characterized in that: It also includes a PCIe module, and the data input module batches out k target data packets from the host machine through the PCIe module in the form of a DMA engine. The data input module is also used to configure the bit width stride of each element of the column vector, and when the bit width stride is not an integer multiple of the fixed bit width w during the matrix-vector multiplication operation, the element corresponding to the last data packet in the column vector is filled with 0.

10. The RS code compilation system based on data stream according to claim 8, characterized in that: It also includes a calculation buffer module for temporarily storing part of the data of the accumulation operation during the matrix-vector multiplication operation. The calculation buffer module and the calculation array module form a ring data flow channel.

Citation Information

Patent Citations

  • Traffic classification method, apparatus and device, and computer readable storage medium

    CN115879019A

  • Code generating, coding and decoding method and device

    CN117271199A

  • Point cloud compression method, point cloud decompression method and point cloud compression device

    CN117409094A

  • Error correction code techniques for matrices with interleaved codewords

    US20130047055A1

  • Error detecting and correcting method and system

    US5563894A