GF(2) Matrix Gaussian Elimination Device, Method, System, Equipment and Medium

By dividing the input matrix into column blocks and into multiple Bigsteps, combined with the operation of exchange, exclusive or and pass-through, the calculation bottleneck problem of Gaussian elimination of the GF(2) matrix is solved, which improves computing efficiency and reduces resource usage, and is suitable for data processing in electronic devices.

CN115329262BActive Publication Date: 2025-07-11TSINGHUA UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210991312.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-17
Publication Date
2025-07-11
Estimated Expiration
2042-08-17

AI Technical Summary

Technical Problem

In the prior art, the GF(2) matrix Gaussian elimination has a long calculation time and a high resource occupancy rate in the McEliece algorithm, which has become a bottleneck in key generation and affects its wide application.

Method used

The input matrix is divided into column blocks, and the Gaussian elimination process is divided into multiple Bigsteps. The calculation array and operation memory are used for playback calculations. Through exchange, XOR and pass-through operations, efficient elimination of column blocks is achieved.

Benefits of technology

It improves the computing efficiency of Gaussian elimination of GF(2) matrix, reduces the clock overhead and computing memory requirements for pipeline startup, and is suitable for data processing in electronic devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115329262B_ABST
    Figure CN115329262B_ABST
Patent Text Reader

Abstract

The present invention provides a GF(2) matrix Gaussian elimination device, which is applied to the field of data processing technology. The device includes: a partitioning unit, configured to partition an input matrix into at least one column block by columns, and divide the process of GF(2) matrix Gaussian elimination into multiple Bigsteps, where one Bigstep corresponds to the calculation process of one column block; a data memory, configured to store the column block; a calculation array, which includes at least one row of calculation units, and the row of calculation units is configured to read the data included in the column block from the data memory and perform a calculation operation on the data to obtain the final calculation result of the column block; and an operation memory, configured to store the operation information. The present invention also provides a GF(2) matrix Gaussian elimination method, system, electronic device, and storage medium, which can save the clock overhead of pipeline startup and save calculation memory at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a GF(2) matrix Gaussian elimination device, method, system, electronic device and medium. Background Art

[0002] McEliece is the only surviving code-based key encapsulation protocol in the third round of the NIST Post-Quantum Cryptography Algorithm Competition. For McEliece, the key generation speed is hundreds or thousands of times slower than encryption and decryption. Therefore, if the speed of the key generation phase can be accelerated, it is of great significance for the wide application of McEliece.

[0003] Large-scale GF(2) matrix Gaussian elimination is the main computational bottleneck of McEliece. Existing technologies have a high resource occupancy rate and a long computational time for large-scale GF(2) matrix Gaussian elimination. Summary of the Invention

[0004] The main objective of the present invention is to provide a GF(2) matrix Gaussian elimination device, method, system, electronic device and medium.

[0005] To achieve the above objective, a first aspect of an embodiment of the present invention provides a GF(2) matrix Gaussian elimination device, including:

[0006] For performing GF(2) matrix Gaussian elimination, the device includes:

[0007] A partitioning unit, configured to partition an input matrix into at least one column block by columns, and divide the process of GF(2) matrix Gaussian elimination into multiple Bigsteps, where one Bigstep corresponds to the calculation process of one column block;

[0008] A data memory, configured to store the column block;

[0009] A calculation array, the calculation array includes at least one row of calculation units, the row of calculation units is configured to read the data included in the column block from the data memory, and perform a calculation operation on the data to obtain a final calculation result of the column block. Wherein, in the case of replaying the calculation for the Bigstep corresponding to the column block, read the operation information stored in the operation memory to implement the replay, and apply the calculation result of the historical Bigstep to the column block; in the case of performing the elimination calculation of the column block, reduce the column block to the form of an identity matrix, and write the operation information corresponding to the elimination calculation into the operation memory;

[0010] The operation memory, configured to store the operation information.

[0011] In one embodiment of the present invention, the types of calculation operations executable by the calculation unit row include a swap type, an exclusive - OR type, and a pass - through type;

[0012] The swap type is used to indicate outputting the data within the calculation unit row to the next calculation unit row, and storing the data input to the calculation unit row into the calculation unit row;

[0013] The exclusive - OR type is used to indicate performing an exclusive - OR operation on the data input to the calculation unit row and the data within the calculation unit row to obtain an exclusive - OR result, and outputting the exclusive - OR result to the next calculation unit row;

[0014] The pass - through type is used to indicate outputting the data within the calculation unit row to the next calculation unit row.

[0015] In one embodiment of the present invention, the calculation unit row includes:

[0016] A data register for storing the main row of the calculation unit row;

[0017] A padding bit register for storing the padding bits of the data input to the calculation unit row when the type of the calculation operation performed by the calculation unit row is the swap type, where the padding bits are used to indicate auxiliary information related to the data, and the auxiliary information is used to determine the type of calculation operation to be performed by the calculation unit row;

[0018] A flag register for indicating that the calculation unit row has stored the main row when the flag bit is valid.

[0019] In one embodiment of the present invention, the data register is further used for:

[0020] When the data stored in itself is not the main row, outputting the data stored in itself to the next calculation unit row, and storing target data;

[0021] Wherein, when the calculation unit row where the data register is located is the first row of the calculation array, the target data is the data read from the data memory, and when the calculation unit row where the data register is located is not the first row of the calculation array, the target data is the data output from the previous calculation unit row of the calculation unit row where the data register is located.

[0022] In one embodiment of the present invention, assuming the size of the input matrix is L rows and K columns, and the bit - width of the number corresponding to each row is n, then dividing the input matrix into at least one column block by columns includes:

[0023] Dividing the input matrix of size L * K into K / n column blocks by columns.

[0024] In one embodiment of the present invention, the i-th Bigstep corresponds to m phases, where m = min(i, L / n).

[0025] In one embodiment of the present invention, the data input to the computing unit row includes diagonal bits, and the computing unit row is used for:

[0026] In the case of being in the last phase of the first L / n Bigsteps, read the diagonal bits of the data input to the computing unit row, determine the type of calculation operation to be performed by the computing unit row based on the flag bit of the computing unit row and the diagonal bits of the data input to the computing unit row, and store the diagonal bits of the data written to the computing unit row into the operation memory;

[0027] In the case of being in a non-last phase of a Bigstep, read the corresponding bit from the operation memory, and determine the type of calculation operation to be performed by the computing unit row based on the flag bit of the computing unit row and the corresponding bit.

[0028] In one embodiment of the present invention, determining the type of calculation operation to be performed by the computing unit row based on the flag bit of the computing unit row and the diagonal bits of the data input to the computing unit row includes:

[0029] In the case where the flag bit of the computing unit row is invalid and the diagonal bits of the data input to the computing unit row are valid, perform the swap-type operation, and set the flag bit of the computing unit row to be valid;

[0030] In the case where both the flag bit of the computing unit row and the diagonal bits of the data input to the computing unit row are invalid, perform the straight-through type operation;

[0031] In the case where both the flag bit of the computing unit row and the diagonal bits of the data input to the computing unit row are valid, perform the exclusive-OR type operation;

[0032] In the case where the flag bit of the computing unit row is valid and the diagonal bits of the data input to the computing unit row are invalid, perform the straight-through type operation;

[0033] Determining the type of calculation operation to be performed by the computing unit row based on the flag bit of the computing unit row and the corresponding bit includes:

[0034] In the case where the flag bit of the computing unit row is invalid and the corresponding bit is valid, perform the swap-type operation, and set the flag bit of the computing unit row to be valid;

[0035] When the flag bit of the computing unit row and the corresponding bit are both invalid, perform the operation of the pass-through type;

[0036] When the flag bit of the computing unit row and the corresponding bit are both valid, perform the operation of the exclusive-or type;

[0037] When the flag bit of the computing unit row is valid and the corresponding bit is invalid, perform the operation of the pass-through type.

[0038] In an embodiment of the present invention, the computing unit row is used for:

[0039] Determine whether the target bit position in the padding bits of the data input to the computing unit row is the same as the target bit position in the padding bits stored in the padding bit register of the computing unit row, where the target bit positions include the parity bit of the data from the phase stage and the parity bit of the data from the Bigstep stage;

[0040] When the target bit position in the padding bits of the data input to the computing unit row is not the same as the target bit position in the padding bits stored in the padding bit register of the computing unit row, perform the operation of the swap type.

[0041] In an embodiment of the present invention, the padding bits are six bits, and each bit of the padding bits is used to represent an auxiliary information related to the data. The six types of auxiliary information include:

[0042] Whether the data is from the last phase of the current Bigstep;

[0043] Whether the data is from the lower triangular part of the input matrix;

[0044] Whether the data is valid;

[0045] Whether the data is from the end phase;

[0046] The parity bit of the data from the phase stage;

[0047] The parity bit of the data from the Bigstep stage.

[0048] In an embodiment of the present invention, the operating memory includes a first operating memory and a second operating memory, and the second operating memory includes an upper triangle and a lower triangle;

[0049] When all the data stored in the computing array is from the last phase of a Bigstep, the computing array is used to write n-bit operation information into the first operating memory in each clock cycle;

[0050] When the data stored in the computing array all come from a non - last phase of a Bigstep, the computing array is used to read n - bit operation information from the first operation memory in each clock cycle;

[0051] When the state of the computing array is switched between two phases of a Bigstep and neither of the two phases is the last phase of the Bigstep, the computing array is used to read n - bit operation information from the first operation memory in each clock cycle;

[0052] When the state of the computing array is switched from a non - last phase of a Bigstep to the last phase, the computing array is used to read the upper triangle in the second operation memory in the non - last phase and write the upper triangle and the diagonal bits of the data of all the computing unit rows in the last phase into the first operation memory together;

[0053] When the state of the computing array is in the transition state from the last phase of a Bigstep to the first phase of the next Bigstep and the previous column block of the transition state is within the range of a square matrix with the same number of rows as the input matrix, the computing array is used to write the diagonal bits of the data written into the computing unit rows in the last phase into a row of the upper triangle of the second operation memory and read the operation information in the same row of the lower triangle of the second operation memory, and the operation information is used for the replay operation in the first phase;

[0054] When the state of the computing array is in the transition state from the last phase of a Bigstep to the first phase of the next Bigstep and both column blocks involved in the transition state are outside the range of a square matrix with the same number of rows as the input matrix, the computing array is used to read a row of the second operation memory, the lower triangle of a row of the second operation memory includes the operation information of the first phase, and the upper triangle of a row of the second operation memory includes the operation information of the last phase;

[0055] When the state of the computing array is the startup state, the computing array is used to write operation information into the lower triangle of the second operation memory;

[0056] When the state of the computing array is the end state, the computing array is used to read the operation information of the upper triangle of the second operation memory.

[0057] In an embodiment of the present invention, the computing array is further configured to determine that the input matrix is a singular matrix if, in any Bigstep, there is any row of computing units that cannot find a main row.

[0058] A second aspect of the embodiments of the present invention provides a GF(2) matrix Gaussian elimination method, which is applied to a GF(2) matrix Gaussian elimination device. The GF(2) matrix Gaussian elimination device includes a computing array, and the computing array includes at least one row of computing units. The method includes:

[0059] Partition the input matrix into at least one column block by columns, and divide the process of GF(2) matrix Gaussian elimination into multiple Bigsteps, where one Bigstep corresponds to the calculation process of one column block;

[0060] Use the row of computing units to read the data included in the column block, and perform a calculation operation on the data to obtain the result of GF(2) matrix Gaussian elimination of the input matrix;

[0061] Among them, in the case of replaying the calculation for the Bigstep corresponding to the column block, read the operation information stored in the operation memory to implement the replay, and apply the calculation results of the historical Bigsteps to the column block; in the case of performing the elimination calculation for the column block, finally eliminate the column block into the form of an identity matrix, and write the operation information corresponding to the elimination calculation into the operation memory.

[0062] In an embodiment of the present invention, the types of calculation operations that the row of computing units can execute include a swap type, an exclusive OR type, and a pass-through type;

[0063] The swap type is used to indicate outputting the data in the row of computing units to the next row of computing units, and storing the data input to the row of computing units in the row of computing units;

[0064] The exclusive OR type is used to indicate performing an exclusive OR on the data input to the row of computing units and the data in the row of computing units to obtain an exclusive OR result, and outputting the exclusive OR result to the next row of computing units;

[0065] The pass-through type is used to indicate outputting the data in the row of computing units to the next row of computing units.

[0066] In an embodiment of the present invention, the row of computing units includes:

[0067] A data register for storing the main row of the row of computing units;

[0068] A padding bit register, configured to store padding bits of data input to the row of computing units when the type of computing operation performed by the row of computing units is the swapping type, where the padding bits are used to indicate auxiliary information related to the data, and the auxiliary information is used to determine the type of computing operation to be performed by the row of computing units;

[0069] A flag register, configured to indicate that the row of computing units has stored the main row when the flag bit is valid.

[0070] A third aspect of an embodiment of the present invention provides a data processing system, including:

[0071] A constant-time sorting device, a data preprocessing device, and a GF(2) matrix Gaussian elimination device as described in the first aspect;

[0072] The constant-time sorting device includes a storage unit, a first-in first-out memory FIFO, and a sorting unit. The storage unit is configured to store data to be processed, where the data to be processed includes a first part of data and a second part of data, and the number of elements in the first part of data is equal to the number of elements in the second part of data. The FIFO includes a first FIFO and a second FIFO. The first FIFO is configured to read the first part of data, and the second FIFO is configured to read the second part of data. The sorting unit is configured to perform an internal sorting on the first part of data and the second part of data respectively in the case of the first iteration, and in the case of a non-first iteration, use a merge sorting method to sort the first part of data and the second part of data to obtain a plurality of intermediate results, and input the plurality of intermediate results as the data to be processed into the storage unit until a final sorting result is obtained, where the number of elements in the intermediate results after each iteration is twice the number of elements in the intermediate results after the previous iteration;

[0073] The data preprocessing device is configured to preprocess the final sorting result to obtain an input matrix, and send the input matrix to the GF(2) matrix Gaussian elimination device as described in the first aspect.

[0074] A fourth aspect of an embodiment of the present invention provides an electronic device, including:

[0075] A processor; and

[0076] The GF(2) matrix Gaussian elimination device as described in the first aspect, where the GF(2) matrix Gaussian elimination device is electrically connected to the processor.

[0077] A fourth aspect of an embodiment of the present invention provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the method according to the second aspect.

[0078] According to an embodiment of the present invention, the provided GF(2) matrix Gaussian elimination device, method, system, electronic device and medium, the GF(2) matrix Gaussian elimination device is used to perform GF(2) matrix Gaussian elimination, and the device includes: a partitioning unit, configured to partition an input matrix into at least one column block by columns, and divide the process of GF(2) matrix Gaussian elimination into multiple Bigsteps, one Bigstep corresponding to the calculation process of one column block; a data memory, configured to store the column block, a calculation array, the calculation array includes at least one row of calculation units, the row of calculation units is configured to read the data included in the column block from the data memory, and perform a calculation operation on the data to obtain the final calculation result of the column block. Wherein, in the case of performing replay calculation for the Bigstep corresponding to the column block, read the operation information stored in the operation memory to implement replay, and apply the calculation result of the historical Bigstep to the column block. In the case of performing elimination calculation for the column block, finally eliminate the column block into the form of an identity matrix, and write the operation information corresponding to the elimination calculation into the operation memory, and the operation memory is configured to store the operation information. It can save the clock overhead of pipeline startup and save calculation memory at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.

[0080] Figure 1 It is a schematic structural diagram of a GF(2) matrix Gaussian elimination device provided by an embodiment of the present invention;

[0081] Figure 2 It is a schematic working flow diagram of a GF(2) matrix Gaussian elimination device provided by an embodiment of the present invention;

[0082] Figure 3 It is a schematic structural diagram of a calculation array provided by an embodiment of the present invention;

[0083] Figure 4 It is a schematic working principle diagram of an operation memory provided by an embodiment of the present invention;

[0084] Figure 5Schematic flowchart of the GF(2) matrix Gaussian elimination method provided by an embodiment of the present invention;

[0085] Figure 6 Schematic structural diagram of a data processing system provided by an embodiment of the present invention;

[0086] Figure 7 Shows a schematic hardware structure diagram of an electronic device. Detailed implementation manners

[0087] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0088] The present invention provides a GF(2) matrix Gaussian elimination device, method, system, electronic device, and medium. The GF(2) matrix Gaussian elimination device is used to perform GF(2) matrix Gaussian elimination. The device includes: a partitioning unit for partitioning an input matrix into at least one column block by columns and dividing the process of GF(2) matrix Gaussian elimination into multiple Bigsteps, where one Bigstep corresponds to the calculation process of one column block; a data memory for storing the column block; a computing array including at least one row of computing units for reading the data included in the column block from the data memory and performing a computing operation on the data to obtain the final calculation result of the column block. Among them, when replay calculation is performed for the Bigstep corresponding to the column block, the operation information stored in the operation memory is read to achieve replay, and the calculation results of historical Bigsteps are applied to the column block; when elimination calculation of the column block is performed, the column block is finally eliminated into the form of an identity matrix, and the operation information corresponding to the elimination calculation is written into the operation memory; the operation memory for storing the operation information. It can save the clock overhead of pipeline startup and save computing memory at the same time.

[0089] The following will describe in detail some embodiments of the present invention in conjunction with the accompanying drawings. Without conflict between the embodiments, the following embodiments and the features in the embodiments can be combined with each other.

[0090] Please refer to Figure 1 , Figure 1 Schematic structural diagram of a GF(2) matrix Gaussian elimination device provided by an embodiment of the present invention. The GF(2) matrix Gaussian elimination device mainly includes: a partitioning unit, a data memory, a computing array, and an operation memory.

[0091] A partitioning unit for partitioning an input matrix into at least one column block by columns, and dividing the process of Gaussian elimination of the GF(2) matrix into multiple Bigsteps, where one Bigstep corresponds to the calculation process of one such column block.

[0092] A data memory for storing the column block.

[0093] A computing array, the computing array includes at least one row of computing units, the row of computing units is used to read the data included in the column block from the data memory, and perform a computing operation on the data to obtain the final computing result of the column block. Wherein, in the case of performing replay calculation for the Bigstep corresponding to the column block, read the operation information stored in the operation memory to implement replay, and apply the calculation results of the historical Bigstep to the column block; in the case of performing elimination calculation for the column block, finally eliminate the column block into the form of an identity matrix, and write the operation information corresponding to the elimination calculation into the operation memory.

[0094] The operation memory for storing the operation information.

[0095] In the present invention, replay means splitting a matrix into several column blocks. One column block has been eliminated, but the elimination effect is only limited to that column block. The column blocks behind this column block need to replay the elimination operation of this column block to ensure that all column blocks behind this column block have completed the same operation.

[0096] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the working process of a GF(2) matrix Gaussian elimination device provided by an embodiment of the present invention. As Figure 2 shown, let the size of the input matrix be L rows and K columns, and the bit width of the data corresponding to each row be n. Then the partitioning unit partitions the input matrix of size L*K into K / n column blocks (that is, K / n Bigsteps) by columns. The size of the column block is Lxn, and the column block is stored in the data memory as the input of the computing array. The computing array writes operation information to the operation memory and reads operation information from the operation memory during the execution of GF(2) Gaussian elimination to implement replay.

[0097] In the present invention, the i-th Bigstep corresponds to m phases, m = min(i, L / n). In each phase, there are n cycles, and the length of the input data in each cycle is L.

[0098] In the present invention, for the computational logic of block division, the computational information obtained by eliminating a column block is stored in the corresponding operating memory. Then, when other block computations are performed, the operating information stored in the operating memory is replayed to achieve the effect of eliminating the entire row. For the column blocks within the LxL square matrix, first, the information of the operating memory recorded by the previous Bigstep is used to replay the computations of each previous Bigstep in sequence (replay the computations according to the operating information stored in each bigstep), and then the column block is eliminated (this part of the operation corresponds to the operating information of this column block), and the related operating messages are recorded and stored in an additional operating memory for the subsequent Bigstep to replay the computations. For the column blocks in the LxK matrix that are not within the above-mentioned LxL square matrix, the replay process is the same, that is, replaying the operations corresponding to all bigsteps in the LxL square matrix. Therefore, after each column block completes its elimination operation, there is no need to store the corresponding operating information.

[0099] According to an embodiment of the present invention, based on the block-based Gaussian elimination calculation, the entire Gaussian elimination process is divided into multiple Bigsteps. One Bigstep corresponds to the calculation of one column block. When the calculation of a certain column block is completely finished, the final calculation result of this column block can be directly obtained. Compared with the prior art where multiple rounds of iteration are required to sequentially obtain the final calculation results of each column block in the last round of iteration, the calculation efficiency can be greatly improved.

[0100] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a computing array provided by an embodiment of the present invention.

[0101] As Figure 3 shown, the computing array includes at least one row of computing units. Once all the rows of computing units are started, they will always be in a computing state, which can save the clock overhead of starting the pipeline. The row of computing units includes: a data register, a padding bit register, and a flag register.

[0102] The data register is used to store the main row of this row of computing units.

[0103] The padding bit register is used to store the padding bits of the data input to this row of computing units when the type of the computing operation performed by this row of computing units is the swap type. The padding bits are used to indicate the auxiliary information related to the data, and the auxiliary information is used to determine the type of computing operation that this row of computing units needs to perform.

[0104] The flag register is used to indicate that this row of computing units has stored the main row when the flag bit is valid. Among them, in the initial state, the flag register bit is 0.

[0105] In the present invention, the main row means that when eliminating a column block, each computing unit row needs to store its corresponding main row in the data register, and then perform elimination between this main row and the subsequent input rows to eliminate all the diagonal bits of the subsequent input rows to 0 (using straight-through operations and exclusive-OR operations).

[0106] In an embodiment of the invention, the types of computing operations executable by the computing unit row include a swap type, an exclusive-OR type, and a straight-through type. The swap type is used to indicate outputting the data within the computing unit row to the next computing unit row and storing the data input to the computing unit row in the computing unit row.

[0107] The exclusive-OR type is used to indicate performing an exclusive-OR operation on the data input to the computing unit row and the data within the computing unit row to obtain an exclusive-OR result and outputting the exclusive-OR result to the next computing unit row. The straight-through type is used to indicate outputting the data within the computing unit row to the next computing unit row.

[0108] As Figure 3 shown, each column block contains n computing unit rows Pline. The design of the computing array is oriented towards the computing unit rows, and the computing array contains n computing unit rows. Each computing unit row further contains three types of registers: an n-bit wide data register (corresponding to Rline in the figure), a 6-bit padding bit register (corresponding to rPads in the figure), and a 1-bit flag register (corresponding to rPivot in the figure).

[0109] In an embodiment of the present invention, the data register is further used for: when the data stored in itself is not the main row, outputting the data stored in itself to the next computing unit row and storing the target data. Wherein, when the computing unit row where the data register is located is the first row of the computing array, the target data is the data read from the data memory, and when the computing unit row where the data register is located is not the first row of the computing array, the target data is the data output from the previous computing unit row of the computing unit row where the data register is located.

[0110] In each phase, the n-bit wide data register finally stores the main row of each computing unit row. If it is found that the data stored in the data register is not the main row, it is necessary to output the data stored in the data register to the next row of computing unit rows, and then store the data output from the previous computing unit row (i.e., the output vector of the previous computing unit row, which is also the input vector of the current computing unit row) or the data read from the data memory (read from the data memory when the current computing unit row is the first row, and receive the output of the previous computing unit row if it is not the first row) into the data register of the computing unit row.

[0111] In an embodiment of the invention, after data is read from the data memory, six-bit padding bits are appended to the data. These six-bit padding bits are used to represent auxiliary information related to the data. Each bit represents a type of auxiliary information. The six types of auxiliary information include whether the data is from the last phase of the current Bigstep (each replay in a Bigstep is called a phase, and the final cancellation operation is the last phase of the Bigstep), whether the data is from the lower triangular part of the input matrix, whether the data is valid, whether the data is from the end phase, the parity bit of the data from the phase, and the parity bit of the data from the Bigstep. These auxiliary information will help the computing unit row determine the type of operation to be performed. The types of operations include swap type, exclusive OR type, and pass-through type.

[0112] In the present invention, a one-bit flag register inside the computing unit row is used to indicate whether the computing unit row stores the main row.

[0113] In an embodiment of the present invention, the data input to the computing unit row includes diagonal bits. The diagonal bit of the nth computing unit row is the nth bit of the input data. The computing unit row is used to, when in the last phase of the first L / n Bigsteps, read the diagonal bits of the data input to the computing unit row, determine the type of computing operation to be performed by the computing unit row based on the flag bit of the computing unit row and the diagonal bits of the data input to the computing unit row, and store the diagonal bits of the data written to the computing unit row in the operation memory for the replay calculation of subsequent Bigsteps. The computing unit row is also used to, when in a non-last phase of a Bigstep, read the corresponding bit from the operation memory, and determine the type of computing operation to be performed by the computing unit row based on the flag bit of the computing unit row and the corresponding bit. In this way, the change pattern of the flag register in the computing unit row is exactly the same as the change pattern of the replayed Bigstep, achieving the cancellation effect in the sense of the entire row.

[0114] In the present invention, the corresponding bits are stored according to cycles (each time data is input to the computing array is recorded as a cycle). In each cycle, one computing unit row corresponds to one corresponding bit. The n bits corresponding to the n computing unit rows are stored in one row and can be stored in the operation memory according to the row number or other methods. The present invention does not limit this. For example, if stored according to the row number, when the computing unit row reads the corresponding bit, it reads the bit corresponding to the row number of the computing unit row from the operation memory.

[0115] In an embodiment of the present invention, when in the last phase of Bigstep, the types of calculation operations to be performed on the calculation unit row determined based on the flag bit of the calculation unit row and the diagonal bit of the data input to the calculation unit row include the following Classifications 1 to 4.

[0116] Classification 1: When the flag bit of the calculation unit row is invalid and the diagonal bit of the data input to the calculation unit row is valid, perform the swap type of operation, and set the flag bit of the calculation unit row to valid. That is, if the flag bit is invalid, determine whether the diagonal bit in the input data is valid. If it is valid, store the value read from the data into the n-bit wide data register of the calculation unit row (swap operation), and set the flag register to valid (that is, the calculation unit row has found its main row, and the data currently stored in the data register is the main row of the calculation unit row).

[0117] Classification 2: When both the flag bit of the calculation unit row and the diagonal bit of the data input to the calculation unit row are invalid, perform the pass-through type of operation. That is, if the flag bit is invalid and the diagonal bit of the input data is invalid (indicating that the data in the current data register is not the main row), directly output the data stored in the data register to the next calculation unit row (pass-through operation).

[0118] Classification 3: When both the flag bit of the calculation unit row and the diagonal bit of the data input to the calculation unit row are valid, perform the exclusive OR type of operation. That is, if both the flag bit and the diagonal bit are valid (indicating that the data in the current data register is not the main row), perform an exclusive OR operation on the input data and the vector stored in the data register.

[0119] Classification 4: When the flag bit of the calculation unit row is valid and the diagonal bit of the data input to the calculation unit row is invalid, perform the pass-through type of operation. That is, if the flag bit is valid and the diagonal bit of the input data is invalid (indicating that the data in the current data register is not the main row), directly pass through the input data and output it to the next calculation unit row.

[0120] In an embodiment of the present invention, when in a non - last phase of Bigstep, the determining of the type of calculation operation to be performed on the calculation unit row based on the flag bit of the calculation unit row and the corresponding bit includes: when the flag bit of the calculation unit row is invalid and the corresponding bit is valid, performing the swap - type operation and setting the flag bit of the calculation unit row to valid; when both the flag bit of the calculation unit row and the corresponding bit are invalid, performing the pass - through type operation; when both the flag bit of the calculation unit row and the corresponding bit are valid, performing the exclusive - OR type operation; when the flag bit of the calculation unit row is valid and the corresponding bit is invalid, performing the pass - through type operation. Details not described in this embodiment can be explained by referring to the embodiment in the case of being in the last phase of Bigstep as described above.

[0121] In an embodiment of the present invention, the calculation unit row is further configured to determine whether the target bit positions in the padding bits of the data input to the calculation unit row are the same as the target bit positions in the padding bits stored in the padding bit register of the calculation unit row. The target bit positions include the check bits of the data from the phase stage and the check bits of the data from the Bigstep stage. When the target bit positions in the padding bits of the data input to the calculation unit row are not the same as the target bit positions in the padding bits stored in the padding bit register of the calculation unit row, the swap - type operation is performed. That is, during the calculation process of the calculation unit row, it is necessary to determine whether the check bits of each input data are the same as the check bits stored in the padding bit register. If they are not the same, it indicates that a phase switch has occurred. At this time, the data in the data register of the calculation unit row needs to be output to the next calculation unit row, and then the input data is stored in the data register (swap - type operation). The check bits in the padding bit register include the information of the Bigstep and phase where the data vector (target bit position) is located. These two bits can be represented by the value of the sequence number of both modulo 2. Then, if these two padding bits of the input data and these two padding bits in the calculation unit row are not exactly the same, it means that these two data come from different phases (i.e., a phase switch has occurred).

[0122] Please refer to Figure 4 , Figure 4 which is a schematic diagram of the working principle of the operating memory provided by an embodiment of the present invention.

[0123] As Figure 4 shown, the operating memory includes a first operating memory and a second operating memory, and the second operating memory includes an upper triangle and a lower triangle. Figure 4The part on the left shows various situations during the entire Gaussian elimination process of the GF(2) matrix. The entire elimination process can be divided into multiple stages.

[0124] Stage case_a: When the data stored in the computing array all come from the last phase of a Bigstep, the computing array is used to write n-bit operation information into the first operation memory in each clock cycle.

[0125] Stage case_b: When the data stored in the computing array all come from a non-last phase of a Bigstep, the computing array is used to read n-bit operation information from the first operation memory in each clock cycle.

[0126] Stage case_c: When the state of the computing array is switching between two phases of a Bigstep, and neither of the two phases is the last phase of the Bigstep, the computing array is used to read n-bit operation information from the first operation memory in each clock cycle.

[0127] Stage case_d: When the state of the computing array is switching from a non-last phase of a Bigstep to the last phase, the computing array is used to read the upper triangle in the second operation memory during the non-last phase and write the upper triangle together with the diagonal bits of the data of all the computing unit rows during the last phase into the first operation memory.

[0128] Stage case_e1: When the state of the computing array is in the transition state of migrating from the last phase of a Bigstep to the first phase of the next Bigstep, and the previous column block of the transition state is within the range of the square matrix with the same number of rows as the input matrix, the computing array is used to write the diagonal bits of the data written into the computing unit rows during the last phase into a row of the upper triangle of the second operation memory, and read the operation information in the same row of the lower triangle of the second operation memory. The operation information is used for the replay operation of the first phase.

[0129] Stage case_e2: When the state of the computing array is in the transition state of migrating from the last phase of a Bigstep to the first phase of the next Bigstep, and both of the two column blocks involved in the transition state are outside the range of the square matrix with the same number of rows as the input matrix, the computing array is used to read a row of the second operation memory. The lower triangle of a row of the second operation memory includes the operation information of the first phase, and the upper triangle of a row of the second operation memory includes the operation information of the last phase.

[0130] Stage case_f: When the state of the computing array is the startup state, the computing array is used to write operation information into the lower triangle of the second operation memory.

[0131] Stage case_g: When the state of the computing array is the end state, the computing array is used to read the operation information of the upper triangle of the second operation memory.

[0132] In an embodiment of the present invention, the computing array is further configured to determine that the input matrix is a singular matrix if any row of computing units cannot find a main row in any Bigstep. After detecting that the input matrix is a singular matrix, the calculation is terminated, which can save unnecessary clock overhead. Compared with the prior art, a singular matrix can be found with less overhead.

[0133] In the present invention, if the input matrix is a singular matrix, then when calculating in the last phase of a certain Bigstep, not all rows of computing units can find their own main rows (the diagonal bit is 1). Specifically, if the data of the input computing unit row already belongs to the lower triangular matrix or comes from the next Bigstep or phase, and the data in the current data register still belongs to the last phase of the current bigstep and the flag register is still not valid, it means that the computing unit row cannot find its own main row.

[0134] The present invention is described below with a specific example:

[0135] Taking L = 2, K = 4, and the input matrix The computing array includes two rows of computing unit rows, and the bit width n of the data vector corresponding to each row is 2 bits. For example, the entire elimination process is divided into 2 Bigsteps (K / n = 2), and the first Bigstep corresponds to the first two columns corresponding to one Phase, and the second BigStep corresponds to the last two columns This Bigstep corresponds to one Phase.

[0136] The calculation process of the GF(2) matrix Gaussian elimination device using the present invention is as follows:

[0137] The first Phase of the first Bigstep:

[0138] The first cycle: The input data is "01". When "01" is given to the first computing unit row, since the diagonal bit is 0, the flag register remains 0. Since it is the first calculation, the input vector "01" is stored in the first computing unit row. Since it is the first and also the last phase, the diagonal bit "0x" is written to the operation memory at this moment.

[0139] The second cycle: The array input data is "11". When "11" passes through the first computing row, since the diagonal bit is 1 and the flag register is 0 (the flag bit is invalid, the diagonal bit is valid), the computing unit row performs an operation of the swap type. At this moment, the data "01" in the data vector register of the first computing unit row is output to the second computing row, the flag register of the first computing unit row becomes 1, and the input data "11" enters the data vector register of the first computing unit row. For the second computing unit row, the input vector is "01" and the diagonal bit is 1. Since the initial state of the flag register of the second computing unit row is 0 (the flag bit is invalid, the diagonal bit is valid), "01" is stored in the second computing unit row, the flag register of the second computing unit row is changed to 1, and the "01" stored in the second computing unit row is output, that is, the array output is "01". At this moment, since it is the last phase, the diagonal bit "11" is written to the operation memory.

[0140] The first Phase of the second Bigstep:

[0141] The first cycle: The input data is "10". The first computing unit row detects that a Bigstep switch has occurred and directly enters the swap state. Then, at this moment, the data "11" in the data vector register of the first computing unit row is output to the second computing row, the input data "10" enters the data vector register of the first computing unit row, and since the read operation information is '0', the flag register becomes 0. For the second computing unit row, the input data is "11", and an exclusive - OR type operation is performed, which is exclusive - ORed with the "01" stored in the data vector register of the second computing unit row, and the exclusive - OR output is "10". The diagonal bit "x1" is written to the operation memory at this moment.

[0142] The second cycle: The input data is "11". Since the current flag register is 0, an operation of the swap type is performed. The originally stored "10" in the data vector register is exported to the next row, the data "11" enters the data vector register, and since the read operation information is "1", the flag register is changed to "1". The second row of computing unit rows receives the input data "10" and the operation information "1", so this input is saved in the data vector register and the flag register is set to valid, and the array output is "10".

[0143] End cycle: The first row of computing units receives an end signal and outputs "11" of the current data vector register to the second row of computing units; for the second row of computing units, upon receiving the operation information "1", an exclusive OR operation is performed to obtain "01", and the array outputs "01".

[0144] Therefore, the final result is to obtain the input matrix of the Gaussian elimination result.

[0145] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of the GF(2) matrix Gaussian elimination method provided by an embodiment of the present invention.

[0146] As Figure 5 shown, the GF(2) matrix Gaussian elimination method is applied to Figures 1 to 4 the GF(2) matrix Gaussian elimination device shown, the GF(2) matrix Gaussian elimination device includes a computing array, the computing array includes at least one row of computing units, and the GF(2) matrix Gaussian elimination method includes operation S510 to operation S520.

[0147] In operation S510, the input matrix is divided into at least one column block by columns, and the process of GF(2) matrix Gaussian elimination is divided into multiple Bigsteps, and one such Bigstep corresponds to the calculation process of one such column block.

[0148] In operation S520, the computing unit row is used to read the data included in the column block and perform a calculation operation on the data to obtain the result of the GF(2) matrix Gaussian elimination of the input matrix.

[0149] Among them, in the case of performing replay calculation for the Bigstep corresponding to the column block, the operation information stored in the operation memory is read to implement replay, and the calculation results of the historical Bigsteps are applied to the column block; in the case of performing elimination calculation for the column block, the column block is finally eliminated into the form of an identity matrix, and the operation information corresponding to the elimination calculation is written into the operation memory.

[0150] In an embodiment of the present invention, the types of calculation operations that the computing unit row can perform include a swap type, an exclusive OR type, and a pass-through type;

[0151] The swap type is used to indicate outputting the data within the computing unit row to the next row of computing units, and storing the data input to the computing unit row into the computing unit row;

[0152] This XOR type is used to indicate that the data input to this row of computing units and the data within this row of computing units are XORed to obtain an XOR result, and this XOR result is output to the next row of computing units;

[0153] This pass-through type is used to indicate that the data within this row of computing units is output to the next row of computing units.

[0154] In an embodiment of the present invention, this row of computing units includes:

[0155] A data register for storing the main row of this row of computing units;

[0156] A padding bit register for storing the padding bits of the data input to this row of computing units when the type of computing operation performed by this row of computing units is this swap type. The padding bits are used to indicate auxiliary information related to the data, and this auxiliary information is used to determine the type of computing operation that this row of computing units needs to perform;

[0157] A flag register for indicating that this row of computing units has stored the main row when the flag bit is valid.

[0158] In an embodiment of the present invention, when the data stored in the data register itself is not the main row, the data stored in the data register is output to the next row of computing units, and the target data is stored;

[0159] Wherein, when the row of computing units where the data register is located is the first row of the computing array, the target data is the data read from the data memory. When the row of computing units where the data register is located is not the first row of the computing array, the target data is the data output from the previous row of computing units of the row of computing units where the data register is located.

[0160] In an embodiment of the present invention, if the size of the input matrix is L rows and K columns, and the bit width of the corresponding quantity for each row is n, then dividing the input matrix into at least one column block by column includes:

[0161] Dividing the input matrix of size L*K into K / n column blocks by column.

[0162] In an embodiment of the present invention, the i-th Bigstep corresponds to m phases, and m = min(i, L / n).

[0163] In an embodiment of the present invention, the data input to this row of computing units includes diagonal bits, and the method further includes:

[0164] In the case of being in the last phase of the first L / n Bigsteps, read the diagonal bits of the data input to this row of computing units, determine the type of computing operation to be performed on this row of computing units based on the flag bits of this row of computing units and the diagonal bits of the data input to this row of computing units, and store the diagonal bits of the data written to this row of computing units in the operation memory;

[0165] In the case of being in a non-last phase of a Bigstep, read the corresponding bit from the operation memory, and determine the type of computing operation to be performed on this row of computing units based on the flag bits of this row of computing units and the corresponding bit.

[0166] In an embodiment of the present invention, the determining the type of computing operation to be performed on this row of computing units based on the flag bits of this row of computing units and the diagonal bits of the data input to this row of computing units includes:

[0167] In the case where the flag bits of this row of computing units are invalid and the diagonal bits of the data input to this row of computing units are valid, perform the operation of the swap type, and set the flag bits of this row of computing units to be valid;

[0168] In the case where the flag bits of this row of computing units and the diagonal bits of the data input to this row of computing units are both invalid, perform the operation of the straight-through type;

[0169] In the case where the flag bits of this row of computing units and the diagonal bits of the data input to this row of computing units are both valid, perform the operation of the exclusive-or type;

[0170] In the case where the flag bits of this row of computing units are valid and the diagonal bits of the data input to this row of computing units are invalid, perform the operation of the straight-through type;

[0171] The determining the type of computing operation to be performed on this row of computing units based on the flag bits of this row of computing units and the corresponding bit includes:

[0172] In the case where the flag bits of this row of computing units are invalid and the corresponding bit is valid, perform the operation of the swap type, and set the flag bits of this row of computing units to be valid;

[0173] In the case where the flag bits of this row of computing units and the corresponding bit are both invalid, perform the operation of the straight-through type;

[0174] In the case where the flag bits of this row of computing units and the corresponding bit are both valid, perform the operation of the exclusive-or type;

[0175] In the case where the flag bits of this row of computing units are valid and the corresponding bit is invalid, perform the operation of the straight-through type.

[0176] In an embodiment of the present invention, the method further includes:

[0177] Determine whether the target bit positions in the padding bits of the data input to this calculation unit row are consistent with the target bit positions in the padding bits stored in the padding bit register of this calculation unit row. The target bit positions include the check bits of the data from the phase stage and the check bits of the data from the Bigstep stage;

[0178] In the case where the target bit positions in the padding bits of the data input to this calculation unit row are inconsistent with the target bit positions in the padding bits stored in the padding bit register of this calculation unit row, perform an operation of the exchange type.

[0179] In an embodiment of the present invention, the padding bit is six bits, and each bit of the padding bit is used to represent an auxiliary information related to the data. The six types of auxiliary information include:

[0180] Whether the data is from the last phase of the current Bigstep;

[0181] Whether the data is from the lower triangular part of the input matrix;

[0182] Whether the data is valid;

[0183] Whether the data is from the end phase;

[0184] The check bits of the data from the phase stage;

[0185] The check bits of the data from the Bigstep stage.

[0186] In an embodiment of the present invention, the operating memory includes a first operating memory and a second operating memory. The second operating memory includes an upper triangle and a lower triangle. The method further includes:

[0187] In the case where the data stored in the calculation array all comes from the last phase of a Bigstep, use the calculation array to write n-bit operation information into the first operating memory in each clock cycle;

[0188] In the case where the data stored in the calculation array all comes from a non-last phase of a Bigstep, use the calculation array to read n-bit operation information from the first operating memory in each clock cycle;

[0189] In the case where the state of the calculation array is switched between two phases of a Bigstep and neither of the two phases is the last phase of the Bigstep, use the calculation array to read n-bit operation information from the first operating memory in each clock cycle;

[0190] When the state of the computing array switches from a non-final phase to the final phase of a Bigstep, the computing array is used to read the upper triangle in the second operation memory during the non-final phase, and write the upper triangle and the diagonal bits of the data of all the computing unit rows during the final phase into the first operation memory;

[0191] When the state of the computing array is in the transition state from the final phase of a Bigstep to the first phase of the next Bigstep, and the previous column block of the transition state is within the range of a square matrix with the same number of rows as the input matrix, the computing array is used to write the diagonal bits of the data written to the computing unit rows during the final phase into a row of the upper triangle of the second operation memory, and read the operation information in the same row of the lower triangle of the second operation memory, where the operation information is used for the replay operation of the first phase;

[0192] When the state of the computing array is in the transition state from the final phase of a Bigstep to the first phase of the next Bigstep, and both of the two column blocks involved in the transition state are outside the range of a square matrix with the same number of rows as the input matrix, the computing array is used to read a row of the second operation memory, where the lower triangle of the row of the second operation memory includes the operation information of the first phase, and the upper triangle of the row of the second operation memory includes the operation information of the final phase;

[0193] When the state of the computing array is the startup state, the computing array is used to write the operation information into the lower triangle of the second operation memory;

[0194] When the state of the computing array is the end state, the computing array is used to read the operation information of the upper triangle of the second operation memory.

[0195] In an embodiment of the present invention, the computing array is further used to determine that the input matrix is a singular matrix if there is any computing unit row that cannot find the main row in any Bigstep.

[0196] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a data processing system provided by an embodiment of the present invention.

[0197] As Figure 6 shown, the data processing system includes:

[0198] A constant-time sorting device 610, a data preprocessing device 620, and as Figures 1 to 6The GF(2) matrix Gaussian elimination device 630 shown;

[0199] The constant-time sorting device 610 includes a storage unit, a first-in first-out memory FIFO, and a sorting unit. The storage unit is used to store data to be processed, and the data to be processed includes a first part of data and a second part of data. The number of elements in the first part of data is equal to the number of elements in the second part of data. The FIFO includes a first FIFO and a second FIFO. The first FIFO is used to read the first part of data, and the second FIFO is used to read the second part of data. The sorting unit is used to perform internal sorting on the first part of data and the second part of data respectively in the case of the first iteration, and in the case of non-first iteration, use the merge sorting method to sort the first part of data and the second part of data to obtain a plurality of intermediate results, and input the plurality of intermediate results as the data to be processed into the storage unit until a final sorting result is obtained. Wherein, the number of elements in the intermediate result after each iteration is twice the number of elements in the intermediate result after the previous iteration;

[0200] The data preprocessing device 620 is used to preprocess the final sorting result to obtain an input matrix, and send the input matrix to the Figures 1 to 4 GF(2) matrix Gaussian elimination device 630 shown.

[0201] According to the embodiment of the present invention, in the block-based Gaussian elimination calculation scheduling, each time the data block involved is only a part of the matrix rather than all of it. Therefore, only this part of the data block needs to be stored on the chip, which greatly reduces the scale of the on-chip memory. Moreover, the calculation scheduling mode of the present invention realizes the parallel execution of the GF(2) Gaussian elimination calculation and the generation of the data matrix. The GF(2) Gaussian elimination of a certain column block and the output of the calculation result of the previous column block and the generation of the matrix of the next column block (executed by the constant-time sorting device) are scheduled to be executed in parallel, so that the time for data generation and the output of the calculation result can be hidden to the greatest extent. In the hardware design, the data memory can include two banks. When the data in one bank is being calculated iteratively, the data in the other bank can output the calculation result and start importing the input data of the next column block at the same time.

[0202] Any combination of a module, sub-module, unit, and sub-unit according to an embodiment of the present invention, or at least some functions of any combination thereof, may be implemented in one module. Any one or more of the modules, sub-modules, units, and sub-units according to an embodiment of the present invention may be split into multiple modules for implementation. Any one or more of the modules, sub-modules, units, and sub-units according to an embodiment of the present invention may be at least partially implemented as a hardware circuit, such as a field-programmable gate array (FPGA), programmable logic array (PLA), system-on-chip, system-on-substrate, system-on-package, application-specific integrated circuit (ASIC), or may be implemented by any other reasonable manner of integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several thereof. Alternatively, one or more of the modules, sub-modules, units, and sub-units according to an embodiment of the present invention may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.

[0203] For example, the partitioning unit and the data memory may be combined and implemented in one module / unit / sub-unit, or any one of the modules / units / sub-units may be split into multiple modules / units / sub-units. Alternatively, at least some functions of one or more of these modules / units / sub-units may be combined with at least some functions of other modules / units / sub-units and implemented in one module / unit / sub-unit. According to an embodiment of the present invention, at least one of the partitioning unit and the data memory may be at least partially implemented as a hardware circuit, such as a field-programmable gate array (FPGA), programmable logic array (PLA), system-on-chip, system-on-substrate, system-on-package, application-specific integrated circuit (ASIC), or may be implemented by any other reasonable manner of integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several thereof. Alternatively, at least one of the partitioning unit and the data memory may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.

[0204] Figure 7 A block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present invention is schematically shown. Figure 7 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0205] As Figure 7As shown, the electronic device 700 according to an embodiment of the present invention includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage section 708 into a random access memory (RAM) 703. The processor 701 can include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), and so on. The processor 701 can also include on-board memory for caching purposes. The processor 701 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0206] In the RAM 703, various programs and data required for the operation of the system 700 are stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The processor 701 performs various operations of the method flow according to an embodiment of the present invention by executing the program in the ROM 702 and / or the RAM 703. It should be noted that the program can also be stored in one or more memories other than the ROM 702 and the RAM 703. The processor 701 can also perform various operations of the method flow according to an embodiment of the present invention by executing the program stored in the one or more memories.

[0207] According to an embodiment of the present invention, the system 700 can further include an input / output (I / O) interface 705, and the input / output (I / O) interface 705 is also connected to the bus 704. The system 700 can further include one or more of the following components connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read from it can be installed into the storage section 708 as needed.

[0208] According to an embodiment of the present invention, the method flow according to the embodiment of the present invention can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, the above functions defined in the system of the embodiment of the present invention are executed. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0209] The present invention also provides a computer-readable storage medium, which can be included in the device / device / system described in the above embodiment; or can exist alone without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present invention is implemented.

[0210] According to an embodiment of the present invention, the computer-readable storage medium can be a non-volatile computer-readable storage medium. For example, it can include but is not limited to: portable computer disks, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, device, or device.

[0211] For example, according to an embodiment of the present invention, the computer-readable storage medium can include the above-described ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703.

[0212] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0213] Those skilled in the art will appreciate that the features recited in the various embodiments and / or claims of the present invention may be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recited in the various embodiments and / or claims of the present invention may be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0214] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in the respective embodiments cannot be used advantageously in combination. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A GF(2) matrix Gaussian elimination device, characterized in that, For performing Gaussian elimination of GF(2) matrices, the apparatus includes: A partitioning unit, configured to partition an input matrix into at least one column block by columns, and divide the process of Gaussian elimination of the GF(2) matrix into multiple Bigsteps, where one Bigstep corresponds to the calculation process of one column block; A data memory, configured to store the column blocks; A computing array, the computing array includes at least one row of computing units, the row of computing units is configured to read the data included in the column block from the data memory, and perform a computing operation on the data to obtain a final computing result of the column block. Wherein, in the case of performing replay calculation for the Bigstep corresponding to the column block, read the operation information stored in the operation memory to implement replay, and apply the computing results of the historical Bigsteps to the column block; in the case of performing elimination calculation on the column block, finally eliminate the column block into the form of an identity matrix, and write the operation information corresponding to the elimination calculation into the operation memory; The operation memory, configured to store the operation information.

2. The GF(2) matrix Gaussian elimination device according to claim 1, wherein The types of computing operations executable by the row of computing units include a swap type, an exclusive OR type, and a pass-through type; The swap type is used to indicate outputting the data within the row of computing units to the next row of computing units, and storing the data input to the row of computing units into the row of computing units; The exclusive OR type is used to indicate performing an exclusive OR operation on the data input to the row of computing units and the data within the row of computing units to obtain an exclusive OR result, and outputting the exclusive OR result to the next row of computing units; The pass-through type is used to indicate outputting the data within the row of computing units to the next row of computing units.

3. The GF(2) matrix Gaussian elimination device according to claim 2, characterized in that, The row of computing units includes: A data register, configured to store the main row of the row of computing units; A padding bit register, configured to store the padding bits of the data input to the row of computing units in the case where the type of computing operation performed by the row of computing units is the swap type, the padding bits are used to indicate auxiliary information related to the data, and the auxiliary information is used to determine the type of computing operation that the row of computing units needs to perform; A flag register, configured to indicate that the row of computing units has stored the main row when the flag bit is valid.

4. The GF(2) matrix Gaussian elimination device according to claim 3, characterized in that, The data register is further configured to: When the data stored in itself is not the main row, output the data stored in itself to the next row of computing units, and store target data; Wherein, when the row of computing units where the data register is located is the first row of the computing array, the target data is the data read from the data memory; when the row of computing units where the data register is located is not the first row of the computing array, the target data is the data output from the previous row of computing units of the row of computing units where the data register is located.

5. The GF(2) matrix Gaussian elimination device according to any one of claims 1 to 4, characterized in that Let the size of the input matrix be L rows and K columns, and the bit width of the corresponding number of each row be n. Then, partitioning the input matrix into at least one column block by columns includes: Partitioning the input matrix of size L*K into K / n column blocks by columns.

6. The GF(2) matrix Gaussian elimination device according to claim 5, characterized in that, The i-th Bigstep corresponds to m phases, where m = min(i, L / n).

7. The GF(2) matrix Gaussian elimination device according to claim 6, characterized in that, The data input to the row of computing units includes diagonal bits, and the row of computing units is used for: In the case of being in the last phase of the first L / n Bigsteps, reading the diagonal bits of the data input to the row of computing units, determining the type of calculation operation to be performed by the row of computing units based on the flag bit of the row of computing units and the diagonal bits of the data input to the row of computing units, and storing the diagonal bits of the data written to the row of computing units into the operation memory; In the case of being in a non-last phase of a Bigstep, reading the corresponding bits from the operation memory, and determining the type of calculation operation to be performed by the row of computing units based on the flag bit of the row of computing units and the corresponding bits.

8. The GF(2) matrix Gaussian elimination device according to claim 7, characterized in that, The determining the type of calculation operation to be performed by the row of computing units based on the flag bit of the row of computing units and the diagonal bits of the data input to the row of computing units includes: In the case where the flag bit of the row of computing units is invalid and the diagonal bits of the data input to the row of computing units are valid, performing an operation of the swap type, and setting the flag bit of the row of computing units to valid; In the case where both the flag bit of the row of computing units and the diagonal bits of the data input to the row of computing units are invalid, performing an operation of the pass-through type; In the case where both the flag bit of the row of computing units and the diagonal bits of the data input to the row of computing units are valid, performing an operation of the exclusive-OR type; In the case where the flag bit of the row of computing units is valid and the diagonal bits of the data input to the row of computing units are invalid, performing the operation of the pass-through type; The determining the type of calculation operation to be performed by the row of computing units based on the flag bit of the row of computing units and the corresponding bits includes: In the case where the flag bit of the row of computing units is invalid and the corresponding bits are valid, performing the operation of the swap type, and setting the flag bit of the row of computing units to valid; In the case where both the flag bit of the row of computing units and the corresponding bits are invalid, performing the operation of the pass-through type; In the case where both the flag bit of the row of computing units and the corresponding bits are valid, performing the operation of the exclusive-OR type; In the case where the flag bit of the row of computing units is valid and the corresponding bits are invalid, performing the operation of the pass-through type.

9. The GF(2) matrix Gaussian elimination device according to claim 5, wherein, The row of computing units is used for: Judging whether the target bit positions in the padding bits of the data input to the row of computing units are consistent with the target bit positions in the padding bits stored in the padding bit register of the row of computing units, where the target bit positions include the parity bits of the data from the phase stage and the parity bits of the data from the Bigstep stage; In the case where the target bit positions in the padding bits of the data input to the row of computing units are inconsistent with the target bit positions in the padding bits stored in the padding bit register of the row of computing units, performing an operation of the swap type.

10. The GF(2) matrix Gaussian elimination device according to claim 5, characterized in that, The padding bits are six bits, and each of the padding bits is used to represent an auxiliary information related to the data. The six types of auxiliary information include: Whether the data comes from the last phase of the current Bigstep; Whether the data comes from the lower triangular part of the input matrix; Whether the data is valid; Whether the data comes from the end phase; The data comes from the parity bit of the phase; The data comes from the parity bit of the Bigstep phase.

11. The GF(2) matrix Gaussian elimination device according to claim 1, characterized in that, The operating memory includes a first operating memory and a second operating memory. The second operating memory includes an upper triangle and a lower triangle; When the data stored in the computing array all comes from the last phase of a Bigstep, the computing array is used to write n-bit operation information into the first operating memory in each clock cycle; When the data stored in the computing array all comes from a non-last phase of a Bigstep, the computing array is used to read n-bit operation information from the first operating memory in each clock cycle; When the state of the computing array is switched between two phases of a Bigstep and neither of the two phases is the last phase of the Bigstep, the computing array is used to read n-bit operation information from the first operating memory in each clock cycle; When the state of the computing array is switched from a non-last phase of a Bigstep to the last phase, the computing array is used to read the upper triangle in the second operating memory in the non-last phase, and write the upper triangle and the diagonal bits of the data of all the computing unit rows in the last phase into the first operating memory together; When the state of the computing array is in a transition state from the last phase of a Bigstep to the first phase of the next Bigstep, and the previous column block of the transition state is within the range of a square matrix with the same number of rows as the input matrix, the computing array is used to write the diagonal bits of the data written into the computing unit rows in the last phase into a row of the upper triangle of the second operating memory, and read the operation information in the same row of the lower triangle of the second operating memory. The operation information is used for the replay operation of the first phase; When the state of the computing array is in a transition state from the last phase of a Bigstep to the first phase of the next Bigstep, and both of the two column blocks involved in the transition state are outside the range of a square matrix with the same number of rows as the input matrix, the computing array is used to read a row of the second operating memory. The lower triangle of a row of the second operating memory includes the operation information of the first phase, and the upper triangle of a row of the second operating memory includes the operation information of the last phase; When the state of the computing array is the startup state, the computing array is used to write operation information into the lower triangle of the second operation memory; When the state of the computing array is the end state, the computing array is used to read the operation information of the upper triangle of the second operation memory.

12. The GF(2) matrix Gaussian elimination device according to claim 1, characterized in that, The computing array is further used to determine that the input matrix is a singular matrix if there is no master row found in any row of the computing units in any Bigstep.

13. A Gaussian elimination method for GF(2) matrices, characterized in that, Applied to a GF(2) matrix Gaussian elimination device, the GF(2) matrix Gaussian elimination device includes a computing array, the computing array includes at least one row of computing units, and the method includes: Dividing the input matrix into at least one column block by columns, and dividing the process of GF(2) matrix Gaussian elimination into multiple Bigsteps, where one Bigstep corresponds to the calculation process of one column block; Using the row of computing units to read the data included in the column block and performing a calculation operation on the data to obtain the result of GF(2) matrix Gaussian elimination of the input matrix; Wherein, when performing replay calculation for the Bigstep corresponding to the column block, reading the operation information stored in the operation memory to implement replay, and applying the calculation result of the historical Bigstep to the column block; when performing elimination calculation on the column block, finally eliminating the column block into the form of an identity matrix, and writing the operation information corresponding to the elimination calculation into the operation memory.

14. The GF(2) matrix Gaussian elimination method according to claim 13, characterized in that The types of calculation operations executable by the row of computing units include a swap type, an exclusive-or type, and a pass-through type; The swap type is used to indicate outputting the data in the row of computing units to the next row of computing units, and storing the data input to the row of computing units into the row of computing units; The exclusive-or type is used to indicate exclusive-or the data input to the row of computing units and the data in the row of computing units to obtain an exclusive-or result, and outputting the exclusive-or result to the next row of computing units; The pass-through type is used to indicate outputting the data in the row of computing units to the next row of computing units.

15. The GF(2) matrix Gaussian elimination method according to claim 14, wherein The row of computing units includes: A data register for storing the master row of the row of computing units; A padding bit register for storing the padding bits of the data input to the row of computing units when the type of the calculation operation executed by the row of computing units is the swap type, the padding bits being used to indicate auxiliary information related to the data, and the auxiliary information being used to determine the type of calculation operation to be executed by the row of computing units; A flag register for indicating that the row of computing units has stored the master row when the flag bit is valid.

16. An electronic device, characterized in that, Includes: A processor; And The GF(2) matrix Gaussian elimination device according to any one of claims 1 to 12, the GF(2) matrix Gaussian elimination device being electrically connected to the processor.

17. A computer-readable storage medium having executable instructions stored thereon, which when executed by a processor cause the processor to execute the method according to any one of claims 13 to 15.

Citation Information

Patent Citations

  • Lost data recovery method and system based on erasure code, terminal and storage medium

    CN111858157A

  • Cryptographic system based on Classic McEliece cryptosystem

    CN114866231A