Constant-Time Sorting Device, Method, System, Electronic Device, and Medium
By using the combination of storage units and sorting units in the McEliece algorithm, combined with FIFO and merge sorting method, the sorting process is optimized, the problems of high resource occupation and long calculation time are solved, and efficient constant time sorting is achieved.
Patent Information
- Application Number
- CN202210985513.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-17
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-08-17
AI Technical Summary
In the prior art, the McEliece algorithm has a high resource occupancy rate and a long calculation time, which has become its main computing bottleneck and affects its wide application.
Using the combination of storage units and sorting units, FIFO and merge sorting methods, the first and second parts of data are sorted internally, and merged and sorted in non-first iterations to obtain multiple intermediate results, and finally the final sorting results are obtained. By keeping the registers controlled data to queue, resource usage and calculation time are optimized.
It reduces resource occupancy, reduces calculation time, improves sorting efficiency, and solves the bottleneck problem of constant time sorting in McEliece algorithm.
Smart Images

Figure CN115310036B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a constant-time sorting device, method, system, electronic device, and medium. Background Art
[0002] McEliece is the only surviving code-based key encapsulation protocol in the third round of the NIST Post-Quantum Cryptography Algorithm Competition. For McEliece, the key generation speed is hundreds or thousands of times slower than encryption and decryption. Therefore, if the speed of the key generation phase can be increased, it is of great significance for the widespread application of McEliece.
[0003] And constant-time sorting is the main computational bottleneck of McEliece. The existing technologies have a high resource occupancy rate for constant-time sorting and also a relatively long computational time. Summary of the Invention
[0004] The main object of the present invention is to provide a constant-time sorting device, method, system, electronic device, and medium.
[0005] To achieve the above object, a first aspect of an embodiment of the present invention provides a constant-time sorting device, including:
[0006] A storage unit, a first-in-first-out memory FIFO, and a sorting unit;
[0007] The storage unit is used to store data to be processed, the data to be processed includes a first part of data and a second part of data, and the number of elements in the first part of data is equal to the number of elements in the second part of data;
[0008] The FIFO includes a first FIFO and a second FIFO. The first FIFO is used to read the first part of data, and the second FIFO is used to read the second part of data;
[0009] The sorting unit is used to perform internal sorting on the first part of data and the second part of data respectively in the case of the first iteration, and in the case of non-first iteration, use the merge sorting method to sort the first part of data and the second part of data to obtain a plurality of intermediate results, and use the plurality of intermediate results as the data to be processed and input them into the storage unit until a final sorting result is obtained;
[0010] Wherein, the number of elements in the intermediate result after each iteration is twice the number of elements in the intermediate result after the previous iteration.
[0011] In an embodiment of the present invention, the storage unit includes:
[0012] A first storage unit, including a first storage subunit and a second storage subunit, where the first storage subunit is used to store the first part of data for odd-numbered iterations, and the second storage subunit is used to store the second part of data for odd-numbered iterations;
[0013] A second storage unit, including a third storage subunit and a fourth storage subunit, where the third storage subunit is used to store the first part of data for even-numbered iterations, and the fourth storage subunit is used to store the second part of data for even-numbered iterations.
[0014] In an embodiment of the present invention, the constant-time sorting device further includes:
[0015] At least one first holding register, the number of the first holding registers being the same as the number of queues in the first FIFO, at least one of the first holding registers corresponding to the queues in the first FIFO one by one, and the first holding register being used to prohibit the elements in the corresponding queue of the first FIFO from dequeuing when the first holding register is valid;
[0016] At least one second holding register, the number of the second holding registers being the same as the number of queues in the second FIFO, at least one of the second holding registers corresponding to the queues in the second FIFO one by one, and the second holding register being used to prohibit the elements in the corresponding queue of the second FIFO from dequeuing when the second holding register is valid.
[0017] In an embodiment of the present invention, both the first FIFO and the second FIFO include N queues, N being an integer greater than 0, and the constant-time sorting device further includes:
[0018] A bias unit, configured to set the first holding register of the corresponding queue to be valid when each element of an intermediate result included in the first part of data has dequeued in the corresponding queue of the first FIFO during each sorting, set the second holding register of the corresponding queue to be valid when each element of an intermediate result included in the second part of data has dequeued in the corresponding queue of the second FIFO during each sorting, and set both the first holding register and the second holding register to be invalid when both the first holding register and the second holding register are valid.
[0019] In an embodiment of the present invention, the constant-time sorting device further includes:
[0020] The first processing module is configured to, when all the queues of the first FIFO have elements at the same length position dequeued each time, subtract one from the effective depth of the first FIFO to obtain the current first effective depth, and determine whether the current first effective depth is lower than a preset threshold. If the current first effective depth is lower than the preset threshold, read the data to be processed in the storage unit, where the same address means that the addresses of the elements in the queue are the same;
[0021] The second processing module is configured to, when all the queues of the second FIFO have elements at the same length position dequeued each time, subtract one from the effective depth of the second FIFO to obtain the current second effective depth, and determine whether the current second effective depth is lower than the preset threshold. If the current second effective depth is lower than the preset threshold, read the data to be processed in the storage unit;
[0022] Wherein, the effective depth indicates the length of the queue.
[0023] In an embodiment of the present invention, the sorting unit includes:
[0024] A comparison subunit is configured to pair and compare the head elements of all the queues in the first FIFO with the head elements of all the queues in the second FIFO according to a specified element pairing relationship, and take the smaller elements from the first FIFO or the second FIFO to obtain a group of smaller elements;
[0025] A sorting subunit is configured to sort all the elements in the group of smaller elements in order of size to obtain an intermediate result.
[0026] In an embodiment of the present invention, the specified element pairing relationship includes:
[0027] The head element of the m-th queue of the first FIFO is paired with the head element of the n-th queue of the second FIFO;
[0028] Wherein, m + n = N + 1, and both m and n are positive integers.
[0029] A second aspect of the embodiments of the present invention provides a writing of data to be processed, where the data to be processed includes a first part of data and a second part of data, and the number of elements in the first part of data is equal to the number of elements in the second part of data;
[0030] Read the first part of data using the first FIFO, and read the second part of data using the second FIFO;
[0031] In the case of the first iteration, the first part of the data and the second part of the data are sorted internally respectively. In the case of a non-first iteration, the merge sort method is used to sort the first part of the data and the second part of the data to obtain a plurality of intermediate results, and the plurality of intermediate results are used as the data to be processed, and the operation of writing the data to be processed is performed until the final sorting result is obtained;
[0032] Wherein, the number of elements in the intermediate result after each iteration is twice the number of elements in the intermediate result after the previous iteration.
[0033] In an embodiment of the present invention, the writing of the data to be processed includes:
[0034] In the case of an odd-numbered iteration, the first part of the data is stored in the first input storage subunit, and the second part of the data is stored in the second input storage subunit;
[0035] In the case of an even-numbered iteration, the first part of the data is stored in the third input storage subunit, and the second part of the data is stored in the fourth input storage subunit.
[0036] In an embodiment of the present invention, the method further includes:
[0037] When all elements of an intermediate result included in the first part of the data have dequeued from a certain queue in the first FIFO during each sorting, the first holding register is set to valid;
[0038] When all elements of an intermediate result included in the second part of the data have dequeued from a certain queue in the second FIFO during each sorting, the second holding register is set to valid;
[0039] When both the first holding register and the second holding register are valid, both the first holding register and the second holding register are set to invalid;
[0040] Wherein, when the first holding register is valid, the dequeue of elements in the first FIFO is prohibited; when the second holding register is valid, the dequeue of elements in the second FIFO is prohibited; when the first holding register is invalid, the dequeue of elements in the first FIFO is allowed; when the second holding register is invalid, the dequeue of elements in the second FIFO is allowed.
[0041] In an embodiment of the present invention, the method further includes:
[0042] When all the elements at the same length position in all the queues of the first FIFO dequeue each time, subtract one from the effective depth of the first FIFO to obtain the current first effective depth, and determine whether the current first effective depth is lower than a preset threshold. If the current first effective depth is lower than the preset threshold, read the data to be processed in the storage unit. The same address means that the addresses of the elements in the queue are the same;
[0043] When all the elements at the same length position in all the queues of the second FIFO dequeue each time, subtract one from the effective depth of the second FIFO to obtain the current second effective depth, and determine whether the current second effective depth is lower than the preset threshold. If the current second effective depth is lower than the preset threshold, read the data to be processed in the storage unit;
[0044] Wherein, the effective depth indicates the length of the queue.
[0045] A third aspect of the embodiments of the present invention provides an electronic device, including:
[0046] A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the constant-time sorting method provided in the first aspect of the embodiments of the present invention.
[0047] A fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the constant-time sorting method provided in the first aspect of the embodiments of the present invention.
[0048] According to an embodiment of the present invention, the constant-time sorting device, method, system, electronic device and medium provided by the present invention include: a storage unit, a first-in first-out memory (FIFO), and a sorting unit. The storage unit, the FIFO, and the sorting unit are configured to store data to be processed. The data to be processed includes a first part of data and a second part of data. The number of elements in the first part of data is equal to the number of elements in the second part of data. The FIFO includes a first FIFO and a second FIFO. The first FIFO is used to read the first part of data, and the second FIFO is used to read the second part of data. The sorting unit is configured to perform internal sorting on the first part of data and the second part of data respectively in the case of the first iteration, and in the case of non-first iteration, use the merge sorting method to sort the first part of data and the second part of data to obtain multiple intermediate results, and use the multiple intermediate results as the data to be processed and input them into the storage unit until the final sorting result is obtained. Wherein, the number of elements in the intermediate result after each iteration is twice the number of elements in the intermediate result after the previous iteration, which can reduce the resource occupancy rate and reduce the calculation duration. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0050] Figure 1 Schematic structural diagram of a constant-time sorting device provided by an embodiment of the present invention;
[0051] Figure 2 Schematic diagram of a constant-time sorting device provided by an embodiment of the present invention;
[0052] Figure 3 Schematic flowchart of a constant-time sorting method provided by an embodiment of the present invention;
[0053] Figure 4 Schematic structural diagram of a data processing system provided by an embodiment of the present invention;
[0054] Figure 5 Schematic diagram showing the hardware structure of an electronic device. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] In order to make the objectives, features, and advantages of the present invention more obvious and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0056] The present invention provides a constant-time sorting device, method, system, electronic device, and medium. The constant-time sorting device includes: a storage unit, a first-in-first-out memory FIFO, and a sorting unit. The storage unit, the first-in-first-out memory FIFO, and the sorting unit are used to store the data to be processed. The data to be processed includes a first part of data and a second part of data. The number of elements in the first part of data is equal to the number of elements in the second part of data. The FIFO includes a first FIFO and a second FIFO. The first FIFO is used to read the first part of data, and the second FIFO is used to read the second part of data. The sorting unit is used to perform internal sorting on the first part of data and the second part of data respectively in the case of the first iteration, and use the merge sorting method to sort the first part of data and the second part of data to obtain a plurality of intermediate results in the case of non-first iteration, and use the plurality of intermediate results as the data to be processed and input them into the storage unit until the final sorting result is obtained. Among them, the number of elements in the intermediate result after each iteration is twice the number of elements in the intermediate result after the previous iteration, which can reduce the resource occupancy rate and reduce the calculation duration.
[0057] The following will describe in detail some embodiments of the present invention with reference to the accompanying drawings. Without conflict between the embodiments, the following embodiments and the features in the embodiments can be combined with each other.
[0058] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of a constant-time sorting device provided by an embodiment of the present invention. The constant-time sorting device mainly includes: a storage unit, a first-in-first-out memory FIFO, and a sorting unit.
[0059] The storage unit is used to store the data to be processed. The data to be processed includes a first part of data and a second part of data. The number of elements in the first part of data is equal to the number of elements in the second part of data. The data to be processed is arranged in ascending order.
[0060] The FIFO includes a first FIFO and a second FIFO. The first FIFO is used to read the first part of data, and the second FIFO is used to read the second part of data.
[0061] The sorting unit is used to perform internal sorting on the first part of data and the second part of data respectively in the case of the first iteration, and in the case of non-first iteration, use the merge sorting method to sort the first part of data and the second part of data to obtain multiple intermediate results, and use the multiple intermediate results as the data to be processed and input them into the storage unit until the final sorting result is obtained.
[0062] Wherein, the number of elements in the intermediate result after each sorting is twice the number of elements in the intermediate result after the previous sorting.
[0063] In an example, taking the data input to the constant-time sorting device as an array [7, 4, 6, 2, 9, 1, 5, 12, 15, 14, 3, 8, 10, 0, 13, 11] with a length of 16 as an example, then according to the parallel number 4, dividing it into groups of 4 elements per row as [7, 4, 2, 6], [9, 1, 5, 12], [15, 14, 3, 8], [10, 0, 13, 11], [7, 4, 2, 6] and [9, 1, 5, 12] can be divided into the first part of data, and [15, 14, 3, 8] and [10, 0, 13, 11] can be divided into the second part of data. In the case of the first iteration, perform internal sorting on the first part of data [7, 4, 2, 6] and [9, 1, 5, 12], and the second part of data [15, 14, 3, 8] and [10, 0, 13, 11] to obtain the intermediate results [2, 4, 6, 7], [3, 8, 14, 15], [1, 5, 9, 12], [0, 10, 11, 13].
[0064] In the present invention, the number of queues N of the FIFO can be an integer of 0, for example, 2, 3, 4, 6, etc. As Figure 2 shown, taking the parallel number as 4, that is, the number of queues in the first FIFO and the second FIFO is 4 as an example to illustrate the present invention schematically. It can be understood that in the case of the parallel number being 4, after the first iteration, the elements in the intermediate result are stored in the form of sub-lists with a length of 4 and internally sorted. After each round of iteration, the length of the elements in the intermediate result doubles, that is, 8, 16, etc.
[0065] In the sorting process of the present invention, the total number of elements in the first FIFO and the second FIFO remains unchanged. The first FIFO and the second FIFO read one element vector (four elements) in each clock cycle and output four elements.
[0066] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the constant-time sorting device provided by an embodiment of the present invention.
[0067] As Figure 2As shown, the storage unit includes: a first storage unit and a second storage unit.
[0068] The first storage unit includes a first storage subunit and a second storage subunit. The first storage subunit is used to store the first part of the data during odd-numbered iterations, and the second storage subunit is used to store the second part of the data during odd-numbered iterations.
[0069] The second storage unit includes a third storage subunit and a fourth storage subunit. The third storage subunit is used to store the first part of the data during even-numbered iterations, and the fourth storage subunit is used to store the second part of the data during even-numbered iterations.
[0070] In the present invention, the storage unit is split into two parts: a first storage unit (input storage part) and a second storage unit (output storage part). The first storage unit and the second storage unit can each be further split into two sub-blocks A and B according to the storage address (the first storage unit is split into a first storage subunit and a second storage subunit, and the second storage unit is split into a third storage subunit and a fourth storage subunit). ReadVec reads the elements to be compared from the input storage part, sends them to the first FIFO and / or the second FIFO, completes the sorting through the sorting unit on the right, and then writes the sorting result (intermediate result) into the output storage part in sequence. In the next round of comparison, the input storage part and the output storage part are swapped, that is, the second storage unit in the previous round of comparison becomes the first storage unit in this round of comparison. Moreover, the storage capacity of each row in the first storage unit and the second storage unit is equal to the parallelism of the constant-time sorting device, that is, the parallelism of the FIFO.
[0071] In the present invention, the result of the input storage part first enters the FIFO and enters the first FIFO and the second FIFO respectively according to the different source storage subunits. For example, in the first (odd-numbered) iteration, the first part of the data in the first storage subunit enters the first FIFO, and the second part of the data in the second storage subunit enters the second FIFO.
[0072] A threshold can be configured for both the first FIFO and the second FIFO. If the effective depth in the first FIFO and the second FIFO is lower than the threshold, the element vector of the corresponding storage subunit is read. In the cold start stage, both the first FIFO and the second FIFO may meet the reading validity, then the first FIFO or the second FIFO can be given priority. If the first FIFO is given priority, the elements of the first storage subunit are read into the first FIFO. After the effective depth of the first FIFO reaches the threshold, the elements of the second storage subunit are started to be read into the second FIFO.
[0073] In one example, taking the data input to the constant-time sorting device as an array of length 16, [7, 4, 6, 2, 9, 1, 5, 12, 15, 14, 3, 8, 10, 0, 13, 11], the first part of the data [7, 4, 2, 6] and [9, 1, 5, 12], and the second part of the data [15, 14, 3, 8] and [10, 0, 13, 11] as an example. In the first (odd-numbered) iteration, the first storage subunit stores the first part of the data [7, 4, 2, 6] and [9, 1, 5, 12], and the second storage subunit stores the second part of the data [15, 14, 3, 8] and [10, 0, 13, 11]. After the first iteration, the intermediate results [2, 4, 6, 7], [3, 8, 14, 15], [1, 5, 9, 12], [0, 10, 11, 13] are obtained. The intermediate results [2, 4, 6, 7] and [3, 8, 14, 15] are written into the third storage subunit as the first part of the data for the second iteration, and the intermediate results [1, 5, 9, 12] and [0, 10, 11, 13] are written into the fourth storage subunit as the second part of the data for the second iteration. After the second iteration, the intermediate results [1, 2, 4, 5, 6, 7, 9, 12] and [0, 3, 8, 10, 11, 13, 14, 15] are obtained. The intermediate result [1, 2, 4, 5, 6, 7, 9, 12] is written into the first storage unit as the first part of the data for the third iteration, and the intermediate result [0, 3, 8, 10, 11, 13, 14, 15] is written into the second storage unit as the second part of the data for the third iteration.
[0074] In an embodiment of the present invention, the constant-time sorting device further includes: at least one first holding register and at least one second holding register. The number of the first holding registers is the same as the number of queues in the first FIFO, and at least one of the first holding registers corresponds to the queues in the first FIFO one by one. The first holding register is used to prohibit the elements in the corresponding queue in the first FIFO from dequeuing when the first holding register is valid. The number of the second holding registers is the same as the number of queues in the second FIFO, and at least one of the second holding registers corresponds to the queues in the second FIFO one by one. The second holding register is used to prohibit the elements in the corresponding queue in the second FIFO from dequeuing when the second holding register is valid. That is, each queue in the first FIFO and the second FIFO has a holding register, and the holding register is used to prohibit the elements in the corresponding queue from dequeuing when it is valid.
[0075] In an embodiment of the present invention, the constant-time sorting device further includes: a bias unit, configured to set the first holding register of the corresponding queue to valid when each element of an intermediate result included in the first part of data has dequeued from the corresponding queue in the first FIFO during each sorting, set the second holding register of the corresponding queue to valid when each element of an intermediate result included in the second part of data has dequeued from the corresponding queue in the second FIFO during each sorting, and set both the first holding register and the second holding register to invalid when both the first holding register and the second holding register are valid.
[0076] In an example, during the second iteration, the third storage subunit stores the first part of data [2, 4, 6, 7] and [3, 8, 14, 15], and the fourth storage subunit stores the second part of data [1, 5, 9, 12] and [0, 10, 11, 13]. Then, during the first sorting process in the second iteration, when each element (2, 4, 6, 7) of an intermediate result [2, 4, 6, 7] included in the first part of data has dequeued from the corresponding queue in the first FIFO, the first holding register of the corresponding queue in the first FIFO for elements 2, 4, 6, 7 is set to valid.
[0077] In an embodiment of the present invention, the constant-time sorting device further includes: a first processing module and a second processing module. The first processing module is configured to subtract one from the effective depth of the first FIFO to obtain the current first effective depth when all queues of the first FIFO have elements dequeued at the same length position during each time, determine whether the current first effective depth is lower than a preset threshold, and if the current first effective depth is lower than the preset threshold, read the data to be processed in the storage unit, where the same address means that the address of the element in the queue is the same. The second processing module is configured to subtract one from the effective depth of the second FIFO to obtain the current second effective depth when all queues of the second FIFO have elements dequeued at the same length position during each time, determine whether the current second effective depth is lower than the preset threshold, and if the current second effective depth is lower than the preset threshold, read the data to be processed in the storage unit. Wherein, the effective depth indicates the length of the queue.
[0078] In one example, taking the threshold as 2 for illustration, during the second iteration, the third storage subunit stores the first part of the data [2, 4, 6, 7] and [3, 8, 14, 15], and the fourth storage subunit stores the second part of the data [1, 5, 9, 12] and [0, 10, 11, 13]. The lengths of the four queues of the first FIFO are 2. Specifically, 2 and 3 are sequentially stored in the first queue of the first FIFO, 4 and 8 are sequentially stored in the second queue, 6 and 14 are sequentially stored in the third queue, and 7 and 15 are sequentially stored in the fourth queue. Then, when 2, 4, 6, and 7 are all dequeued, that is, when all the elements at the same length position in all the queues of the first FIFO are dequeued, subtract 1 from the effective depth 2 of the first FIFO to obtain the current first effective depth 1. Since the current first effective depth 1 is lower than the threshold 2, the data to be processed in the storage unit is read.
[0079] In an embodiment of the present invention, the sorting unit includes: a comparison subunit and a sorting subunit. The comparison subunit is configured to pair and compare the head elements of all the queues in the first FIFO with the head elements of all the queues in the second FIFO according to a specified element pairing relationship to obtain a smaller element group by taking the smaller elements from the first FIFO and / or the second FIFO. The sorting subunit is configured to sort all the elements in the smaller element group in order of size.
[0080] In one example, during the second iteration, the third storage subunit stores the first part of the data [2, 4, 6, 7] and [3, 8, 14, 15], and the fourth storage subunit stores the second part of the data [1, 5, 9, 12] and [0, 10, 11, 13]. 2 and 3 are sequentially stored in the first queue of the first FIFO, 4 and 8 are sequentially stored in the second queue, 6 and 14 are sequentially stored in the third queue, and 7 and 15 are sequentially stored in the fourth queue. 1 and 0 are sequentially stored in the first queue of the second FIFO, 5 and 10 are sequentially stored in the second queue, 9 and 11 are sequentially stored in the third queue, and 12 and 13 are sequentially stored in the fourth queue. Then, the comparison subunit pairs and compares the head elements (2, 4, 6, 7) of all the queues in the first FIFO with the head elements (1, 5, 9, 12) of all the queues in the second FIFO according to the specified element pairing relationship to compare their sizes.
[0081] In an embodiment of the present invention, the specified element pairing relationship includes: the head element of the m-th queue of the first FIFO is paired with the head element of the n-th queue of the second FIFO. Wherein, m + n = N + 1, and both m and n are positive integers.
[0082] In one example, N is 4, and both the first FIFO and the second FIFO have 4 queues, as Figure 2As shown, the head element of the first queue of the first FIFO is paired with the head element of the fourth queue of the second FIFO, the head element of the second queue of the first FIFO is paired with the head element of the third queue of the second FIFO, the head element of the third queue of the first FIFO is paired with the head element of the second queue of the second FIFO, and the head element of the fourth queue of the first FIFO is paired with the head element of the first queue of the second FIFO. According to the pairing relationship in this embodiment, the above example comparison subunit pairs the head elements (2, 4, 6, 7) of all queues in the first FIFO with the head elements (1, 5, 9, 12) of all queues in the second FIFO according to the specified element pairing relationship ([2 - 12], [4 - 9], [6 - 5], [7 - 1]) to compare the sizes, takes the smaller elements from the first FIFO or this FIFO, and gets the smaller element group as 2, 4, 5, 1. Then the sorting subunit sorts all the elements in the smaller element group 2, 4, 5, 1 in order of size, and gets the intermediate result [1, 2, 4, 5].
[0083] Taking the data input to the constant-time sorting device as an array of length 16 [7, 4, 6, 2, 9, 1, 5, 12, 15, 14, 3, 8, 10, 0, 13, 11], the parallel number of the constant-time sorting device is 4, and the threshold of the effective depth of the first FIFO and the second FIFO is 2 as an example, the present invention will be described exemplarily.
[0084] According to the parallel number 4, divide it into groups of 4 elements per row [7, 4, 2, 6], [9, 1, 5, 12], [15, 14, 3, 8], [10, 0, 13, 11]. [7, 4, 2, 6] and [9, 1, 5, 12] can be divided into the first part of the data, and [15, 14, 3, 8] and [10, 0, 13, 11] can be divided into the second part of the data.
[0085] In the case of the first iteration, perform internal sorting on the first part of the data [7, 4, 2, 6] and [9, 1, 5, 12], and the second part of the data [15, 14, 3, 8] and [10, 0, 13, 11], and get the intermediate results [2, 4, 6, 7], [3, 8, 14, 15], [1, 5, 9, 12], [0, 10, 11, 13]. In the storage unit, the intermediate results [2, 4, 6, 7] and [3, 8, 14, 15] are written into the third storage subunit, and the intermediate results [1, 5, 9, 12] and [0, 10, 11, 13] are written into the fourth storage subunit.
[0086] In the first sorting of the second iteration, the first FIFO reads the first part of the data [2, 4, 6, 7] and [3, 8, 14, 15] from the third storage subunit, and the second FIFO reads the second part of the data [1, 5, 9, 12] and [0, 10, 11, 13] from the fourth storage subunit. Then, when the elements in the FIFO are output at this iteration, the head elements of the queues are compared according to the specified element pairing relationship (the head element of the m-th queue of the first FIFO is paired with the head element of the n-th queue of the second FIFO). The pairing relationships are [2 - 12, [4 - 9], [6 - 5], [7 - 1]. After passing through the comparison subunit, the minimum values 2, 4, 5, and 1 are output respectively. These four minimum values are output to the sorting subunit for sorting to obtain the array [1, 2, 4, 5]. At this time, two elements (1 and 5) in the second FIFO dequeue, and two elements (2 and 4) in the first FIFO dequeue. At the same time, the hold registers of the queues where 1, 2, 4, and 5 are located are set to valid. Then, the array [1, 2, 4, 5] is written into the first storage subunit. Since the length of the array in the example is too short, on the input side of the FIFO, the current array [1, 2, 4, 5] enters the first FIFO.
[0087] In the second sorting of the second iteration, due to the hold registers, although the head of some queues points to the elements in the next column, these elements cannot dequeue due to the limitation of the valid hold registers. So the remaining four elements 12, 9, 6, and 7 dequeue and pass through the sorting unit to output the intermediate result [6, 7, 9, 12]. At this time, two elements (9 and 12) in the second FIFO dequeue, and two elements (6 and 7) in the first FIFO dequeue. At the same time, the hold registers of the queues where 12, 9, 6, and 7 are located are set to valid. At this time, the hold registers of all 4 queues of the first FIFO and the second FIFO are valid, then all the first hold registers and all the second hold registers are set to invalid. The array [6, 7, 9, 12] is written into the first storage subunit. Similarly, the array [6, 7, 9, 12] enters the first FIFO.
[0088] In the third and fourth sortings of the first iteration, the remaining arrays participating in the comparison are [3, 8, 14, 15] and [0, 10, 11, 13]. Similar to the previous two sortings, [3, 8, 10, 0] and [13, 11, 14, 15] are output to the sorting unit respectively, and finally the arrays [0, 3, 8, 10] and [11, 13, 14, 15] are obtained. The arrays [0, 3, 8, 10] and [11, 13, 14, 15] are written into the second storage subunit and enter the second FIFO.
[0089] Therefore, the intermediate result stored in the first storage subunit obtained in this iteration is [1, 2, 4, 5, 6, 7, 9, 12], and the intermediate result stored in the second storage subunit is [0, 3, 8, 10, 11, 13, 14, 15].
[0090] In the first sorting of the third iteration, the first FIFO reads the first part of the data [1, 2, 4, 5, 6, 7, 9, 12] from the first storage subunit, and the second FIFO reads the second part of the data [0, 3, 8, 10, 11, 13, 14, 15] from the second storage subunit. However, since the parallelism is 4, the vector pairs processed by the sorting unit at this time are [1, 2, 4, 5] and [0, 3, 8, 10]. Among them, the output of the sorting subunit is [1, 2, 3, 0], and then the output of the comparison subunit is [0, 1, 2, 3]. The array [0, 1, 2, 3] is written into the third storage subunit. In the second sorting of the third iteration, the vector pairs processed by the sorting unit are [6, 7, 4, 5] and [11, 13, 8, 10]. Among them, the output of the sorting subunit is [6, 7, 4, 5], and then the output of the comparison subunit is [4, 5, 6, 7]. The array [4, 5, 6, 7] is written into the third storage subunit. In the third sorting of the third iteration, the vector pairs processed by the sorting unit are [X, x, 9, 12] and [11, 13, 8, 10]. Among them, the output of the sorting subunit is [10, 8, 9, 11], and then the output of the comparison subunit is [8, 9, 10, 11]. The array [8, 9, 10, 11] is written into the fourth storage subunit. In the fourth sorting of the third iteration, the vector pairs processed by the sorting unit are [x, x, x, 12] and [x, 13, 14, 15]. Among them, the output of the sorting subunit is [15, 14, 13, 12], and then the output of the comparison subunit is [12, 13, 14, 15]. The array [12, 13, 14, 15] is written into the fourth storage subunit. Here, x represents empty, and no element dequeues at this bit.
[0091] Therefore, the intermediate result obtained in this iteration is [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15].
[0092] Please refer to Figure 3 , Figure 3 which is the flowchart of the constant-time sorting method provided by an embodiment of the present invention.
[0093] As Figure 3 shown, the constant-time sorting method includes operations S310 to S330.
[0094] In operation S310, the data to be processed is written, and the data to be processed includes the first part of the data and the second part of the data.
[0095] Among them, the number of elements in the first part of the data is equal to the number of elements in the second part of the data, and the elements of the arrays in the first part of the data and the second part of the data are arranged in ascending order in sequence.
[0096] In operation S320, the first part of the data is read using the first FIFO, and the second part of the data is read using the second FIFO.
[0097] In operation S330, in the case of the first iteration, the first part of the data and the second part of the data are respectively internally sorted. In the case of a non-first iteration, the merge sort method is used to sort the first part of the data and the second part of the data to obtain multiple intermediate results, and the multiple intermediate results are used as the data to be processed, and operation S310 is executed until the final sorting result is obtained.
[0098] Among them, the number of elements in the intermediate result after each sorting is twice the number of elements in the intermediate result after the previous sorting.
[0099] In an embodiment of the present invention, writing the data to be processed in operation S310 includes: in the case of an odd-numbered iteration, storing the first part of the data in the first input storage subunit, and storing the second part of the data in the second input storage subunit; in the case of an even-numbered iteration, storing the first part of the data in the third input storage subunit, and storing the second part of the data in the fourth input storage subunit.
[0100] In an embodiment of the present invention, Figure 3 The method shown further includes: when all elements of an intermediate result included in the first part of the data have dequeued from a certain queue in the first FIFO during each sorting, setting the first holding register to be valid; when all elements of an intermediate result included in the second part of the data have dequeued from a certain queue in the second FIFO during each sorting, setting the second holding register to be valid; when both the first holding register and the second holding register are valid, setting both the first holding register and the second holding register to be invalid.
[0101] Among them, when the first holding register is valid, dequeueing of elements in the first FIFO is prohibited; when the second holding register is valid, dequeueing of elements in the second FIFO is prohibited; when the first holding register is invalid, dequeueing of elements in the first FIFO is allowed; when the second holding register is invalid, dequeueing of elements in the second FIFO is allowed.
[0102] In an embodiment of the present invention, Figure 3The method further includes: when all elements at the same length position in all queues of the first FIFO dequeue each time, subtracting one from the effective depth of the first FIFO to obtain the current first effective depth, and determining whether the current first effective depth is lower than a preset threshold. If the current first effective depth is lower than the preset threshold, read the data to be processed in the storage unit, where the same address means that the addresses of the elements in the queue are the same;
[0103] When all elements at the same length position in all queues of the second FIFO dequeue each time, subtracting one from the effective depth of the second FIFO to obtain the current second effective depth, and determining whether the current second effective depth is lower than the preset threshold. If the current second effective depth is lower than the preset threshold, read the data to be processed in the storage unit.
[0104] Wherein, the effective depth indicates the length of the queue.
[0105] In an embodiment of the present invention, in operation S330, using the merge sort method, sort the first part of data and the second part of data to obtain multiple intermediate results, including: pairing and comparing the head elements of all queues in the first FIFO and the head elements of all queues in the second FIFO according to the specified element pairing relationship, taking the smaller elements from the first FIFO or the second FIFO to obtain a smaller element group; sorting all elements in the smaller element group according to size to obtain an intermediate result.
[0106] In an embodiment of the present invention, the specified element pairing relationship includes: pairing the head element of the m-th queue of the first FIFO and the head element of the n-th queue of the second FIFO, where m + n = N + 1, and m and n are both positive integers.
[0107] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a data processing system provided by an embodiment of the present invention.
[0108] As Figure 4 shown, the data processing system includes: a constant-time sorting device 410, a data preprocessing device 420, and a GF(2) matrix Gaussian elimination device 430 as Figures 1 to 2 shown.
[0109] The data preprocessing device 420 is used to preprocess the final sorting result to obtain an input matrix, and send the input matrix to the GF(2) matrix Gaussian elimination device.
[0110] The GF(2) matrix Gaussian elimination device 430 includes: a partitioning unit, a data memory, a computing array, an operation memory, and a splicing unit.
[0111] The partitioning unit is used to partition the input matrix into at least one column block by columns, and divide the process of Gaussian elimination of the GF(2) matrix into multiple Bigsteps, where one Bigstep corresponds to the calculation process of one column block.
[0112] The data memory is used to store the column block.
[0113] The computing array includes at least one row of computing units. The row of computing units is used to read the data included in the column block from the data memory, and perform a computing operation on the data to obtain the final computing result of the column block. Wherein, in the case of replaying the calculation for the Bigstep corresponding to the column block, the operation information stored in the operation memory is read to implement the replay, and the calculation results of the historical Bigsteps are applied to the column block; in the case of performing the elimination calculation of the column block, the column block is finally eliminated into the form of an identity matrix, and the operation information corresponding to the elimination calculation is written into the operation memory.
[0114] The operation memory is used to store the operation information.
[0115] According to the present invention, in the block-based Gaussian elimination calculation scheduling, each time the data block involved is only a part of the matrix rather than the whole, so only this part of the data block needs to be stored on-chip, greatly reducing the scale of the on-chip memory. Moreover, the calculation scheduling mode of the present invention realizes the parallel execution of the calculation of GF(2) Gaussian elimination and the generation of the data matrix. The GF(2) Gaussian elimination of a certain column block, the output of the calculation results of the previous column block, and the generation of the matrix of the next column block (executed by the constant-time sorting device 410) are scheduled to be executed in parallel, so that the time for data generation and the output of calculation results can be hidden to the greatest extent. In the hardware design, the data memory can include two banks. When the data in one bank is being calculated iteratively, the data in the other bank can output the calculation results and start importing the input data of the next column block at the same time.
[0116] Any of a plurality of modules, sub-modules, units, and sub-units according to embodiments of the present invention, or at least part of the functions of any of them, may be implemented in one module. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present invention may be split into multiple modules for implementation. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present invention may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system in a package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means of integrating or packaging circuits, or in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present invention may be at least partially implemented as a computer program module, which can execute corresponding functions when the computer program module is run.
[0117] For example, the first processing module and the second processing module may be combined and implemented in one module / unit / sub-unit, or any one of them may be split into multiple modules / units / sub-units. Alternatively, at least part of the functions of one or more of these modules / units / sub-units may be combined with at least part of the functions of other modules / units / sub-units and implemented in one module / unit / sub-unit. According to embodiments of the present invention, at least one of the first processing module and the second processing module may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system in a package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means of integrating or packaging circuits, or in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the first processing module and the second processing module may be at least partially implemented as a computer program module, which can execute corresponding functions when the computer program module is run.
[0118] Figure 5 A block diagram of an electronic device suitable for implementing the method described above according to embodiments of the present invention is schematically shown. Figure 5 The shown electronic device is only an example and should not impose any limitation on the functions and usage scope of embodiments of the present invention.
[0119] As Figure 5As shown, the electronic device 500 according to an embodiment of the present invention includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage section 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), and so on. The processor 501 may also include on-board memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0120] In the RAM 503, various programs and data required for the operation of the system 500 are stored. The processor 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. The processor 501 performs various operations of the method flow according to an embodiment of the present invention by executing a program in the ROM 502 and / or the RAM 503. It should be noted that the program may also be stored in one or more memories other than the ROM 502 and the RAM 503. The processor 501 may also perform various operations of the method flow according to an embodiment of the present invention by executing a program stored in the one or more memories.
[0121] According to an embodiment of the present invention, the system 500 may further include an input / output (I / O) interface 505, and the input / output (I / O) interface 505 is also connected to the bus 504. The system 500 may further include one or more of the following components connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that a computer program read from it can be installed into the storage section 508 as needed.
[0122] According to an embodiment of the present invention, the method flow according to the embodiment of the present invention can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the processor 501, the above functions defined in the system of the embodiment of the present invention are executed. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0123] The present invention also provides a computer-readable storage medium, which may be included in the device / device / system described in the above embodiment; or may exist alone without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present invention is implemented.
[0124] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, device, or device.
[0125] For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 502 and / or RAM 503 and / or one or more memories other than ROM 502 and RAM 503.
[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutively represented blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0127] Those skilled in the art will appreciate that the features recited in the various embodiments and / or claims of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recited in the various embodiments and / or claims of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0128] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A constant-time sorting device, characterized in that, Comprising: A storage unit, a first-in-first-out memory FIFO, and a sorting unit; The storage unit is used to store data to be processed, the data to be processed includes a first part of data and a second part of data, and the number of elements in the first part of data is equal to the number of elements in the second part of data; The FIFO includes a first FIFO and a second FIFO, the first FIFO is used to read the first part of data, and the second FIFO is used to read the second part of data; The sorting unit is used to perform internal sorting on the first part of data and the second part of data respectively in the case of the first iteration, and in the case of non-first iteration, use the merge sorting method to sort the first part of data and the second part of data to obtain a plurality of intermediate results, and use the plurality of intermediate results as the data to be processed and input them into the storage unit until the final sorting result is obtained; Wherein, the number of elements in the intermediate result after each iteration is twice the number of elements in the intermediate result after the previous iteration.
2. The constant time sorting device according to claim 1, wherein The storage unit includes: A first storage unit, including a first storage subunit and a second storage subunit, the first storage subunit is used to store the first part of data for odd-numbered iterations, and the second storage subunit is used to store the second part of data for odd-numbered iterations; A second storage unit, including a third storage subunit and a fourth storage subunit, the third storage subunit is used to store the first part of data for even-numbered iterations, and the fourth storage subunit is used to store the second part of data for even-numbered iterations.
3. The constant time sorting device according to claim 1, wherein The constant-time sorting device further includes: At least one first holding register, the number of the first holding registers is the same as the number of queues in the first FIFO, at least one of the first holding registers corresponds to the queues in the first FIFO one by one, and the first holding register is used to prohibit the elements in the corresponding queue in the first FIFO from dequeuing when the first holding register is valid; At least one second holding register, the number of the second holding registers is the same as the number of queues in the second FIFO, at least one of the second holding registers corresponds to the queues in the second FIFO one by one, and the second holding register is used to prohibit the elements in the corresponding queue in the second FIFO from dequeuing when the second holding register is valid.
4. The constant-time sorting device according to claim 3, wherein Both the first FIFO and the second FIFO include N queues, N is an integer greater than 0, and the constant-time sorting device further includes: A bias unit is used to set the first holding register of the corresponding queue to valid when each element of an intermediate result included in the first part of data has dequeued from the corresponding queue in the first FIFO during each sorting operation, set the second holding register of the corresponding queue to valid when each element of an intermediate result included in the second part of data has dequeued from the corresponding queue in the second FIFO during each sorting operation, and set both the first holding register and the second holding register to invalid when both the first holding register and the second holding register are valid.
5. The constant-time sorting device according to any one of claims 1 to 4, characterized in that, The constant-time sorting device further includes: A first processing module is used to subtract one from the effective depth of the first FIFO to obtain the current first effective depth when all queues in the first FIFO have dequeued elements at the same length position during each operation, determine whether the current first effective depth is lower than a preset threshold, and if the current first effective depth is lower than the preset threshold, read the data to be processed in the storage unit; the same address means that the addresses of the elements in the queue are the same. A second processing module is used to subtract one from the effective depth of the second FIFO to obtain the current second effective depth when all queues in the second FIFO have dequeued elements at the same length position during each operation, determine whether the current second effective depth is lower than the preset threshold, and if the current second effective depth is lower than the preset threshold, read the data to be processed in the storage unit. Wherein, the effective depth indicates the length of the queue.
6. The constant-time sorting device according to claim 5, wherein The sorting unit includes: A comparison sub-unit is used to pair and compare the head elements of all queues in the first FIFO with the head elements of all queues in the second FIFO according to a specified element pairing relationship, take the smaller elements from the first FIFO or the second FIFO to obtain a smaller element group. A sorting sub-unit is used to sort all elements in the smaller element group in order of size to obtain an intermediate result.
7. The constant-time sorting apparatus according to claim 6, wherein The specified element pairing relationship includes: Pairing the head element of the m-th queue in the first FIFO with the head element of the n-th queue in the second FIFO. Wherein, m + n = N + 1, and both m and n are positive integers.
8. A constant-time sorting method, characterized in that, Includes: Write the data to be processed, where the data to be processed includes a first part of data and a second part of data, and the number of elements in the first part of data is equal to the number of elements in the second part of data. Read the first part of data using the first FIFO, and read the second part of data using the second FIFO. In the case of the first iteration, perform internal sorting on the first part of data and the second part of data respectively. In the case of non-first iteration, use the merge sorting method to sort the first part of data and the second part of data to obtain multiple intermediate results, and use the multiple intermediate results as the data to be processed, and execute the operation of writing the data to be processed until the final sorting result is obtained. Among them, the number of elements in the intermediate result after each iteration is twice the number of elements in the intermediate result after the previous iteration.
9. The constant time sorting method according to claim 8, wherein, The writing of the data to be processed includes: In the case of odd-numbered iterations, storing the first part of the data in the first input storage subunit, and storing the second part of the data in the second input storage subunit; In the case of even-numbered iterations, storing the first part of the data in the third input storage subunit, and storing the second part of the data in the fourth input storage subunit.
10. The constant-time sorting method according to claim 8, wherein The method further includes: When all elements of an intermediate result included in the first part of the data have dequeued from a certain queue in the first FIFO during each sorting, setting the first holding register to valid; When all elements of an intermediate result included in the second part of the data have dequeued from a certain queue in the second FIFO during each sorting, setting the second holding register to valid; When both the first holding register and the second holding register are valid, setting both the first holding register and the second holding register to invalid; Among them, when the first holding register is valid, dequeuing of elements in the first FIFO is prohibited; when the second holding register is valid, dequeuing of elements in the second FIFO is prohibited; when the first holding register is invalid, dequeuing of elements in the first FIFO is allowed; when the second holding register is invalid, dequeuing of elements in the second FIFO is allowed.
11. The constant time sorting method according to claim 8, wherein The method further includes: When all queues of the first FIFO have dequeued elements at the same length position, subtracting one from the effective depth of the first FIFO to obtain the current first effective depth, and determining whether the current first effective depth is lower than a preset threshold. If the current first effective depth is lower than the preset threshold, reading the data to be processed in the storage unit, where the same address means the address of the element in the queue is the same; When all queues of the second FIFO have dequeued elements at the same length position, subtracting one from the effective depth of the second FIFO to obtain the current second effective depth, and determining whether the current second effective depth is lower than the preset threshold. If the current second effective depth is lower than the preset threshold, reading the data to be processed in the storage unit; Among them, the effective depth indicates the length of the queue.
12. A data processing system, characterized in that, It includes: The constant-time sorting device, data preprocessing device, and GF(2) matrix Gaussian elimination device according to any one of claims 1 to 7; The data preprocessing device is used to preprocess the final sorting result to obtain an input matrix, and send the input matrix to the GF(2) matrix Gaussian elimination device; The GF(2) matrix Gaussian elimination device includes: a partitioning unit, a data memory, a computing array, an operation memory, and a splicing unit; The partitioning unit is configured to partition an input matrix into at least one column block by columns, and divide the process of Gaussian elimination of the GF(2) matrix into multiple Bigsteps, where one Bigstep corresponds to the calculation process of one column block; The data memory is configured to store the column blocks; The computing array, the computing array includes at least one row of computing units, the row of computing units is configured to read the data included in the column block from the data memory, and perform a computing operation on the data to obtain the final computing result of the column block. Wherein, in the case of performing replay calculation for the Bigstep corresponding to the column block, read the operation information stored in the operation memory to implement replay, and apply the calculation results of historical Bigsteps to the column block; in the case of performing elimination calculation on the column block, finally eliminate the column block into the form of an identity matrix, and write the operation information corresponding to the elimination calculation into the operation memory; The operation memory is configured to store the operation information.
13. An electronic device, characterized in that, Comprising: A processor; And A constant-time sorting device according to any one of claims 1 to 7, the constant-time sorting device being electrically connected to the processor.
14. A computer-readable storage medium having executable instructions stored thereon, which when executed by a processor cause the processor to execute the method according to any one of claims 8 to 11.
Citation Information
Patent Citations
Method and device for achieving data convolution operation based on FPGA, and medium
CN112464150A
Systems and methods for post-quantum cryptography optimization
US11322050B1