Data sorting method and hardware-accelerated parallel sorting circuit
Through parallel processing and hardware acceleration circuits, the problem of long data sorting time in the prior art is solved, and fast and efficient data sorting is achieved.
Patent Information
- Application Number
- CN202210245825.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-14
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-03-14
AI Technical Summary
The existing data sorting algorithm needs to read all data into the complete storage space before it can start sorting, resulting in a long processing time and it is difficult to meet the low-latency real-time requirements such as modern AI computing and 5G network scheduling.
By constructing storage sequences and intermediate sequences, using parallel processing methods to sort while reading data, combining hardware acceleration circuits to realize parallel sorting of data, including comparing logic and filling method of storage sequences, and using latch circuits and offset modules to update data.
The sorting time overhead is significantly reduced, and the sorting is completed as soon as the data is read, which improves the overall computing speed.
Smart Images

Figure CN114706554B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to integrated circuit technology. Background Art
[0002] Data sorting is to arrange a group of disordered data in descending or ascending order. Data sorting is widely used in fields such as computer vision computing, network scheduling, and AI computing. Sorting is a computationally complex and time-consuming operation. Therefore, there are numerous sorting algorithms. Common sorting algorithms include bubble sort, simple selection sort, direct insertion sort, heap sort, merge sort, quick sort, and so on. These sorting algorithms usually need to read all the sorted data into the complete storage space before starting the sorting process, and the processing time is relatively long. The present invention utilizes the concurrent characteristics implemented by integrated circuit hardware to invent a real-time sorting integrated circuit implementation scheme, which can achieve the real-time effect of sorting while reading data. When the data reading is completed, the sorting is completed, and the sorting calculation is quickly completed, which has very important application value for scenarios with high requirements for low latency and real-time performance such as modern AI computing, big data analysis, and 5G network scheduling.
[0003] The implementation timing of traditional sorting algorithms is as Figure 1 shown. It is usually divided into two stages: data reading (time-consuming T1) and data sorting (time-consuming T2). In the first step, the data needs to be stored in a complete array or queue. In the second step, various sorting calculations and processing are performed on the data in the array. The final sorting time is T1 + T2, where T1 is the data reading time and T2 is the algorithm processing time. According to different algorithms, the processing time is different, and usually T2 > T1. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a data sorting method and a hardware-accelerated parallel sorting circuit, which can greatly improve the sorting efficiency by executing data reading and data sorting in parallel.
[0005] The technical solution adopted by the present invention to solve the above technical problem is to provide a data sorting method, including the following steps:
[0006] 1) Construct a storage sequence with the same length as the sequence to be sorted, and determine the comparison logic and the storage sequence filling method according to the predetermined sorting method:
[0007] If the sorting method is decreasing along the address order, fill the storage sequence with the maximum value and determine the comparison logic as ">" or the comparison logic as "≥";
[0008] If the sorting method is increasing along the address order, fill the storage sequence with the minimum value and determine the comparison logic as "<" or the comparison logic as "≤";
[0009] 2) Construct an intermediate sequence with the same length as the sequence to be sorted, where each digit of the intermediate sequence is the first identifier; construct an offset sequence with the same length as the intermediate sequence, where the last digit of the offset sequence is the second identifier and all other addresses are the first identifier;
[0010] 3) Read an unread address in the sequence to be sorted and store the value therein at the end of the storage sequence;
[0011] 4) Read an unread address in the sequence to be sorted, use its value as the newly entered value, and make a traversal comparison with the values of each digit in the storage sequence.
[0012] If the relationship between the newly entered value and the value of the current digit in the storage sequence conforms to the comparison logic corresponding to the predetermined sorting method, then set the current digit of the intermediate sequence to the first identifier; otherwise, set the current digit of the intermediate sequence to the second identifier;
[0013] After the traversal is completed, write each digit of the intermediate sequence except the first digit to the offset sequence in a way that shifts one position towards the lower address direction;
[0014] [[ID=!16]]5) Update the storage sequence:
[0015] If the value of the i-th digit in the intermediate sequence is the second identifier and the i-th digit in the offset sequence is the second identifier, then keep the value of the i-th digit in the storage sequence unchanged;
[0016] If the i-th digit in the intermediate sequence is the first identifier and the i-th digit in the offset sequence is the second identifier, then update the value of the i-th digit in the storage sequence to the newly entered value;
[0017] If the i-th digit in the intermediate sequence is the first identifier and the i-th digit in the offset sequence is the first identifier, then update the value of the i-th digit in the storage sequence to the value of the (i + 1)-th digit in the storage sequence;
[0018] Wherein i is the address serial number;
[0019] 6) If all addresses in the sequence to be sorted have been read, the sorting is completed; otherwise, return to step 4).
[0020] The present invention also provides a hardware-accelerated parallel sorting circuit, which is characterized by including:
[0021] N comparison branches, each comparison branch includes a selector, a latch circuit, a comparison module, and an offset module connected in series. The output end of the comparison module is connected to the first input end of the bit-splicing module in this branch, and the output end of the offset module is connected to the second input end of the bit-splicing module in this branch, where N is the number of numbers to be sorted;
[0022] It should be noted that in the above translation, for the sake of better understanding, the Chinese punctuation marks in the original text are replaced with English punctuation marks in the translation. If you have any other special requirements, please feel free to let me know.In each comparison branch, the output terminal of the selector is connected to the input terminal of the latch circuit, and the output terminal of the latch circuit is connected to the second input terminal of the comparator and the first input terminal of the selector;
[0023] In each comparison branch except the first comparison branch, the output terminal of the latch circuit is also connected to the second input terminal of the selector in the adjacent previous comparison branch;
[0024] The offset module is used to connect the output terminal of the comparator in the adjacent subsequent branch to the second input terminal of the bit-splicing module in this branch, and the output of the offset module in the last branch is constant;
[0025] The data input terminal is connected to the third input terminal of each selector and the first input terminal of the comparison module;
[0026] The output terminal of the latch circuit in the first comparison branch is connected to the output terminal of the sorting circuit.
[0027] The latch circuit is composed of a combination of D flip-flops.
[0028] The present invention sorts through a parallel processing method, which can significantly reduce the time overhead and improve the overall operation speed. Description of the Drawings
[0029] Figure 1 is a timing schematic diagram of the prior art.
[0030] Figure 2 is a timing schematic diagram of the present invention.
[0031] Figure 3 is a schematic diagram of the sorting algorithm flow.
[0032] Figure 4 is a schematic diagram of the initialization of the embodiment.
[0033] Figure 5 is a schematic diagram of the first data writing of the embodiment.
[0034] Figure 6 is a schematic diagram of the overall function division of the accelerated sorting implementation of the present invention.
[0035] Figure 7 is a schematic diagram of the implementation scheme of the accelerated sorting algorithm of the present invention.
[0036] Figure 8 Schematic diagram of the interface block diagram of the accelerated sorting module.
[0037] Figure 9 is the schematic diagram of the input data latch circuit of the accelerated sorting module.
[0038] Figure 10 is the schematic diagram of the circuit for implementing the core sorting function of the accelerated sorting module.
[0039] Figure 11 It is a schematic diagram comparing the parallel hardware acceleration of the present invention with the traditional bubble processing time.
[0040] Figure 12 It is the first simulation schematic diagram of the parallel hardware acceleration of the present invention.
[0041] Figure 13 It is the second simulation schematic diagram of the parallel hardware acceleration of the present invention. Detailed implementation manner
[0042] The present invention provides a data sorting method, including the following steps:
[0043] 1) Construct a storage sequence with the same number of digits as the sequence to be sorted, and determine the comparison logic and the storage sequence filling method according to a predetermined sorting method:
[0044] If the sorting method is decreasing along the address order, fill the storage sequence with the maximum value, and determine the comparison logic as ">" or the comparison logic as "≥";
[0045] If the sorting method is increasing along the address order, fill the storage sequence with the minimum value, and determine the comparison logic as "<" or the comparison logic as "≤";
[0046] Detailed explanation of step 1):
[0047] For example, the sequence to be sorted is a sequence composed of 4 numbers (i.e., the length is 4) [5, 9, 3, 6], and the predetermined sorting is in descending order, and the corresponding comparison logic is ">". In step 1), construct a storage sequence with a length of 4. In hexadecimal, the maximum value of a number is F and the minimum value is 0. For the descending sorting method, fill the storage sequence with the maximum value F, and the filled storage sequence is [F, F, F, F]; for the ascending sorting method, fill the storage sequence with the minimum value 0, and the filled storage sequence is [0, 0, 0, 0]. The sequence [5, 9, 3, 6] in the present invention is only an example.
[0048] 2) Construct an intermediate sequence with the same length as the sequence to be sorted, and each digit of the intermediate sequence is the first identifier; construct an offset sequence with the same length as the intermediate sequence, the last digit of the offset sequence is the second identifier, and other addresses are the first identifier;
[0049] Detailed explanation of step 2):
[0050] Similarly, taking the sequence [5, 9, 3, 6] as an example, construct an intermediate sequence and an offset sequence with a length of 4. Fill the intermediate sequence with the first identifier, and fill the offset sequence with the first identifier and the second identifier. The first identifier and the second identifier are different. For example, the first identifier is 0 and the second identifier is 1; or the first identifier is X and the second identifier is Y. In computer technology, both X and Y need to be transcoded into numbers, so they can still be placed in the sequence. In essence, it is the same as 0 and 1. As an example, here 0 is used as the first identifier and 1 is used as the second identifier. The intermediate sequence is [0, 0, 0, 0], and the offset sequence is [0, 0, 0, 1].
[0051] 3) Read an unread address in the sequence to be sorted, and store the value therein at the end of the storage sequence;
[0052] Detailed explanation of step 3):
[0053] Read the number corresponding to an unread address from the sequence to be sorted [5, 9, 3, 6]. The present invention does not limit the reading order to be increasing or decreasing by address, and only requires that the read address has not been read before. For a certain address, for example, the number 5 at address number 0, once read in this step, this address will become a read address and will be avoided in the next read operation. At this time, the storage sequence is updated to [F, F, F, 5].
[0054] 4) Read an unread address in the sequence to be sorted, use its value as the newly entered value, and make a traversal comparison with the values of each bit in the storage sequence.
[0055] If the relationship between the newly entered value and the value of the current bit in the storage sequence conforms to the comparison logic corresponding to the predetermined sorting method, set the current bit of the intermediate sequence to the second identifier 1, otherwise set the current bit of the intermediate sequence to the first identifier 0;
[0056] After the traversal is completed, write the bits other than the first bit of the intermediate sequence to the offset sequence in a way that shifts one bit in the direction of the lower address.
[0057] Detailed explanation of step 4):
[0058] After reading the number 9 at address 1 in step 3), in this step, read one of the numbers corresponding to the remaining 3 addresses, for example, the number 9 at address 1, as the newly entered value, and compare it with each bit in the storage sequence [F, F, F, 5]. Since the comparison logic corresponding to sorting from large to small is ">", the newly entered value 9 is compared with each bit of the storage sequence [F, F, F, 5] one by one:
[0059]
[0060] The middle sequence is updated to [0, 0, 0, 1]. After the update of the middle sequence, except for the first address 0, each digit is shifted one place towards the lower address direction, and the offset sequence is written according to the shifted addresses. The written offset sequence is [0, 0, 1, 1].
[0061] 5) Update the storage sequence:
[0062] If the value of the i-th digit in the middle sequence is the second identifier and the i-th digit in the offset sequence is the second identifier, then keep the value of the i-th digit in the storage sequence unchanged;
[0063] If the i-th digit in the middle sequence is the first identifier and the i-th digit in the offset sequence is the second identifier, then update the value of the i-th digit in the storage sequence to the newly entered value;
[0064] If the i-th digit in the middle sequence is the first identifier and the i-th digit in the offset sequence is the first identifier, then update the value of the i-th digit in the storage sequence to the value of the (i + 1)-th digit in the storage sequence;
[0065] Wherein, i is the address serial number;
[0066] 6) If all addresses in the sequence to be sorted have been read, the sorting is completed; otherwise, return to step 4).
[0067] In this step, i is the address serial number, the starting bit is the 0th bit, and each value of i is judged one by one.
[0068] Embodiment
[0069] The timing sequence of the sorting algorithm of the present invention is as Figure 2 shown. Data reading and data sorting are executed in parallel. When the data reading is completed, the sorting is completed. The data reading time T1 = the sorting processing time T2, and the overall sorting time is T1.
[0070] The idea of the present invention is to store the input data through an array unit of data storage, and the incoming data is stored in ascending or descending order according to the ascending or descending requirements. The specific algorithm flow is as Figure 6 shown.
[0071] All storage units need to be initialized before the data enters. The initialized value is set according to the sorting requirements. If sorting from large to small is required, the initialized value is the maximum value of the data; on the contrary, if sorting from small to large is required, the initialized value is the minimum value of the data.
[0072] In this embodiment, the second column of the storage array is the storage sequence, and the first column is the address column.
[0073] The first data enters the space with the largest fixed storage address. Each time a new data data_new enters, the corresponding storage address is addr_new. Based on the comparison result between the new data data_new and all the data data_array[N] in the storage array, including all the previously entered data, the largest starting address addr_gt greater than or equal to this data is found. If addr_new = addr_gt, then data_new is written into the storage space at addr_gt. At the same time, all the data in the space greater than or equal to addr_gt is shifted down to the next lower address space. For example, the data originally at addr_gt is stored at addr_gt - 1, the data at addr_gt - 1 is stored at addr_gt - 2, and so on.
[0074] If all the data is stored in the above - mentioned manner, after all the data is read, the data is stored in the corresponding order, and the sorting function is automatically completed. The data can be read out in order.
[0075] Taking the four numbers 5, 9, 3, and 6 as an example, where the data is an unsigned number with a 4 - bit width, and sorting from largest to smallest, the implementation process of the present invention is introduced below.
[0076] Since the data is an unsigned data with a 4 - bit width, the 4x4 storage column is initialized to the maximum value F of the data range. The corresponding schematic diagram is as Figure 4 shown.
[0077] The first data written is 5, which is fixed to be written in the highest - address unit, that is, address 3, as Figure 5 shown.
[0078] The second data written is 9. After comparing 9 with all the data in the units, a result is obtained and represented in bitmap form. If the corresponding bit is 1, it means the data in the corresponding unit is greater than or equal to 9; otherwise, the corresponding bit is 0, which means the data in the corresponding unit is less than 9. The bitmap is shifted one bit to the right to get bitmap_rsf (the "right shift" referred to in the present invention means shifting horizontally towards the lower - address direction) and the highest bit is filled with 1. Combining the shifted - right bitmap_rsf and the original result bitmap for joint judgment (i represents the address unit, i = {0, 1, 2, 3}):
[0079] If {bitmap[i], bitmap_rsf[i]} is equal to 00, then:
[0080] The data in the corresponding i unit is updated to the data in the i + 1 unit;
[0081] If {bitmap[i], bitmap_rsf[i]} is equal to 11, then:
[0082] The data corresponding to the i-th unit remains unchanged.
[0083] If {bitmap[i], bitmap_rsf[i]} is equal to 01, then:
[0084] The data corresponding to the i-th unit is updated to the newly incoming data.
[0085] According to the algorithm principle, the second incoming number 9 is automatically written into the second data unit. At the same time, the data in the first unit and the zero-th unit will be updated to the original data in the second unit and the first unit, and the data in the third unit remains unchanged. Table 1 shows the process of writing the second data 9 in the example of the hardware-accelerated sorting algorithm.
[0086] Table 1
[0087]
[0088] Similarly, the update process of the input data of the third data 3 is shown in Table 2.
[0089] Table 2
[0090]
[0091] Similarly, the update process of the input data of the fourth data 6 is shown in Table 3. When the four data are read in, they are arranged in descending order from low address to high address in the entire storage column.
[0092] Table 3
[0093]
[0094] The functional division of the sorting algorithm of the present invention is shown as Figure 6 shown. It is mainly divided into four functional blocks: data selection, data storage, data comparison, and result recording to complete.
[0095] In the implementation process of the integrated circuit accelerator, data selection is completed by a selector, data storage is completed by D flip-flops, the comparison unit is implemented by a comparator, and result recording is completed by bitmap and its right-shifted signal. The overall hardware scheme of the sorting algorithm is shown as Figure 7 shown. Figure 7 In, the selection of the dout output is realized through a modulo counter to obtain the sorted result.
[0096] The specific circuit implementation interface block diagram is shown as Figure 8 shown.
[0097] After the data enters, the data is latched to ensure the timing of subsequent data processing. The specific circuit schematic diagram of the latch is shown as Figure 9 shown.
[0098] The circuit implementation schematic diagram of input latch data queuing, queue sorting, comparing the input latch data with the queue data, result recording and post-processing is as follows Figure 10 shown
[0099] The Verilog RTL description corresponding to the sorting module is as follows
[0100]
[0101]
[0102]
[0103] I. Technical effects
[0104] Compared with the traditional bubble method, the sorting is carried out by means of parallel processing hardware acceleration. The longer the sorted data is, the more time is saved by parallel processing. As Figure 11 shown, for data with a length of 60, the parallel hardware acceleration method only needs 60 clock cycles to complete, while the bubble method needs 3660 clock cycles to complete
[0105] II. Embodiment
[0106] Write a Testbench to simulate the sorting circuit. The testbench code is as follows
[0107]
[0108]
[0109]
[0110]
[0111] As Figure 12 shown, through simulation, it can be seen that 10 din data (0, 1, 2, 3, 4, 5, 6, 7, 8, 9) are serially input from small to large. When the 10th data is input, dout_arr will arrange the 10 data in descending order (9, 8, 7, 6, 5, 4, 3, 2, 1, 0). The sorting of 10 data is completed in 10 cycles. As Figure 13 shown, through simulation, it can be seen that 10 din random data (269, 769, 530, 101, 397, 781, 611, 521, 641, 292) are serially input from small to large. When the 10th data is input, dout_arr will arrange the 10 data in descending order (781, 769, 641, 611, 530, 521, 397, 292, 269, 101).
[0112] The description and drawings of the present invention have fully described the necessary technical content of the present invention, and those of ordinary skill in the art can implement it accordingly. For more specific details (such as the structures of units such as selectors and latch circuits), they will not be elaborated further.
Claims
1. A hardware-accelerated parallel sorting circuit, characterized in that Including: N comparison branches, each comparison branch includes a selector, a latch circuit, a comparison module, and an offset module connected in series. The output end of the comparison module is connected to the first input end of the bit-splicing module in this branch, and the output end of the offset module is connected to the second input end of the bit-splicing module in this branch. The N is the number of numbers to be sorted; In each comparison branch, the output end of the selector is connected to the input end of the latch circuit, and the output end of the latch circuit is connected to the second input end of the comparator and the first input end of the selector; In each comparison branch except the first comparison branch, the output end of the latch circuit is also connected to the second input end of the selector in the adjacent previous comparison branch; The offset module is used to connect the output end of the comparator in the adjacent next branch to the second input end of the bit-splicing module in this branch, and the output of the offset module of the last branch is constant; The data input end is connected to the third input end of each selector and the first input end of the comparison module; The output end of the latch circuit of the first comparison branch is connected to the output end of the sorting circuit; The hardware-accelerated parallel sorting circuit is used to implement a data sorting method including the following steps: 1) Construct a storage sequence with the same length as the sequence to be sorted, and determine the comparison logic and the storage sequence filling method according to the predetermined sorting method: If the sorting method is decreasing along the address order, fill the storage sequence with the maximum value, and determine the comparison logic as ">", or the comparison logic as "≥"; If the sorting method is increasing along the address order, fill the storage sequence with the minimum value, and determine the comparison logic as "<", or the comparison logic as "≤"; 2) Construct an intermediate sequence with the same length as the sequence to be sorted, and each bit of the intermediate sequence is the first identifier; construct an offset sequence with the same length as the intermediate sequence, the last bit of the offset sequence is the second identifier, and other positions are all the first identifier; 3) Read an unread address in the sequence to be sorted, and store the value therein to the end of the storage sequence; 4) Read an unread address in the sequence to be sorted, use its value as the newly entered value, and make a traversal comparison with the values of each bit in the storage sequence, If the relationship between the newly entered value and the value of the current bit in the storage sequence conforms to the comparison logic corresponding to the predetermined sorting method, set the current bit of the intermediate sequence to the first identifier, otherwise set the current bit of the intermediate sequence to the second identifier; After the traversal is completed, write each bit of the intermediate sequence except the first bit to the offset sequence in a way of shifting one bit in the direction of the lower address; 5) Update the storage sequence: If the value of the i-th bit of the intermediate sequence is the second identifier and the i-th bit of the offset sequence is the second identifier, keep the value of the i-th bit of the storage sequence unchanged; If the i-th bit of the intermediate sequence is the first identifier and the i-th bit of the offset sequence is the second identifier, update the value of the i-th bit of the storage sequence to the newly entered value; If the i-th bit of the intermediate sequence is the first identifier and the i-th bit of the offset sequence is the first identifier, update the value of the i-th bit of the storage sequence to the value of the (i + 1)-th bit of the storage sequence; The i is the address serial number; 6) If all addresses in the sequence to be sorted have been read, the sorting is completed, otherwise return to step 4).
2. The hardware-accelerated parallel sorting circuit according to claim 1, wherein The latch circuit is composed of a combination of D flip-flops.
Citation Information
Patent Citations
Data sorting device, method and data processing chip achieved by hardware
CN105512179A
Super-parallel comparison method and system
CN110647665A