Data sorting method and hardware accelerator
By alternately using multiple sorting operations in registers, the problem of inefficiency of existing data sorting algorithms in parallel calculations is solved, and efficient data sorting is achieved.
Patent Information
- Application Number
- CN202510859900.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-25
AI Technical Summary
The existing data sorting algorithms have problems such as corruption of calculation rules and large memory access in parallel computing, resulting in inefficient parallel computing.
The target data is stored in the register using a register based arrangement order, and the first sorting operation and intermediate sorting operation (such as the second sorting operation and the third sorting operation) are alternately performed until the target data is sorted from large to small or from small to large by numerical value.
It improves the efficiency of data sorting, reduces the time complexity in the sorting process, and simplifies the sorting process.
Smart Images

Figure CN120371257A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and particularly to a data sorting method and a hardware accelerator. Background Art
[0002] Currently, in the fields of high-performance computing, big data analysis, artificial intelligence, etc., the demand for data sorting operations is increasing day by day. To save the time of data operations and processing, it is necessary to improve the performance of data sorting operations. Currently, the following data sorting methods are widely used in the industry. For example, the sorting algorithms applied to parallel computing mainly include bubble sort, direct sort, insertion sort, quick sort, and heap sort, etc. For different types of sorting algorithms in the scenario of parallel computing, the following drawbacks generally exist: during the parallel computing process, the calculations of multiple computing units disrupt the regularity of the calculations, reducing the efficiency of parallel computing. In addition, during the calculation process, the amount of data accessed from the associated memory (such as DRAM) is large, resulting in a decline in the performance of data sorting. Summary of the Invention
[0003] An embodiment of this application provides a data sorting method, including: Based on the arrangement order of registers, store multiple target data into their respective corresponding registers; Perform a first sorting operation on the target data in the registers, which includes: starting from the first register with the first serial number, sequentially compare the target data in two adjacent registers, store the target data with a larger value into the register with a higher or lower sorting order among the two registers participating in the comparison, and store the target data with a smaller value into the register with a lower or higher sorting order among the two registers participating in the comparison; Perform a sub-operation in the intermediate sorting operation on the target data in the registers, where the sub-operation includes a second sorting operation and a third sorting operation. The second sorting operation includes: starting from the second register with the second serial number, sequentially compare the target data in two adjacent registers, and compare the first register with the register with the last serial number, store the target data with a larger value into the register with a higher or lower sorting order among the two registers participating in the comparison, and store the target data with a smaller value into the register with a lower or higher sorting order among the two registers participating in the comparison; The third sorting operation includes: starting from the second register, sequentially compare the target data in two adjacent registers, store the target data with a larger value into the register with a higher or lower sorting order among the two registers participating in the comparison, and store the target data with a smaller value into the register with a lower or higher sorting order among the two registers participating in the comparison; Repeat the first sorting operation and the second sorting operation alternately for multiple times, or repeat the first sorting operation and the third sorting operation alternately for multiple times, until all the target data are sorted in descending or ascending order of values.
[0004] Optionally, when there are multiple data units in the target data and they are not sorted, the method further includes: Based on the arrangement order of the registers, store the multiple data units into their respective corresponding registers; Take the data units in the registers as objects and perform the first sorting operation on them; Take the data units in the registers as objects and perform one of the sub-operations in the intermediate sorting operation on them; Repeat the first sorting operation and the second sorting operation alternately for multiple times, or repeat the first sorting operation and the third sorting operation alternately for multiple times, until all the data units are sorted in descending or ascending order of values.
[0005] Optionally, when there are multiple data units in the target data and they are already sorted, the method further includes: Group the data units based on the numerical magnitudes of the data units to form multiple data unit groups; Group the registers based on the arrangement order of the registers to form multiple register groups; Based on the arrangement order of the register groups, store the data unit groups into their respective corresponding register groups; Take the data unit groups in the register groups as objects and perform the first sorting operation on them; Take the data unit groups in the register groups as objects and perform one of the sub-operations in the intermediate sorting operation on them, Repeat the first sorting operation and the second sorting operation alternately for multiple times, or repeat the first sorting operation and the third sorting operation alternately for multiple times, until all the data unit groups are sorted in descending or ascending order of values.
[0006] Optionally, the grouping the data units based on the numerical magnitudes of the data units to form multiple data unit groups includes: dividing the data units into 2 data unit groups based on the numerical magnitudes of the data unit groups; Correspondingly, the grouping the registers to form multiple register groups includes: constructing 1 register group based on 4 registers.
[0007] Optionally, the method further includes: Sorting the target data based on a preset time interval; Sending the sorted target data to a memory through a memory access controller.
[0008] An embodiment of the present application further provides a hardware accelerator connected to a memory. The hardware accelerator includes: Multiple sequentially arranged registers; A scheduling unit connected to the registers, configured to store target data into their respective corresponding registers; Multiple basic sorting units, each connected to two adjacent corresponding registers. The basic sorting unit is configured to: perform a first sorting operation on the target data in the registers under the control of the scheduling unit; perform one of the sub-operations of an intermediate sorting operation on the target data in the registers, where the sub-operations include a second sorting operation and a third sorting operation; repeatedly alternate between the first sorting operation and the second sorting operation multiple times, or repeatedly alternate between the first sorting operation and the third sorting operation multiple times until all the target data is sorted in descending or ascending order of values. Wherein, The first sorting operation includes: starting from a first register with a first serial number, sequentially comparing the target data in two adjacent registers, storing the target data with a larger value in the register with a higher or lower rank among the two registers participating in the comparison, and storing the target data with a smaller value in the register with a lower or higher rank among the two registers participating in the comparison. Wherein, The second sorting operation includes: starting from a second register with a second serial number, sequentially comparing the target data in two adjacent registers, and comparing the first register with the register with the last serial number, storing the target data with a larger value in the register with a higher or lower rank among the two registers participating in the comparison, and storing the target data with a smaller value in the register with a lower or higher rank among the two registers participating in the comparison; The third sorting operation includes: starting from the second register, sequentially comparing the target data in two adjacent registers, storing the target data with a larger value in the register with a higher or lower rank among the two registers participating in the comparison, and storing the target data with a smaller value in the register with a lower or higher rank among the two registers participating in the comparison.
[0009] Optionally, the hardware accelerator further includes: A tightly coupled memory for storing the target data; A register prefetch unit, which is respectively connected to the scheduling unit, the register, and the tightly coupled memory. After receiving an instruction from the scheduling unit, the register prefetch unit fetches the target data from the tightly coupled memory and loads the target data into their respective corresponding registers according to the instruction.
[0010] Optionally, a first data multiplexer is provided between the register and the basic sorting unit. The first data multiplexer is connected to the scheduling unit and, under the control of the scheduling unit, selectively determines one of its input data as the output data and sends it to the basic sorting unit; A second data multiplexer is provided between the register prefetch unit and the register. The second data multiplexer is connected to the scheduling unit and, under the control of the scheduling unit, selectively determines one of its input data as the output data and sends it to the register.
[0011] Optionally, the hardware accelerator further includes: A result write-back unit, which is respectively connected to the basic sorting unit, the scheduling unit, and the tightly coupled memory. The result write-back unit stores the sorting result of the sorting operation on the target data under the control of the scheduling unit.
[0012] Optionally, the hardware accelerator further includes: A memory access controller, which is respectively connected to the tightly coupled memory and the memory. It is used to obtain the target data from the memory according to a data call request of the tightly coupled memory, or send the sorting result of the sorting operation on the target data to the memory.
[0013] The data sorting method according to the embodiment of the present application alternately uses a first sorting operation and an intermediate sorting operation on the target data, simply and quickly realizes the sorting of the target data, improves the sorting efficiency of the target data, and reduces the time complexity in the sorting process. Description of the Drawings
[0014] Figure 1 It is a flowchart of the data sorting method according to the embodiment of the present application; Figure 2 It is a flowchart of an embodiment of the data sorting method according to the embodiment of the present application; Figure 3 It is a flowchart of another embodiment of the data sorting method according to the embodiment of the present application; Figure 4 It is a schematic diagram of the first sorting operation and the second sorting operation according to the embodiment of the present application; Figure 5 Schematic diagram of the first sorting operation and the third sorting operation of the embodiment of the present application; Figure 6 Schematic diagram of the first sorting operation and the second sorting operation after grouping data units in the embodiment of the present application; Figure 7 Schematic diagram of the first sorting operation and the third sorting operation after grouping data units in the embodiment of the present application; Figure 8 Schematic diagram of the structure of the hardware accelerator of the embodiment of the present application. Detailed implementation manners
[0015] Reference is made herein to the various solutions and features of the present application with reference to the accompanying drawings.
[0016] It should be understood that various modifications can be made to the embodiments applied herein. Therefore, the above specification should not be regarded as a limitation, but merely as an example of the embodiments. Those skilled in the art will think of other modifications within the scope and spirit of the present application.
[0017] The accompanying drawings included in and constituting a part of this specification illustrate embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, are used to explain the principles of the present application.
[0018] These and other features of the present application will become apparent from the following description of the preferred forms of the embodiments given by way of non-limiting example with reference to the accompanying drawings.
[0019] It should also be understood that although the present application has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present application.
[0020] When combined with the accompanying drawings, the above and other aspects, features and advantages of the present application will become more apparent in view of the following detailed description.
[0021] The specific embodiments of the present application are hereinafter described with reference to the accompanying drawings; however, it should be understood that the embodiments applied are merely examples of the present application and can be implemented in various ways. Well-known and / or repeated functions and structures are not described in detail to avoid unnecessary or redundant details from obscuring the present application. Therefore, the specific structural and functional details applied herein are not intended to be limiting, but merely serve as a basis for the claims and a representative basis for teaching those skilled in the art to use the present application in substantially any suitable detailed structure in a variety of ways.
[0022] This specification may use phrases such as "in one embodiment", "in another embodiment", "in yet another embodiment", or "in other embodiments", which may each refer to one or more of the same or different embodiments according to the present application.
[0023] A data sorting method according to an embodiment of the present application. This method can be applied in scenarios such as high-performance computing, big data analysis, and artificial intelligence. In the process of processing target data in these scenarios, numerical sorting processing needs to be performed on the target data to facilitate sequential reading and use of the target data. Specifically, this data sorting method can be applied to an electronic device. Multiple target data to be processed are respectively stored in multiple registers of the electronic device, and then multiple sorting operations are performed on the target data in the registers to gradually achieve the sorting of the target data. This data sorting method can be sorting the target data from largest to smallest in terms of numerical value, or it can be sorting the target data from smallest to largest in terms of numerical value.
[0024] The following will specifically describe this data sorting method according to an embodiment of the present application in conjunction with specific embodiments. Figure 1 It is a flowchart of the data sorting method according to an embodiment of the present application, as Figure 1 shown. This method includes the following steps: S100, based on the arrangement order of the registers, store multiple target data into their respective corresponding registers.
[0025] Exemplarily, the target data is the object to be processed and needs to be sorted according to its numerical value. The registers of the electronic device are used to store the target data, and they have an arrangement order so that the corresponding target data can be stored respectively according to the arrangement order. During the process of processing the target data, the arrangement order of the registers does not need to be adjusted, but only the storage positions of the target data need to be adjusted.
[0026] In one embodiment, one register can store one target data. After all the target data are stored in the registers, since the registers have an arrangement order, the target data forms an arrangement order.
[0027] S200, perform a first sorting operation on the target data in the registers, which includes: starting from the first register with the first serial number, sequentially compare the target data in two adjacent registers respectively, store the target data with a larger numerical value into the register with a higher or lower sorting position among the two registers participating in the comparison, and store the target data with a smaller numerical value into the register with a lower or higher sorting position among the two registers participating in the comparison.
[0028] Exemplarily, in combination with Figure 4 and Figure 5, according to the arrangement order of the registers, the registers can be numbered. For example, the target data are 0, 1, 2, 3, 4, 5, 6, 7 respectively. The serial number of the first register x(0) arranged at the first position is the first serial number, the serial number of the second register x(1) arranged at the second position is the second serial number, the serial number of the third register x(2) arranged at the third position is the third serial number, …, the serial number of the eighth register x(7) arranged at the last position is the penultimate serial number. Among them, 0 is stored in the first register x(0), 1 is stored in the second register x(1), 2 is stored in the third register x(2), 3 is stored in the fourth register x(3), 4 is stored in the fifth register x(4), 5 is stored in the sixth register x(5), 6 is stored in the seventh register x(6), and 7 is stored in the eighth register x(7).
[0029] Starting from the first register x(0) with the first serial number, a first sorting operation is performed on each target data in the registers, including comparing the target data 0 in the first register x(0) with the target data 1 in the second register x(1), storing the target data 1 with the larger value in the first register x(0) with the earlier sorting among the two registers participating in the comparison, and storing the target data 0 with the smaller value in the second register x(1) with the later sorting among the two registers participating in the comparison. Similarly, compare the target data 2 in the third register x(2) with the target data 3 in the fourth register x(3), store the target data 3 with the larger value in the third register x(2) with the earlier sorting among the two registers participating in the comparison, and store the target data 2 with the smaller value in the fourth register x(3) with the later sorting among the two registers participating in the comparison. Use the above similar sorting operations to arrange all the target data.
[0030] After one first sorting operation, the arrangement order of the target data is: 1, 0, 3, 2, 5, 4, 7, 6. They are stored in the first register x(0), the second register x(1), the third register x(2), the fourth register x(3), the fifth register x(4), the sixth register x(5), the seventh register x(6), and the eighth register x(7) respectively.
[0031] S300 performs a sub-operation in the intermediate sorting operation on the target data in the register. The sub-operation includes a second sorting operation and a third sorting operation. The second sorting operation includes: starting from the second register with the second serial number, sequentially comparing the target data in two adjacent registers, and comparing the target data in the first register with the register with the last serial number. Store the target data with a larger value in the register with a higher or lower sort order among the two registers participating in the comparison, and store the target data with a smaller value in the register with a lower or higher sort order among the two registers participating in the comparison. The third sorting operation includes: starting from the second register, sequentially comparing the target data in two adjacent registers, storing the target data with a larger value in the register with a higher or lower sort order among the two registers participating in the comparison, and storing the target data with a smaller value in the register with a lower or higher sort order among the two registers participating in the comparison.
[0032] Exemplarily, the intermediate sorting operation is a sorting operation for another intermediate step after the first sorting operation on the target data. The intermediate sorting operation includes multiple sub-operations, and the sub-operation includes a second sorting operation and a third sorting operation. In this embodiment, after the first sorting operation on the target data, it is necessary to perform a sub-operation in the intermediate sorting operation on the target data. For example, after the first sorting operation on the target data, the second sorting operation can be performed on the target data, or after the first sorting operation on the target data, the third sorting operation (without performing the second sorting operation) can be performed on the target data. Thus, a sorting operation is performed on the target data again.
[0033] In one embodiment, continue to describe the second sorting operation in combination with the above embodiment. As Figure 4 shown, after the first sorting operation on the target data, the second sorting operation is performed on the target data. Among them, the second sorting operation includes starting from the second register x(1) with the second serial number, sequentially comparing the target data in two adjacent registers. First, compare the target data 0 in the second register x(1) with the target data 3 in the third register x(2). Store the target data 3 with a larger value in the second register x(1) with a higher sort order, and store the target data 0 with a smaller value in the third register x(2) with a lower sort order. Similarly, perform the above similar operations on the fourth register to the seventh register.
[0034] In addition, the second sorting operation further includes comparing the target data 1 in the first register x(0) with the target data 6 in the eighth register x(7) having the last serial number, storing the target data 6 with a larger value in the first register x(0) with a higher sorting order, and storing the target data 1 with a smaller value in the eighth register x(7) with a lower sorting order.
[0035] In another embodiment, continuing with the above embodiment to illustrate the third sorting operation, as Figure 5 shown, after the first sorting operation on the target data, the second sorting operation is not performed on the target data, but the third sorting operation is performed on the target data. Among them, the third sorting operation includes starting from the second register x(1) with the second serial number, sequentially comparing the target data in two adjacent registers. First, compare the target data 0 in the second register x(1) with the target data 3 in the third register x(2), store the target data 3 with a larger value in the second register x(1) with a higher sorting order, and store the target data 0 with a smaller value in the third register x(2) with a lower sorting order. Similarly, perform the above similar operations on the fourth to seventh registers.
[0036] In addition, the target data 1 in the first register x(0) is not adjusted in the third sorting operation. Thus, the target data 1 is kept stored in the first register x(0). The target data 6 in the eighth register x(7) is not adjusted in the third sorting operation. Thus, the target data 6 is kept stored in the eighth register x(7).
[0037] In the above embodiments, it is illustrated that when comparing the target data, the target data with a larger value is stored in the register with a higher sorting order, and the target data with a smaller value is stored in the register with a lower sorting order. Of course, when comparing the target data, according to the specific sorting target, the target data with a larger value can also be stored in the register with a lower sorting order, and the target data with a smaller value can be stored in the register with a higher sorting order.
[0038] S400, repeatedly perform the first sorting operation and the second sorting operation alternately for multiple times, or repeatedly perform the first sorting operation and the third sorting operation alternately for multiple times until all the target data are sorted in descending or ascending order of values.
[0039] Exemplarily, as Figure 4As shown, after a sub-operation in the intermediate sorting operation on the target data in the register, the first sorting operation and the second sorting operation (a sub-operation in the intermediate sorting operation) need to be alternately repeated multiple times until all the target data are sorted in descending or ascending order of numerical value. For example, a total of 7 sorting operations (including multiple first sorting operations and multiple second sorting operations) are required, so as to achieve the on-demand arrangement of the target data, such as the arranged order being 7, 6, 5, 4, 3, 2, 1, 0.
[0040] Or, as Figure 5 shown, after a sub-operation in the intermediate sorting operation on the target data in the register, the first sorting operation and the third sorting operation (a sub-operation in the intermediate sorting operation) need to be alternately repeated multiple times until all the target data are sorted in descending or ascending order of numerical value. For example, a total of 8 sorting operations (including multiple first sorting operations and multiple third sorting operations) are required, so as to achieve the on-demand arrangement of the target data, such as the arranged order being 7, 6, 5, 4, 3, 2, 1, 0.
[0041] In this embodiment, the sorted target data can be stored in the memory of the electronic device for easy calling.
[0042] The data sorting method of this embodiment of the present application realizes the sorting of the target data simply and quickly by alternately using the first sorting operation and the intermediate sorting operation on the target data, improves the sorting efficiency of the target data, and reduces the time complexity in the sorting process.
[0043] In an embodiment of the present application, in the case where there are multiple data units in the target data and they are not sorted, as Figure 2 shown, the method further includes the following steps: S10, based on the arrangement order of the register, store the multiple data units into their respective corresponding registers.
[0044] Exemplarily, the target data can be a single data unit or have multiple data units. The numerical values of the multiple data units may be different, which requires sorting the multiple data units, including sorting in descending or ascending order of numerical value. After sorting the data units, sort the multiple target data. The specific sorting method can be based on the above content.
[0045] S20, take the data units in the register as objects and perform the first sorting operation on them.
[0046] Exemplarily, the register can not only be used to store target data for sorting the target data, but also be used to store data units for sorting the data units. Similar to sorting the target data, the data units in the register are used as objects and a first sorting operation is performed on them. The first sorting operation includes: starting from the first register with the first serial number, sequentially comparing the data units in two adjacent registers, storing the data unit with a larger value in the register with a higher or lower sort order among the two registers participating in the comparison, and storing the data unit with a smaller value in the register with a lower or higher sort order among the two registers participating in the comparison.
[0047] S30. Use the data units in the register as objects and perform the intermediate sorting operation on them.
[0048] Exemplarily, the sub-operations include a second sorting operation and a third sorting operation. The second sorting operation includes: starting from the second register with the second serial number, sequentially comparing the data units in two adjacent registers, and comparing the first register with the register with the last serial number, storing the data unit with a larger value in the register with a higher or lower sort order among the two registers participating in the comparison, and storing the data unit with a smaller value in the register with a lower or higher sort order among the two registers participating in the comparison; The third sorting operation includes: starting from the second register, sequentially comparing the data units in two adjacent registers, storing the data unit with a larger value in the register with a higher or lower sort order among the two registers participating in the comparison, and storing the data unit with a smaller value in the register with a lower or higher sort order among the two registers participating in the comparison.
[0049] S40. Repeatedly alternate between the first sorting operation and the second sorting operation multiple times, or repeatedly alternate between the first sorting operation and the third sorting operation multiple times until all the data units are sorted from largest to smallest or from smallest to largest in value.
[0050] Exemplarily, as Figure 4 shown, after performing one of the sub-operations in the intermediate sorting operation on the data units in the register, it is necessary to repeatedly alternate between the first sorting operation and the second sorting operation multiple times until all the data units are sorted from largest to smallest or from smallest to largest in value.
[0051] Or, as Figure 5 shown, after performing one of the sub-operations in the intermediate sorting operation on the data units in the register, it is necessary to repeatedly alternate between the first sorting operation and the third sorting operation multiple times until all the data units are sorted from largest to smallest or from smallest to largest in value.
[0052] In one embodiment of the present application, when there are multiple data units in the target data and they have been sorted, as Figure 3 shown and in combination with Figure 6 and Figure 7 , the method further includes: S51, group the data units based on the numerical magnitudes of the data units to form multiple data unit groups.
[0053] Exemplarily, after the multiple data units are sorted, all the data units can be regarded as one target data, and the multiple target data can be sorted using the sorting method described above. Alternatively, the multiple data units can be grouped to form multiple data unit groups, with each data unit group being used as a processing object and sorted.
[0054] For example, after the multiple data units are sorted as "7, 6, 5, 4, 3, 2, 1, 0". Based on the numerical magnitudes of the data units "7, 6, 5, 4, 3, 2, 1, 0", the data units "7, 6, 5, 4, 3, 2, 1, 0" are divided into two groups, forming the first data unit group "7, 6, 5, 4" and the second data unit group "3, 2, 1, 0".
[0055] S52, group the registers based on the arrangement order of the registers to form multiple register groups.
[0056] Exemplarily, the registers can be grouped based on the number of data unit groups and the number of data units therein, and based on the arrangement order of the registers. For example, the registers can be divided into 8 register groups, namely DG0, DG1, DG2, DG3, DG4, DG5, DG6, DG7, and each register group stores the data of one data unit group. Combining the above embodiments, specifically, each register group can store 4 data units.
[0057] Preferably, the step of grouping the data units based on the numerical magnitudes of the data units to form multiple data unit groups includes: dividing the data units into 2 data unit groups based on the numerical magnitudes of the data unit groups; such as the first data unit group "7, 6, 5, 4" and the second data unit group "3, 2, 1, 0" formed above. Correspondingly, the step of grouping the registers to form multiple register groups includes: constructing 1 register group based on 4 registers. Each register group among DG0, DG1, DG2, DG3, DG4, DG5, DG6, DG7 described above includes 4 registers.
[0058] S53, store the data unit groups into their respective corresponding register groups based on the arrangement order of the register groups.
[0059] Exemplarily, the register bank has an arrangement order, such as sorting in the order of the first register bank DG0, the second register bank DG1, the third register bank DG2, the fourth register bank DG3, the fifth register bank DG4, the sixth register bank DG5, the seventh register bank DG6, and the eighth register bank DG7. When storing the data unit groups into the register bank respectively, the storage operation can be performed according to the arrangement order of the register bank and the arrangement order of the data unit groups. For example, store the first data unit group into the first register bank DG0, store the second data unit group into the second register bank DG1, …, and store the eighth data unit group into the eighth register bank DG7.
[0060] S54. Take the data unit group in the register bank as an object and perform the first sorting operation on it.
[0061] Exemplarily, similar to the process of processing the target data in the register in the above embodiment, take the data unit group in the register bank as an object and perform the first sorting operation on it, which includes: starting from the first register bank with the first serial number, sequentially compare the data unit groups in two adjacent register banks, store the data unit group with a larger value into the register bank with a higher or lower sort order among the two register banks participating in the comparison, and store the data unit group with a smaller value into the register bank with a lower or higher sort order among the two register banks participating in the comparison.
[0062] S55. Take the data unit group in the register bank as an object and perform a sub - operation of the intermediate sorting operation on it.
[0063] Exemplarily, similar to the process of processing the target data in the register in the above embodiment, after performing the first sorting operation on the data unit group in the register bank, take the data unit group as an object and perform a sub - operation of the intermediate sorting operation on it. This sub - operation includes a second sorting operation and a third sorting operation.
[0064] S56. Repeatedly alternate between the first sorting operation and the second sorting operation multiple times, or repeatedly alternate between the first sorting operation and the third sorting operation multiple times until all the data unit groups are sorted from largest to smallest or from smallest to largest in terms of value.
[0065] Exemplarily, such as Figure 6As shown, after a sub - operation in the intermediate sorting operation on the data unit groups in the register bank, the first sorting operation and the second sorting operation (a sub - operation in the intermediate sorting operation) need to be alternately repeated multiple times until all the data unit groups are sorted in descending or ascending order of values. The value of a data unit group can be a comprehensive value, such as the value obtained by accumulating all data unit groups, or a value obtained through other calculations.
[0066] Or, as Figure 7 shown, after a sub - operation in the intermediate sorting operation on the data unit groups in the register bank, the first sorting operation and the third sorting operation (a sub - operation in the intermediate sorting operation) need to be alternately repeated multiple times until all the data unit groups are sorted in descending or ascending order of values.
[0067] The following combines Figures 4 to 7 to illustrate a specific embodiment of the data sorting method of this application.
[0068] When performing the first sorting operation and the second sorting operation on the target data: The X() sequence is the sorting input array.
[0069] The sorting unit is composed of multiple sub - units. Each sub - unit receives two data inputs X(i) and X(j), and assigns the maximum value of X(i) and X(j) to X(i) and the minimum value to X(j) through comparison. If the two data are equal, the values of X(i) and X(j) remain unchanged. If a value swap occurs, a value swap signal is output simultaneously.
[0070] Taking the sequence length of 8 as an example. The flow of signals is divided into two states: The first sorting operation (A state), pairwise sorting of adjacent data from the first data to the last data.
[0071] The second sorting operation (B state), pairwise sorting of adjacent data from the second data to the second - last data. At the same time, pairwise sorting of the first and the last data.
[0072] Continuously repeat states A and B. After 7 rounds of sorting, the sorting is completed. If in two consecutive rounds of sorting, no order swap occurs in each pairwise sorting unit, the sorting can end prematurely.
[0073] The time complexity of the algorithm is Ot = Nx(N - 1) / 2, that is, for a sequence of length N, at most Ot comparisons are required to complete the sorting. The space complexity is Os = N, that is, N data need to be stored during the sorting of a sequence of length N.
[0074] In another embodiment, when performing the first sorting operation and the third sorting operation on the target data: In the third sorting operation (B state), the sorting of the first and last data is discarded, but the number of sorting rounds needs to be increased from 7 rounds to 8 rounds. At the same time, the early exit condition needs to be modified to the B state (third sorting operation) and no data exchange occurs in the A state (first sorting operation) before this B state.
[0075] The space complexity of the variant algorithm is the same as the basic form, and the time complexity increases to Ot = NxN / 2.
[0076] In yet another embodiment, when there are multiple data units in the target data and they are not sorted, group sorting is performed on the data units: This group sorting meets the requirements of sorting ultra-long sequences.
[0077] First, the entire sequence needs to be evenly divided into multiple data unit groups. When the sequence length cannot be evenly divided by the number of groups, the last group can be filled with the minimum or maximum value to make each group of equal length.
[0078] Taking the example of dividing the sequence into 8 equal parts, after grouping, the sequence is divided into 8 groups corresponding to 8 register groups (DG0~DG7), and each group contains 4 data units.
[0079] After grouping, merge sorting is first performed between every two groups. For example, the 8 data units contained in DG0 and DG1 can use the above data sorting method.
[0080] The operations between each data unit group can also be divided into two states, A and B states.
[0081] The first sorting operation (A state) performs merge sorting on adjacent groups from the first group to the last group.
[0082] The second sorting operation (B state) performs merge sorting on adjacent groups from the second group to the second-to-last group. At the same time, the first and last groups are merged and sorted, with the larger number written back to the first group and the smaller number written back to the last group.
[0083] States A and B are continuously repeated, and the sorting is completed after 7 rounds of sorting. If there is no order swap in each pairwise sorting unit for two consecutive rounds of sorting, the sorting can be ended early.
[0084] Group sorting of data unit groups can also use the third sorting operation. In the third sorting operation (B state), the first and last data unit groups do not participate in the sorting, and one more round of sorting operation is required overall.
[0085] The grouping and sorting method of this data unit group can be further evolved into a hierarchical grouping structure.
[0086] First, the entire sequence can be divided into multiple first-level groups of equal length. Each first-level group can be further divided into multiple second-level groups of equal length, and the second-level groups can continue to be subdivided until a hierarchical structure that matches the hardware storage structure is formed. Taking the structure of a digital signal processing vector machine (DSP) as an example, its typical storage mechanism is as follows: 16 x 32-bit operation core registers, 128KB core tightly coupled cache, and 1GB external DRAM storage. Based on the above structure, for example, when using a DSP to perform a sorting operation on 1024 x 1024 (1M) data, it can be done as follows: A sequence with a length of 1024 x 1024 is first divided into 64 first-level groups of 16 x 1024 data. Each first-level group requires 16 x 1024 x 32bit = 64KB of storage space (assuming each data occupies 32bit of storage space). So, during the operation process, two 64KB first-level groups can be stored in the 128KB tightly coupled cache of the core first and then processed.
[0087] Each first-level group can be further divided into 2048 second-level groups of 8 data. Each second-level group requires 8 x 32bit = 256bit of storage space. Every two 256bit second-level groups can be stored using the core register with a capacity of 16 x 32-bit.
[0088] Through the above division, when sorting the second-level groups, only access to the core register is required, and when sorting the first-level groups, only access to the on-chip tightly coupled cache is required. During the entire calculation process, the demand for different types of storage access bandwidth is greatly reduced.
[0089] After the grouping is completed, first, the basic wave sorting algorithm is used to sort the lowest-level groups. Then, the above data sorting algorithm is used to sort the groups above the lowest level until the highest-level group.
[0090] In an embodiment of this application, the method further includes: sorting the target data based on a preset time interval; sending the sorted target data to the memory through a memory access controller.
[0091] Exemplarily, during the use of the target data, its value may change, such that the sorted target data is not sorted in a descending or ascending order. In this embodiment, the target data can be sorted multiple times based on a preset time interval, so as to dynamically adjust the sorting result of the target data, and to a certain extent, keep the sorting result of the target data meeting the usage requirements. In addition, after sorting, the target data can be sent to a memory storage (such as a DRAM memory) through the memory access controller of the electronic device, and the memory storage stores and calls the target data.
[0092] The embodiment of the present application further provides a hardware accelerator, which is used to perform a sorting operation on target data. As Figure 8 shown, the hardware accelerator is connected to the memory storage, The hardware accelerator includes: A plurality of sequentially arranged registers; A scheduling unit, which is connected to the registers, and is used to store the target data into their respective corresponding registers; A plurality of basic sorting units, which are respectively connected to two adjacent registers corresponding thereto. The basic sorting unit is configured to: perform a first sorting operation on the target data in the registers under the control of the scheduling unit; perform one of the sub-operations in the intermediate sorting operation on the target data in the registers, and the sub-operations include a second sorting operation and a third sorting operation; repeatedly perform the first sorting operation and the second sorting operation alternately for multiple times, or repeatedly perform the first sorting operation and the third sorting operation alternately for multiple times, until all the target data is sorted in a descending order or an ascending order of values.
[0093] Exemplarily, the target data is the object to be processed, and it needs to be sorted according to its value. The registers of the electronic device are used to store the target data, and they have an arrangement order so that the corresponding target data can be stored respectively according to the arrangement order. During the process of processing the target data, the arrangement order of the registers can be not adjusted, but only the storage position of the target data is adjusted.
[0094] In one embodiment, one register can store one target data. After all the target data are stored in the registers, since the registers have an arrangement order, the target data forms an arrangement order.
[0095] The hardware accelerator further includes a scheduling unit, which is connected to a plurality of registers, and is used to store the target data obtained from the memory storage of the electronic device into their respective corresponding registers.
[0096] Each basic sorting unit is respectively connected to two corresponding adjacent registers and is connected to the scheduling unit. For example, basic sorting unit 0 is respectively connected to register 0 and register 1, basic sorting unit 1 is respectively connected to register 2 and register 3, basic sorting unit 2 is respectively connected to register 4 and register 5, …, basic sorting unit N-2 is respectively connected to register 2N-4 and register 2N - 3, and basic sorting unit N-1 is respectively connected to register 2N-2 and register 2N-1. At the same time, it is connected to the scheduling unit. Under the control of the scheduling unit, each basic sorting unit performs a first sorting operation on the target data in the corresponding register. For example, basic sorting unit 0 performs a sorting operation on the target data in register 0 and the target data in register 1, basic sorting unit 1 performs a sorting operation on the target data in register 2 and the target data in register 3, and so on.
[0097] The basic sorting unit can use a selector to compare the target data in two adjacent registers. Continuing with the above embodiment, the first sorting operation will be described.
[0098] Starting from having register 0, basic sorting unit 0 compares the target data in two adjacent registers 0 and 1, stores the target data with a larger value in register 0, and stores the target data with a smaller value in register 1. Basic sorting unit 1 compares the target data in two adjacent registers 2 and 3, stores the target data with a larger value in register 2, and stores the target data with a smaller value in register 3, …, basic sorting unit N-1 compares the target data in two adjacent registers 2N-2 and 2N-1, stores the target data with a larger value in register 2N-2, and stores the target data with a smaller value in register 2N-1.
[0099] After each basic sorting unit performs the first sorting operation on the target data in the corresponding register under the control of the scheduling unit; it then performs a sub-operation of an intermediate sorting operation on the target data in the corresponding register. The intermediate sorting operation is a sorting operation of another intermediate step after the first sorting operation on the target data. The intermediate sorting operation includes multiple sub-operations, and this sub-operation includes a second sorting operation and a third sorting operation. In this embodiment, each basic sorting unit needs to perform a sub-operation of an intermediate sorting operation on the target data after performing the first sorting operation on the target data. For example, after the basic sorting unit performs the first sorting operation on the target data, it can then perform the second sorting operation on the target data, or after performing the first sorting operation on the target data, it can then perform the third sorting operation (without performing the second sorting operation) on the target data. Thus, the target data is sorted again.
[0100] For the second sorting operation, it is described in combination with the above embodiments. The basic sorting unit performs the second sorting operation as follows: Starting from register 1, the basic sorting unit 1 compares the target data in two adjacent registers, register 1 and register 2, stores the target data with a larger value in register 1, and stores the target data with a smaller value in register 2. The basic sorting unit 2 compares the target data in two adjacent registers, register 3 and register 4, stores the target data with a larger value in register 3, and stores the target data with a smaller value in register 4,... The basic sorting unit N - 1 compares the target data in two adjacent registers, register 2N - 3 and register 2N - 2, stores the target data with a larger value in register 2N - 3, and stores the target data with a smaller value in register 2N - 2.
[0101] The basic sorting unit performing the second sorting operation further includes: The basic sorting unit 0 compares the target data in two adjacent registers, register 0 and register 2N - 1, stores the target data with a larger value in register 0, and stores the target data with a smaller value in register 2N - 1.
[0102] For the third sorting operation, it is described in combination with the above embodiments. The basic sorting unit performs the third sorting operation as follows: Starting from register 1, the basic sorting unit 1 compares the target data in two adjacent registers, register 1 and register 2, stores the target data with a larger value in register 1, and stores the target data with a smaller value in register 2. The basic sorting unit 2 compares the target data in two adjacent registers, register 3 and register 4, stores the target data with a larger value in register 3, and stores the target data with a smaller value in register 4,... The basic sorting unit N - 1 compares the target data in two adjacent registers, register 2N - 3 and register 2N - 2, stores the target data with a larger value in register 2N - 3, and stores the target data with a smaller value in register 2N - 2.
[0103] Among them, in the third sorting operation, the basic sorting unit 0 does not adjust the target data in register 0, but keeps the target data in register 0. The basic sorting unit N - 1 does not adjust the target data in register 2N - 1, but keeps the target data in register 2N - 1.
[0104] In an embodiment of the present application, the hardware accelerator further includes: A tightly - coupled memory for storing the target data; A register prefetch unit is respectively connected to the scheduling unit, the register, and the tightly coupled memory. After receiving an instruction from the scheduling unit, the register prefetch unit obtains the target data from the tightly coupled memory and loads the target data into their respective corresponding registers according to the instruction.
[0105] Exemplarily, the tightly coupled memory is respectively connected to the scheduling unit and the memory access controller. The tightly coupled memory obtains and stores the target data from the memory through the memory access controller. The register prefetch unit is respectively connected to the scheduling unit, the register, and the tightly coupled memory. When performing a sorting operation on the target data, the scheduling unit sends relevant instructions to the register prefetch unit. After receiving the instruction from the scheduling unit, the register prefetch unit obtains the target data from the tightly coupled memory and loads the target data into their respective corresponding registers according to the instruction. For example, the target data obtained by the register prefetch unit is "0, 1, 2, 3, 4, 5, 6, …, 2N - 2, 2N - 1". The target data 0 is loaded into register 0, the target data 1 is loaded into register 1, the target data 2 is loaded into register 2, …, the target data 2N - 2 is loaded into register 2N - 2, and the target data 2N - 1 is loaded into register 2N - 1.
[0106] In an embodiment of the present application, a first data multiplexer is provided between the register and the basic sorting unit. The first data multiplexer is connected to the scheduling unit and, under the control of the scheduling unit, selectively determines one of its input data as the output data and sends it to the basic sorting unit; A second data multiplexer is provided between the register prefetch unit and the register. The second data multiplexer is connected to the scheduling unit and, under the control of the scheduling unit, selectively determines one of its input data as the output data and sends it to the register.
[0107] Exemplarily, under the control of the scheduling unit, the first data multiplexer can selectively select one of the two input data for output. For example, the first data multiplexer outputs the input data to the corresponding basic sorting unit, and the basic sorting unit can correspond to one or two first data multiplexers. The basic sorting unit then compares the received target data.
[0108] Illustratively, the first data multiplexer in the first position sends one of the target data in register 1 and register 2 to the basic sorting unit 1, and the first data multiplexer in the second position sends one of the target data in register 2 and register 3 to the basic sorting unit 1. The basic sorting unit 1 compares the two received target data and outputs the target data with the larger value.
[0109] A second data multiplexer is provided between the register prefetch unit and the register. The second data multiplexer is connected to the scheduling unit and the output end of the basic sorting unit. Under the control of the scheduling unit, the second data multiplexer alternatively determines one of its input data as the output data and sends it to the register for storage.
[0110] In an embodiment of the present application, the hardware accelerator further includes a result writing-back unit. The result writing-back unit is respectively connected to the basic sorting unit, the scheduling unit, and the tightly coupled memory. Under the control of the scheduling unit, the result writing-back unit stores the sorting result of the sorting operation for the target data, and sends the sorting result to the memory through the tightly coupled memory and the memory access controller.
[0111] Next Figures 4 to 8 , a specific embodiment of the hardware accelerator in the embodiment of the present application will be described in detail.
[0112] The register is used to temporarily store the sorted input packet data (data unit or data unit group) and the sorted intermediate result data.
[0113] The second data multiplexer at the front end of the register selects the stored data according to the register load enable signal output by the scheduling unit.
[0114] When the load enable signal is high, the input data of the register prefetch unit is loaded into the register.
[0115] When the load enable signal is low, the result data of the basic sorting unit is loaded into the register for the next sorting use.
[0116] The basic sorting unit is used to compare the input data (such as target data or data unit): When in_a >= in_b, the output max is selected as in_a, and at the same time the output min is selected as in_b, and the "not swapped" output signal is kept as 1.
[0117] When in_a < in_b, the output max is selected as in_b, and at the same time the output min is selected as in_a, and the "not swapped" output signal is changed to 0.
[0118] The first data multiplexer in the front stage of the basic sorting unit selects the input data according to the first sorting operation (type A) or the intermediate sorting operation (type B) selection signal output by the scheduling unit (subsequently, the first sorting operation can be represented by type A, and the intermediate sorting operation can be represented by type B).
[0119] When the sorting is in the first sorting operation (type A), select the corresponding register output as the input of the basic sorting unit. Taking the basic sorting unit N-2 as an example, in type A, select the "register 2N-4" and "register 2N-3" with the serial numbers 2x(N-2) and 2x(N-2)+1 as the inputs.
[0120] When the sorting is in the intermediate sorting operation (type B), select the outputs of the registers adjacent to the corresponding registers as the inputs of the basic sorting unit. Similarly, taking the basic sorting unit N-2 as an example, in type B, select the "register 2N-5" and "register 2N-4" with the serial numbers 2x(N-2)-1 and 2x(N-2) as the inputs. The first basic sorting unit is slightly different: The in_a and max data paths of the first basic sorting unit remain unchanged; the in_b input is selected as the output of the last register, and at the same time, the min output is selected as the max output of the next basic sorting unit.
[0121] The third data multiplexer after the basic sorting unit selects the output data according to the selection signal of the first sorting operation (type A) or the intermediate sorting operation (type B) output by the scheduling unit.
[0122] When the sorting is in type A, keep the order of the output results unchanged.
[0123] When the sorting is in type B, the max output result of the basic sorting unit is replaced by the min output of this basic sorting unit; at the same time, the min output result is replaced by the max output of the adjacent basic sorting unit. Taking the basic sorting unit N-2 as an example, in type B, replace the max output of the basic sorting unit N-2 with the min output, and at the same time, use the max output of the basic sorting unit N-1 to replace the min output of the basic sorting unit N-2. The first and the last basic sorting units are slightly different: The max output of the first basic sorting unit remains unchanged in the type B sorting operation.
[0124] The min value of the last basic sorting unit is replaced by the min value of the first basic sorting unit in the type B sorting operation.
[0125] The register prefetch unit receives the preloading enable signal from the scheduling unit to preload the sorting data.
[0126] After receiving the preloading enable signal, the register prefetch unit requests data from the tightly coupled cache according to the preset secondary packet data length (LEN2), primary packet data length (LEN1), and the base address (BA2) of the data in the tightly coupled memory.
[0127] The register prefetch unit prefetches data from the tightly coupled cache according to the indication of "LD2_AB type selection" of the scheduling unit: When in the A-type secondary sorting cycle, starting from the base address BA2, after receiving a load request, two secondary groups, that is, LEN2x2 data, are sequentially requested from the tightly coupled cache each time, and the corresponding data is placed at the register input port for the register to use. After each request is completed, the requested data address is incremented by LEN2x2. When the amount of requested data reaches LEN1, the requested address is restored to the initial BA2.
[0128] When in the B-type secondary sorting cycle, starting from the base address BA2 + LEN2, after receiving a load request, two secondary groups, that is, LEN2x2 data, are sequentially requested from the tightly coupled cache each time, and the corresponding data is placed at the register input port for the register to use. After each request is completed, the requested data address is incremented by LEN2x2. When the amount of requested data reaches LEN1 - LEN2x2, the requested address is restored to the initial BA2. For the sake of simplifying the hardware complexity, a variant of the grouped wave sorting algorithm is used for secondary group sorting.
[0129] The result write-back unit receives the result write-back enable signal, saves the result output by the basic comparison unit, and writes it back to the tightly coupled cache.
[0130] After receiving the result write-back enable signal, the result write-back unit first holds the calculation result, and then writes back the data to the tightly coupled cache according to the pre-set secondary group data length (LEN2), primary group data length (LEN1), and the base address (BA2) of the data in the tightly coupled memory.
[0131] The result write-back unit writes back the result data to the tightly coupled cache according to the indication of "ST2_AB type selection" of the scheduling unit: When in the A-type secondary sorting cycle, starting from the base address BA2, after receiving a write-back request, the result data is first saved, and two secondary groups, that is, LEN2x2 data, are sequentially written back to the tightly coupled cache. After each write-back is completed, the write-back data address is incremented by LEN2x2. When the amount of write-back data reaches LEN1, the write-back address is restored to the initial BA2.
[0132] When in the B-type secondary sorting cycle, starting from the base address BA2 + LEN2x2, after receiving a write-back request, the result data is first saved, and two secondary groups, that is, LEN2x2 data, are sequentially written back to the tightly coupled cache. After each write-back is completed, the write-back data address is incremented by LEN2x2. When the amount of write-back data reaches LEN1 - LEN1x2, the write-back address is restored to the initial BA2.
[0133] The tightly coupled cache is used to store the first-level packet data.
[0134] After receiving the memory load enable signal, the tightly coupled cache requests data from the DRAM according to the pre-set first-level packet data length (LEN1), the overall data length (LEN), and the base address (BA1) of the data in the DRAM memory.
[0135] The tightly coupled cache reads data from the DRAM memory through the memory access controller according to the indication of the scheduling unit "LD1_AB type selection": When in the A-type sorting cycle, starting from the base address BA1, after receiving the load request, two first-level packets, that is, LEN1x2 data, are sequentially requested from the DRAM memory each time, and the corresponding data is placed in the tightly coupled cache. After each request is completed, the request data address is incremented by LEN1x2. When the amount of requested data reaches LEN, the request address is restored to the initial BA1.
[0136] When in the B-type sorting cycle, starting from the base address BA1+LEN1, after receiving the load request, two first-level packets, that is, LEN1x2 data, are sequentially requested from the DRAM memory each time, and the corresponding data is placed in the tightly coupled cache. After each request is completed, the request data address is incremented by LEN1x2. When the amount of requested data reaches LEN-LEN1x2, the request address is restored to the initial BA1. For the sake of simplifying the hardware complexity, a variant packet wave sorting algorithm is adopted for the first-level packet sorting.
[0137] After receiving the memory write-back enable signal, the tightly coupled cache writes back data to the DRAM according to the indication of the scheduling unit "ST1_AB type selection", based on the pre-set first-level packet data length (LEN1), the overall data length (LEN), and the base address (BA1) of the data in the DRAM memory.
[0138] When in the A-type sorting cycle, starting from the base address BA1, after receiving the write-back request, two first-level packets, that is, LEN1x2 data, are sequentially written back to the DRAM memory each time. After each write-back is completed, the write-back data address is incremented by LEN1x2. When the amount of write-back data reaches LEN, the request address is restored to the initial BA1.
[0139] When in the B-type sorting cycle, starting from the base address BA1+LEN1, after receiving the write-back request, two first-level packets, that is, LEN1x2 data, are sequentially written back to the DRAM memory each time. After each write-back is completed, the write-back data address is incremented by LEN1x2. When the amount of write-back data reaches LEN-LEN1x2, the request address is restored to the initial BA1.
[0140] The memory access controller is used to convert the data requests of the sorter into access requests for the DRAM.
[0141] The DRAM memory is a dynamic random access memory, which has the characteristic of large storage capacity. However, limited by the bandwidth of the data transmission interface, its data transmission rate is often much lower than that of the tightly coupled cache.
[0142] The scheduling unit coordinates and sorts each internal module to complete the sorting function.
[0143] The scheduling unit schedules the sorting process according to the lengths LEN2, LEN1, LEN of each level of grouped data, the number GC2 of secondary groups, and the number GC1 of primary groups.
[0144] In this embodiment, a variant sorting mode is adopted to simplify the design complexity for scheduling the sorting of primary and secondary groups. The basic sorting uses the basic sorting algorithm.
[0145] After the operation starts, for the tightly coupled memory, the scheduling unit first clears the primary sorting counter (CNT1), and then: Sends a memory load enable request to the tightly coupled memory to load two primary group data.
[0146] Determines the current A or B type cycle according to the value of the current CNT1. When CNT1%2 = 0, the "LD1_AB type selection" signal is marked as the A type cycle (the signal is 0); when CNT1%2 = 1, the "LD1_AB type selection" signal is marked as the B type cycle (the signal is 1).
[0147] Determines the current A or B type cycle according to the value of the current CNT1. When CNT1%2 = 0, the "ST1_AB type selection" signal is marked as the A type cycle (the signal is 0); when CNT1%2 = 1, the "ST1_AB type selection" signal is marked as the B type cycle (the signal is 1).
[0148] When two primary groups in the tightly coupled memory are processed, CNT1 is incremented by one, and then the operation of the next two primary groups is started. And the above operation is repeated until CNT1 is equal to GC1, then the sorting is completed.
[0149] For the register prefetch unit, the scheduling unit first clears the secondary sorting load counter (LCNT2), and then: Sends a preloading enable request to the register prefetch unit to load two secondary group data.
[0150] Determine the current A or B type cycle according to the value of the current LCNT2. When LCNT1 % 2 = 0, mark the "LD2_AB type selection" signal as the A type cycle (the signal is 0); when LCNT2 % 2 = 1, mark the "LD2_AB type selection" signal as the B type cycle (the signal is 1).
[0151] After two secondary packet processes in the register prefetch unit are completed, increment LCNT2 by one, and then start the operations of the next two primary packets. Repeat the above operations until LCNT2 is equal to GC2, then the sorting is completed.
[0152] For the register and the basic sorting unit: Before starting the sorting after every two secondary packets are merged, the scheduling unit first clears the basic sorting counter (BCNT).
[0153] After the register prefetch unit finishes loading, first issue a register load enable indication to load the data prefetched by the register prefetch unit into the register, and then disable the register load enable. At the same time, BCNT starts counting, and the basic sorting unit increments BCNT by one every time a basic sorting is completed.
[0154] Determine the current A or B type cycle according to the value of the current BCNT. When BCNT1 % 2 = 0, mark the "AB type selection" signal as the A type cycle (the signal is 0); when BCNT2 % 2 = 1, mark the "AB type selection" signal as the B type cycle (the signal is 1).
[0155] Repeat the above operations until BCNT is equal to LEN2x2 - 1, then the sorting is completed. At the same time, issue a result write-back enable indication to the result write-back unit.
[0156] Regarding the early termination of sorting: For the basic sorting: The scheduling unit determines whether the sorting can be terminated early by detecting the result of the "AND" operation of the "not occurred" indications of all basic sorting units (only when no exchanges occur in all basic sorting units, the result of the "AND" operation is 1, otherwise it is 0).
[0157] When no data exchanges occur in two consecutive sorting operations, the sorting is terminated early.
[0158] When no data exchanges occur in the current two sortings, it indicates that the sorting has been completed in the initial state, and mark the status of this sorting as the "basic sorting without exchange" status.
[0159] For the secondary sorting: The scheduling unit determines whether the sorting can be terminated early by detecting the "basic sorting without exchange" detected during the basic sorting.
[0160] When there is no data exchange in all basic sorts during three consecutive secondary sorting operations, the sorting terminates prematurely.
[0161] When there is no data exchange in all basic sorts during the current three secondary sorting operations, it indicates that the secondary sorting has been completed in the initial state, and the status of this sorting is marked as the "secondary sorting completed" status.
[0162] For the primary sorting: The scheduling unit determines whether the sorting can be terminated prematurely by detecting "no exchange in secondary sorting" detected during the basic sorting process.
[0163] When there is no data exchange in all secondary sorts during three consecutive primary sorting operations, the sorting terminates prematurely and the sequence sorting is completed.
[0164] Compared with the traditional sorting algorithm, the advantages of the data sorting method in the embodiments of the present application for sorting ultra-long sequences mainly include: The sorting algorithm has extremely high regularity, can effectively remove data dependencies between the data participating in the operation, and is convenient for using high-parallel vector machines and tensor machines to accelerate the operation. At the same time, it can relatively naturally integrate and solidify relevant operations in vector and tensor processors.
[0165] Adopting a hierarchical grouping structure can effectively utilize the storage structure of modern vector and tensor processors to optimize the access bandwidth of the sorting operation to the system DRAM, and greatly shorten the sorting time. Taking the sorting of 1024x1024 32-bit width data as an example: Using a traditional sorting algorithm such as bubble sort requires accessing 4x(1024x1024)x(1024x1024) / 2x2 = 4396GB of data from the DRAM.
[0166] Adopting a hierarchical fluctuation sorting method such as in the example: A sequence with a length of 1024x1024 is first divided into 64 first-level groups of 16x1024 data. During the operation, the first-level groups are stored in the tightly coupled cache, and the memory access generated by the subsequent first-level group sorting can be completed in the tightly coupled cache without generating additional data access to the DRAM.
[0167] According to the characteristics of the sorting algorithm, only 4x(1024x1024)x64x2 = 0.54GB of data needs to be accessed from the DRAM. Compared with the traditional sorting algorithm, the amount of data accessed has been greatly reduced.
[0168] An embodiment of the present application further provides a storage medium carrying one or more computer programs, and when the one or more computer programs are executed by a processor, the steps of the method described above are implemented.
[0169] An embodiment of the present application further provides a computer program product, including a computer program / instructions, characterized in that when the computer program / instructions are executed by a processor, the steps of the method described above are implemented.
[0170] It should be understood that in the embodiment of the present application, the processor may be a central processing unit (Central Processing Unit, abbreviated as CPU), and the processor may also be other general-purpose processors, digital signal processors (Digital Signal Processing, abbreviated as DSP), application-specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), field-programmable gate arrays (Field-Programmable Gate Array, abbreviated as FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0171] It should also be understood that the memory mentioned in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0172] It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, the memory (storage module) is integrated in the processor.
[0173] It should be noted that the memory described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0174] It should also be understood that the first, second, third, fourth, and various numerical numbers involved herein are only for the convenience of description and are not used to limit the scope of the present application.
[0175] It should be understood that the term "and / or" herein is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally indicates that the associated objects before and after are in an "or" relationship.
[0176] In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0177] In various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution is prior or subsequent. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0178] Those of ordinary skill in the art can realize that the various illustrative logical blocks (ILB) and steps described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0179] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the hardware accelerator embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical or other form.
[0180] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0181] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, may exist physically separately for each unit, or two or more units may be integrated into one unit.
[0182] In the above embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server, data center, etc. that contains one or more integrated available media. The available media may be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media (such as solid-state drives), etc.
[0183] As described above, the foregoing are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all such changes or substitutions should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A data sorting method, characterized in that, Including: Based on the arrangement order of the registers, respectively store multiple target data into their corresponding registers; Perform a first sorting operation on the target data in the registers, which includes: starting from the first register with the first serial number, sequentially compare the target data in two adjacent registers respectively, store the target data with a larger value into the register with a higher or lower order in the two registers participating in the comparison, and store the target data with a smaller value into the register with a lower or higher order in the two registers participating in the comparison; Perform a sub-operation in the intermediate sorting operation on the target data in the registers, where the sub-operation includes a second sorting operation and a third sorting operation. The second sorting operation includes: starting from the second register with the second serial number, sequentially compare the target data in two adjacent registers respectively, and compare the first register with the register with the last serial number, store the target data with a larger value into the register with a higher or lower order in the two registers participating in the comparison, and store the target data with a smaller value into the register with a lower or higher order in the two registers participating in the comparison; The third sorting operation includes: starting from the second register, sequentially compare the target data in two adjacent registers respectively, store the target data with a larger value into the register with a higher or lower order in the two registers participating in the comparison, and store the target data with a smaller value into the register with a lower or higher order in the two registers participating in the comparison; Repeatedly perform the first sorting operation and the second sorting operation alternately for multiple times, or repeatedly perform the first sorting operation and the third sorting operation alternately for multiple times until all the target data are sorted from large to small or from small to large in terms of value.
2. The data sorting method according to claim 1, wherein When there are multiple data units in the target data and they are not sorted, the method further includes: Based on the arrangement order of the registers, respectively store multiple data units into their corresponding registers; Take the data units in the registers as objects and perform the first sorting operation on them; Take the data units in the registers as objects and perform a sub-operation in the intermediate sorting operation on them; Repeatedly perform the first sorting operation and the second sorting operation alternately for multiple times, or repeatedly perform the first sorting operation and the third sorting operation alternately for multiple times until all the data units are sorted from large to small or from small to large in terms of value.
3. The data sorting method according to claim 1, wherein When there are multiple data units in the target data and they are already sorted, the method further includes: Group the data units based on the numerical size of the data units to form multiple data unit groups; Group the registers based on the arrangement order of the registers to form multiple register groups; Store the groups of data units into their corresponding register groups respectively based on the arrangement order of the register groups; Take the groups of data units in the register groups as objects and perform the first sorting operation on them; Take the groups of data units in the register groups as objects and perform a sub - operation of the intermediate sorting operation; Repeatedly alternate between the first sorting operation and the second sorting operation, or repeatedly alternate between the first sorting operation and the third sorting operation multiple times until all the groups of data units are sorted from largest to smallest or from smallest to largest in terms of numerical value.
4. The data sorting method according to claim 3, wherein Group the data units based on the numerical size of the data units to form multiple groups of data units, including: divide the data units into 2 groups of data units based on the numerical size of the groups of data units; Correspondingly, group the registers to form multiple register groups, including: construct 1 register group based on 4 registers.
5. The data sorting method according to claim 1, wherein The method further includes: Sort the target data based on a preset time interval; Send the sorted target data to the memory through the memory access controller.
6. A hardware accelerator, characterized in that, Connected to the memory, the hardware accelerator includes: Multiple sequentially arranged registers; A scheduling unit, which is connected to the registers and is used to store the target data into their corresponding registers; Multiple basic sorting units, which are respectively connected to two adjacent corresponding registers. The basic sorting unit is configured to: perform the first sorting operation on the target data in the registers under the control of the scheduling unit; perform a sub - operation of the intermediate sorting operation on the target data in the registers, and the sub - operation includes the second sorting operation and the third sorting operation; repeatedly alternate between the first sorting operation and the second sorting operation, or repeatedly alternate between the first sorting operation and the third sorting operation multiple times until all the target data are sorted from largest to smallest or from smallest to largest in terms of numerical value; where The first sorting operation includes: starting from the first register with the first serial number, sequentially compare the target data in two adjacent registers respectively, store the target data with a larger numerical value into the register with a higher or lower sorting order among the two registers participating in the comparison, and store the target data with a smaller numerical value into the register with a lower or higher sorting order among the two registers participating in the comparison; where The second sorting operation includes: starting from the second register with the second serial number, sequentially compare the target data in two adjacent registers respectively, and compare the first register with the register with the last serial number, store the target data with a larger numerical value into the register with a higher or lower sorting order among the two registers participating in the comparison, and store the target data with a smaller numerical value into the register with a lower or higher sorting order among the two registers participating in the comparison; The third sorting operation includes: starting from the second register, sequentially comparing the target data in two adjacent registers, storing the target data with a larger value in the register with a higher or lower order among the two registers participating in the comparison, and storing the target data with a smaller value in the register with a lower or higher order among the two registers participating in the comparison.
7. The hardware accelerator according to claim 6, wherein It further includes: A tightly coupled memory for storing the target data; A register prefetch unit, which is respectively connected to the scheduling unit, the register, and the tightly coupled memory. After receiving the instruction from the scheduling unit, the register prefetch unit obtains the target data from the tightly coupled memory and loads the target data into their respective corresponding registers according to the instruction.
8. The hardware accelerator according to claim 6, wherein A first data multiplexer is provided between the register and the basic sorting unit. The first data multiplexer is connected to the scheduling unit and, under the control of the scheduling unit, selectively determines one of its input data as the output data and sends it to the basic sorting unit; A second data multiplexer is provided between the register prefetch unit and the register. The second data multiplexer is connected to the scheduling unit and, under the control of the scheduling unit, selectively determines one of its input data as the output data and sends it to the register.
9. The hardware accelerator according to claim 7, characterized in that It further includes: A result write-back unit, which is respectively connected to the basic sorting unit, the scheduling unit, and the tightly coupled memory. Under the control of the scheduling unit, the result write-back unit stores the sorting result of the sorting operation on the target data.
10. The hardware accelerator according to claim 7, wherein It further includes: A memory access controller, which is respectively connected to the tightly coupled memory and the memory. It is used to obtain the target data from the memory according to the data call request of the tightly coupled memory, or send the sorting result of the sorting operation on the target data to the memory.
Citation Information
Patent Citations
Hardware circuit for accomplishing paralleling data ordering and method
CN101261576A
Quick sorting method and device based on FPGA (Field Programmable Gate Array)
CN114416020A
Sequencing circuit, sequencing method and electronic equipment
CN115543254A
Device, system and method for parallel data sorting
WO2019170961A1