Double-tuning sorting accelerator
By designing a hardware accelerator, using the comparison switching circuit and FIFO buffer, the efficient double-tuning sorting operation is achieved at the hardware level, solving the problem of intensive sorting operations in the existing technology, and improving the sorting speed and throughput.
Patent Information
- Application Number
- CN201980058652.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-31
- Filing Date
- 2019-07-11
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2039-07-11
AI Technical Summary
The prior art, when performing sorting operations, relies on software implementations of the central processing unit (CPU) or graphics processing unit (GPU), results in computation-intensive and reduces the ability of the CPU or GPU to perform other tasks.
A hardware accelerator is designed to perform double-tuning sorting operations through pipeline using multiple comparison switching circuits and associated first-in-first-out (FIFO) buffers. The hardware accelerator sorts N binary numbers in a total (N*log2N) clock cycles and further improves throughput by increasing parallelism.
A hardware solution that enables more efficient sorting speed and smaller circuit area, sorts arrays of data values, is more efficient than software implementation, and has three times the throughput.
Smart Images

Figure CN112654962B_ABST
Abstract
Description
Summary of the invention
[0001] According to at least one example of the present disclosure, a hardware accelerator for bitonic sorting includes a plurality of comparison exchange circuits and a first-in-first-out (FIFO) buffer associated with each of the comparison exchange circuits. The output of each FIFO buffer is a FIFO data value. The comparison exchange circuit is configured to: in a first operating mode, store a previous data value from a previous comparison exchange circuit or memory to its associated FIFO buffer, and pass the FIFO data value from its associated FIFO buffer to a subsequent comparison exchange circuit or memory; in a second operating mode, compare the previous data value with the FIFO data value, store the larger of the data values to its associated FIFO buffer, and pass the smaller of the data values to a subsequent comparison exchange circuit or memory; and in a third operating mode, compare the previous data value with the FIFO data value, store the smaller of the data values to its associated FIFO buffer, and pass the larger of the data values to a subsequent comparison exchange circuit or memory.
[0002] According to another example of the present disclosure, a hardware accelerator for bitone sorting includes four multiplexers (mux), each multiplexer includes an output and a first input configured to be coupled to a memory. The hardware accelerator also includes a four-input comparison exchange circuit having four inputs and four outputs, wherein the output of each multiplexer is coupled to one of the inputs of the four-input comparison exchange circuit. The hardware accelerator also includes four bitone sorting accelerators, which include a first bitone sorting accelerator, a second bitone sorting accelerator, a third bitone sorting accelerator, and a fourth bitone sorting accelerator. Each of the four bitone sorting accelerators has an input and an output, and each output of the four-input comparison exchange circuit is coupled to one of the bitone sorting accelerator inputs. The output of each bitone sorting accelerator is coupled to a second input of one of the multiplexers.
[0003] According to another example of the present disclosure, a method for bitonic sorting includes: for each of a plurality of comparison exchange circuits, receiving a control signal and operating in one of a first operation mode, a second operation mode, and a third operation mode in response to the control signal. In the first operation mode, the method also includes storing, by the comparison exchange circuit, a previous data value from a previous comparison exchange circuit or a memory to an associated FIFO buffer, wherein the output of the associated FIFO buffer is a FIFO data value; and passing the FIFO data value from the associated FIFO buffer to a subsequent comparison exchange circuit or memory. In the second operation mode, the method also includes: comparing, by the comparison exchange circuit, a previous data value with a FIFO data value; storing the larger of the data values to the associated FIFO buffer; and passing the smaller of the data values to a subsequent comparison exchange circuit or memory. In the third operation mode, the method also includes: comparing, by the comparison exchange circuit, a previous data value with a FIFO data value; storing the smaller of the data values to the associated FIFO buffer; and passing the larger of the data values to a subsequent comparison exchange circuit or memory. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] For a detailed description of various examples, reference will now be made to the accompanying drawings, in which:
[0005] Figure 1 A signal flow diagram showing a bitonic sequencing network according to various examples;
[0006] Figure 2 A block diagram showing a bitune sorting accelerator according to various examples;
[0007] Figure 3 A circuit schematic diagram showing a comparison exchange circuit according to various examples;
[0008] Figure 4 shows a signal flow graph of a bitonic sequencing network including a flow-through operation according to various examples;
[0009] Figure 5 shows data flow and timing diagrams for a bitune sorting accelerator according to various examples;
[0010] Figure 6 A block diagram illustrating a bitune sort accelerator with improved data parallelism according to various examples;
[0011] Fig. 7A and Figure 7B A signal flow diagram illustrating a bitonic sorting network with improved data parallelism according to various examples; and
[0012] Figure 8A flow chart of a method for bitonic sorting according to various examples is shown. DETAILED DESCRIPTION
[0013] Various algorithms such as those used for signal processing, radar tracking, image processing, etc. often use sorting operations. Sorting operations are usually implemented using software executed by a central processing unit (CPU) or a graphics processing unit (GPU), which requires a lot of calculations and thus reduces the ability of the CPU or GPU to perform other tasks. Hardware accelerators are used to perform certain mathematical operations (such as sorting) more efficiently than software executed on a general host processor such as a CPU or GPU. However, improvements in sorting speed and circuit area are expected.
[0014] According to the disclosed examples, a hardware accelerator for bitonic sorting (bitonic sorting accelerator) and a method for bitonic sorting provide a hardware solution for sorting an array of data values with improved sorting speed and reduced circuit area. Compared with software executed by a host processor, for example, the bitonic sorting accelerator of the present disclosure performs bitonic sorting more efficiently. In particular, the bitonic sorting accelerator of the present disclosure performs bitonic sorting of an array of data values in a pipelined manner using a structure similar to a radix-2 single delay feedback (R2SDF) architecture. The bitonic sorting accelerator sorts N binary numbers serially fed into the accelerator in a total of (N*log2N) clock cycles, which is equal to the theoretical upper limit of the sorting speed that can be achieved with any comparison-based sorting algorithm. In some examples, by increasing the parallelism of the hardware accelerator, the throughput of the bitonic sorting accelerator is further increased by three times.
[0015] A bitonic sequence is a sequence of elements (a0, a1, ..., a N-1 ). The first condition is that there exists an index i, 0≤i≤N-1, such that (a0,...,a i ) increases monotonically and (a i+1 ,...,a N-1 ) is monotonically decreasing. The second condition is that there is a cyclic shift of the exponent such that the first condition is satisfied. For example, {1,4,6,8,3,2} (which is monotonically increasing and then monotonically decreasing), {6,9,4,2,3,5} (a sequence for which the cyclic shift produces a monotonically increasing and then monotonically decreasing (starting from {2}) or a monotonically decreasing and then monotonically increasing (starting from {9})), and {9,8,3,2,4,6} (which is monotonically decreasing and then monotonically increasing) are bitonic sequences.
[0016] In an example of the present disclosure, a hardware accelerator sorts a bitonic sequence of size N by recursively applying a compare-exchange (CE) operation to the elements of the bitonic sequence. The hardware accelerator enables input data of size N to be sorted in a total of (N*log2N) clock cycles (which is equal to the theoretical upper limit of any comparison-based sorting algorithm) while reusing some parts of the R2SDF architecture. The CE operation compares two elements and then selectively swaps or exchanges the positions of the two elements based on which element has a larger value. For example, if the CE operation attempts to place the largest element in the second position, the CE operation compares a first value and a second value, and if the first value is greater than the second value, the two elements are swapped. However, if the second value is greater than the first value, no swap occurs.
[0017] Figure 1 An example signal flow graph of a bitonic sorting network 100 for sorting a data sequence of size N=8 with random input is shown. Typically, the input data is an N-element vector of data values. In the signal flow graph 100, the arrows indicate the two elements being compared (the elements at the "head" and "tail" of each arrow) and the direction in which the elements are swapped or permuted. Figure 1 In the example of , the smaller of the two elements being compared is positioned at the tail of the arrow after the comparison. The bitonic sorting network 100 first rearranges the unsorted data sequence (sequence A) into a bitonic sequence (sequence C), which occurs in the first log2N-1 stage, in this case, stages S1 and S2. Subsequently, in the final stage S3, the bitonic sorting network 100 rearranges the bitonic sequence (sequence C) into a sorted sequence (sequence D).
[0018] The input data or unsorted data sequence (sequence A) is considered as a combination of bitonic sequences of length 2. In stage S1, parallel CE operations are applied to adjacent bitonic sequences (pairs) in opposite directions, as marked by adjacent arrows facing opposite directions. The result of stage S1 is the conversion of the input data (sequence A) into a combination of bitonic sequences of length 4 (sequence B). In stage S2, as shown in the figure, similar parallel CE operations are applied to adjacent bitonic sequences in the opposite direction, and in the case where the input data size is greater than 8, the subsequent stages will continue in a similar manner until a bitonic sequence of length N is generated. In this example, the result of stage S2 is the generation of a bitonic sequence (sequence C) of length N=8. As shown in the figure, in the last stage, stage S3 in this example, the bitonic sequence (sequence C) is converted into a sorted sequence (sequence D).
[0019] Figure 2A bitonic sort accelerator 200 according to an example of the present disclosure is shown. The bitonic sort accelerator 200 receives input data (Di, which is an N-element vector of data values) from a memory 208 as an input to a dual-input multiplexer (mux) 202. As described above, the input data elements are received serially by the multiplexer 202 of the bitonic sort accelerator 200. The bitonic sort accelerator 200 also includes one or more pipeline compare exchange (CE) circuits 204. In Figure 2 In the example of , the CE circuit 204 includes a first CE circuit 204a and a last CE circuit 204c. For the CE circuit 204b, the CE circuit 204a is referred to as the previous CE circuit 204a, and the CE circuit 204c is referred to as the subsequent CE circuit 204c. Generally, each CE circuit 204 between the first CE circuit 204a and the last CE circuit 204c has one previous CE circuit 204 and one subsequent CE circuit 204.
[0020] For a bitonic sort accelerator 200 configured to sort input data of size N (generally assumed to be a power of 2), the bitonic sort accelerator 200 includes at least log2N CE circuits 204. In an example where N is not a power of 2, zero padding is used to increase the input data size to the next power of 2. Figure 2 In the example and in order to Figure 1 For example, assume that the size of the input data is N = 8. Therefore, Figure 2 In the example of , the bitonic sorting accelerator 200 includes three CE circuits 204a, 204b, 204c. The multiplexer 202 includes two inputs, one input coupled to the memory 208 as described above, and the other input coupled to the output data (Do) generated by the last CE circuit 204c. The output data (Do) is also provided to the memory 210, which is the same as the memory 208 in some examples and separate from the memory 208 in other examples.
[0021] Each CE circuit 204a, 204b, 204c is associated with a first-in-first-out (FIFO) buffer 206a, 206b, 206c, respectively. The FIFO buffers 206a, 206b, 206c act as delay elements and are implemented as memories or shift registers in some examples. For a bitonic sorting accelerator 200 having M CE circuits 204a, 204b, 204c, where the M CE circuits can be indexed using M', where M' ranges from 0 to log2N-1, the size of the FIFO buffers 206a, 206b, 206c is or in this case the sizes are 4, 2, and 1 respectively. The size of the FIFO buffer 206 associated with a particular CE circuit 204 specifies the "distance" of the comparisons performed by that particular CE circuit 204. Back Figure 1 For example, in stage S1, all comparisons have adjacent values with a distance of 1; similarly, in stage S2, comparisons have values with a distance of 2 and a distance of 1; finally, in stage S3, comparisons have values with a distance of 4, then 2, and then 1. Each CE circuit 204a, 204b, 204c also receives a control signal C2, C1, C0, respectively, which will be further described in detail below.
[0022] Figure 3 CE circuit 204 is shown in more detail. CE circuit 204 includes a first input 302 coupled to the output of its associated FIFO buffer 206. For ease of reference, the output data of each FIFO buffer 206 may be referred to as a FIFO data value. CE circuit 204 also includes a first output 306 coupled to the input of its associated FIFO buffer 206. CE circuit 204 also includes a second input 304 and a second output 308. Second input 304 is coupled to a previous CE circuit 204 (e.g., Figure 2 ) or a memory (e.g., coupled to memory 208 via multiplexer 202, such as Figure 2 The second output 308 is coupled to the subsequent CE circuit 204 (eg, Figure 2 ) or a second input or memory (e.g., as shown) Figure 2 Memory 210 shown).
[0023] The CE circuit 204 also includes a comparator 310 that receives the first input 302 and the second input 304 as inputs and generates an output based on a comparison of the first input 302 and the second input 304. Figure 3 In the example of , the output of the comparator 310 is asserted (e.g., to “1”) when the first input 302 is greater than the second input 304, and is de-asserted (e.g., to “0”) when the first input 302 is less than the second input 304.
[0024] CE circuit 204 receives the n [0] and C n The output of comparator 310 and the least significant bit C n [0] is provided as input to an exclusive OR (XOR) gate 312. The output of the exclusive OR gate 312 and the most significant bit C n[1] is provided as an input to an AND gate 314. The output of the AND gate 314 is a control of a first output multiplexer 316 and a second output multiplexer 318, the outputs of which include a first output 306 and a second output 308, respectively. In response to the output of the AND gate 314 being asserted, the first output multiplexer 316 passes the first input 302 therethrough as the first output 306, and the second output multiplexer 318 passes the second input 304 therethrough as the second output 308. In response to the output of the AND gate 314 being deasserted, the first output multiplexer 316 passes the second input 304 therethrough as the first output 306, and the second output multiplexer 318 passes the first input 302 therethrough as the second output 308.
[0025] Due to the above logic of CE circuit 204, the compare exchange operation is controlled by the control signal C n Specify as follows:
[0026] 0 (or 1): In the first operation mode, the compare-exchange operation will bypass the CE circuit 204, which corresponds to the following about Figure 4 The flow-through operation described in more detail; data from the previous CE circuit (second input 304) is stored in the FIFO buffer 206 (as the first output 306), and the earliest data from the FIFO buffer 206 (first input 302) is passed to the next CE circuit (as the second output 308).
[0027] 2: In the second operation mode, the compare-and-exchange operation is: in the case of the first CE circuit 204a, the data from the previous CE circuit or memory 208 (the second input 304) is compared with the FIFO data value (which is the data from the FIFO buffer 206 (the first input 302)); in the case of the last CE circuit 204c, the larger data value is stored in the FIFO buffer 206 (as the first output 306), and the smaller data value is passed to the next CE circuit or memory 210 (as the second output 308).
[0028] 3: In the third operation mode, the compare-and-exchange operation is: in the case of the first CE circuit 204a, the data from the previous CE circuit or memory 208 (the second input 304) is compared with the FIFO data value (which is the earliest data (the first input 302) from the FIFO buffer 206); in the case of the last CE circuit 204c, the smaller data value is stored in the FIFO buffer 206 (as the first output 306), and the larger data value is passed to the next CE circuit or memory 210 (as the second output 308).
[0029] As will be described further below, the difference in direction between control signal "2" and control signal "3" allows for Figure 1 The directionality of the arrow.
[0030] Figure 4 Another example signal flow diagram for bitonic sequencing 400 including a flow-through operation (eg, corresponding to control signal 0 as described above) is shown. In particular, Figure 1 The exemplary signal flow graph 100 of is shown to include flow-through operations 402, which are shown as data elements at the ends connected by dashed lines (not arrows). The flow-through operations are implemented, for example, to maintain a steady flow of data across the pipeline stages of the bitonic sort engine 200. For example, in stage S1, the comparison between elements with a distance of 4 and 2 is represented as a flow-through operation. Similarly, in stage S2, the comparison between elements with a distance of 4 is represented as a flow-through operation. In stage S3, since the comparison with a distance of 4 is required (for performing the final bitonic sort operation, as described above with respect to Figure 1 As described above), there is no flow-through operation for the special example of N=8.
[0031] Figure 5 An example data flow and timing diagram 500 of input data (Di) and output data (Do) of a bitune sorting accelerator 200 for N=8 is shown, where the input pattern corresponds to Figure 1 and Figure 4 Typically, the operation of the bitone sort accelerator 200 requires N*log2N clock cycles (24 clock cycles in this case) to complete from the time the last input data value (“1” in this example) is fed to the bitone sort accelerator 200 .
[0032] Return to reference Figure 2 The feedback connection from the output data Do to the input of the multiplexer 202 allows the above description of the bitune sorting accelerator 200 to be implemented iteratively (e.g., log2N times). Figure 1 and Figure 4 In the first iteration corresponding to stage S1, CE circuit 204a (corresponding to distance 4) and CE circuit 204b (corresponding to distance 2) operate in flow-through mode because stage S1 only performs compare-and-exchange operations on adjacent values with a distance of 1.
[0033] The compare-and-exchange operation for CE circuit 204c (corresponding to a distance of 1) begins at 0 in the seventh clock cycle to cause the first value (in this example, "8") to flow to the associated FIFO buffer 206c. At this point in time, ordered from oldest to newest, FIFO buffer 206a contains the values 5, 4, 3, 2; FIFO buffer 206b contains the values 7, 6; and FIFO buffer 206c contains the value 8.
[0034] In the eighth clock cycle, the compare-exchange operation of the CE circuit 204c is 2, which causes the CE circuit 204c to compare the data from the previous CE circuit 204b (value 7, as the earliest data in the FIFO buffer 206b and subject to the flow-through operation) with the earliest data from the FIFO buffer 206c (value 8). The larger data value 8 is stored back to the FIFO buffer 206c, while the smaller data value 7 is passed as the output data Do, which is reflected as the first element of Do (sequence B) in the timing diagram 500. In addition, the control signal to the multiplexer 202 is changed at this point in time so that the output data Do is used as the input data to the CE circuit 204a to start the second iteration to implement the next stage, in this case, stage S2.
[0035] In the ninth clock cycle, the compare-exchange operation of the CE circuit 204c is again 0 (flow-through), which causes the CE circuit 204c to pass the data value 8 from its associated FIFO buffer 206c as the output data Do, which is reflected as the second element of Do (sequence B) in the timing diagram 500. In the tenth clock cycle, the compare-exchange operation of the CE circuit 204c is 3, which causes the CE circuit 204c to compare the comparison data (value 5, as the earliest data in the FIFO buffer 206b and subject to the flow-through operation) from the previous CE circuit 204b with the earliest data (value 6) from the FIFO buffer 206c. The smaller data value 5 is stored in the FIFO buffer 206c, and the larger data value 6 is passed as the output data Do, which is reflected as the third element of Do (sequence B) in the timing diagram 500. The above process is repeated to compare data values 4 and 3 (using compare-exchange operation 2) and data values 2 and 1 (using compare-exchange operation 3) to complete the stage S1 compare-exchange operation for adjacent values with a distance of 1.
[0036] In addition to modifying the control signal C n Phase S2 is implemented in a manner similar to that described above with respect to phase S1, except that the directional change of the compare-and-exchange operation is taken into account. The remainder of timing diagram 500 reflects control signal C corresponding to the result of phase S1 (sequence B), the result of phase S2 (sequence C), and the result of phase S3 (sequence D). n And output data Do.
[0037] In addition, the control signal C n A pattern is followed that is generated, for example, using counter bits from a modulo-N binary counter (counting from 0 to N-1) and a modulo-log2N binary counter (counting from 0 to log2N-1) associated with each CE circuit 204a, 204b, 204c. The modulo-log2N binary counter is incremented at each iteration, and the modulo-N binary counter is incremented at each clock cycle. Each of the CE circuits 204a, 204b, 204c is asserted (e.g., control signal C is asserted) when the modulo-log2N binary counter reaches a particular value. n =2 or C n =3). For example, for N=8, when the modulo log2N counter is equal to 2, C2 is valid; when the modulo log2N counter is greater than or equal to 1, C1 is valid; and when the modulo log2N counter is greater than or equal to 0, C0 is valid. Using the bits from the modulo-N counter, C is determined for each CE circuit 204a, 204b, 204c based on combinational logic. n In other examples, the control signal C is accessed from a control signal buffer in memory. n .
[0038] Figure 2 The bitone sort accelerator 200 shown in and described above is serial in nature, because the bitone sort accelerator 200 receives serial input data (Di) and generates output data (Do) serially after a fixed latency. However, in some examples, the computer system on which the bitone sort accelerator is implemented includes a processor, a bus structure, and a memory access (e.g., direct memory access (DMA)) with a wider bandwidth, and is therefore capable of handling a higher throughput. In such a computer system, the overall system performance is reduced by a hardware accelerator that consumes data and generates data relatively slowly, such as the serial input and serial output of the bitone sort accelerator 200.
[0039] Figure 6 A bitonic sort accelerator 600 is shown with a high level of data parallelism, which reduces the number of clock cycles required to perform a sort on an N-element vector of data values. The bitonic sort accelerator 600 receives input data from four parallel streams (denoted as x1-x4) from the memory 208, where each stream is an input to a dual-input multiplexer (mux) 602. As described above with respect to Figure 2As described above, the input data elements are received serially by the multiplexer 602 of the bitune sort accelerator 600, but with a 4-fold parallelism. The bitune sort accelerator 600 also includes a four-input CE circuit 604, which includes four CE circuits 204a-204d, which are connected to the bitune sort accelerator 600. Figure 2 and Figure 3 The CE circuits are the same as those shown in and described above.
[0040] The first CE circuit 204a includes a first input coupled to the output of the first multiplexer 602a and a second input coupled to the output of the second multiplexer 602b. The second CE circuit 204b includes a first input coupled to the output of the third multiplexer 602c and a second input coupled to the output of the fourth multiplexer 602d. The third CE circuit 204c includes a first input coupled to the first output of the first CE circuit 204a and a second input coupled to the first output of the second CE circuit 204b. The fourth CE circuit 204d includes a first input coupled to the second output of the first CE circuit 204a and a second input coupled to the second output of the second CE circuit 204b. As described above, the CE circuits 204a-204d are configured to: operate in a flow-through mode, wherein the first output and the second output correspond to the second input and the first input, respectively; operate in a comparison mode, wherein the larger data value of the input is the first output and the smaller data value of the input is the second output; and operate in a comparison mode, wherein the smaller data value of the input is the first output and the larger data value of the input is the second output.
[0041] The first output and the second output of the third CE circuit 204c and the fourth CE circuit 204d are respectively coupled to the above Figure 2 The output (y1) of the first dual-tone sorting accelerator 200a is coupled to the input of the multiplexer 602d. The output (y2) of the second dual-tone sorting accelerator 200b is coupled to the input of the multiplexer 602b. The output (y3) of the third dual-tone sorting accelerator 200c is coupled to the input of the multiplexer 602c. The output (y4) of the fourth dual-tone sorting accelerator 200d is coupled to the multiplexer 602a.
[0042] Fig. 7A and Figure 7B An example signal flow graph of bitonic sequencing 700 is shown, which does not include flow-through operations for simplicity. Figure 2 The example of the bitone sorting accelerator 200 is directed to an 8-point bitone sorting accelerator, but the disclosure can be extended to other numbers of points by adding additional CE circuits and associated FIFO buffers as described above. Figure 2As an example, as described above, the function of the bitone sorting accelerator 600 is described as a 32-point bitone sorting accelerator using four 8-point bitone sorting accelerators 200. In the signal flow graph 700, rows 701, 703, 705, and 707 correspond to the functions of the 8-point bitone sorting accelerators 200a, 200b, 200c, and 200d, respectively.
[0043] In the first stage 702, the CE circuits 204a-204d of the four-input CE circuit 604 operate in the flow-through mode, so that the x1 input data is provided to the 8-point bitonic sort accelerator 200d, the x2 input data is provided to the 8-point bitonic sort accelerator 200b, the x3 input data is provided to the 8-point bitonic sort accelerator 200c, and the x4 input data is provided to the 8-point bitonic sort accelerator 200a. In the first stage 702, the 8-point bitonic sort accelerators 200a-200d implement the flow-through operation for comparison between elements with distances of 4 and 2, while comparing elements with distances of 1 as described above. In this case, only the final CE circuits of the 8-point bitonic sort accelerators 200a-200d are not operated in the flow-through mode.
[0044] In the second stage 704 and the third stage 706, the CE circuits 204a-204d of the four-input CE circuit 604 also operate in the flow-through mode, but after reading 8 elements (in this example) from the memory 208, the multiplexers 602a-602d are configured to provide the output of the 8-point bitone sort accelerator 200a-200d as input to the four-input CE circuit 604. In the second stage 704, the 8-point bitone sort accelerator 200a-200d implements the flow-through operation for comparison between elements with a distance of 4, while comparing elements with distances of 2 and 1 as described above. In this case, only the last two CE circuits of the 8-point bitone sort accelerator 200a-200d do not operate in the flow-through mode. In the third stage 706, the 8-point bitone sort accelerator 200a-200d does not perform the flow-through operation, and compares elements with distances of 4, 2, and 1 as described above.
[0045] In the fourth stage 708, CE circuits 204c and 204d operate in the comparison mode (corresponding to 708a) to perform comparisons between elements with a distance of 8. The 8-point bitune sort accelerators 200a-200d do not implement flow-through operations, and compare elements with distances of 4, 2, and 1 as described above (corresponding to 708b). CE circuits 204a and 204b operate in the flow-through mode.
[0046] Finally, in the fifth stage 710, the CE circuits 204a-204d are all operated in the comparison mode (corresponding to 710a) to perform comparisons between elements with distances of 16 and 8. The 8-point bitonic sort accelerators 200a-200d do not implement the flow-through operation, and compare elements with distances of 4, 2, and 1 as described above (corresponding to 710b). Neither the CE circuits 204a-204d nor the CE circuits in the 8-point bitonic sort accelerators 200a-200d implement the flow-through operation. In this example, the fourth cycle and the fifth cycle are exemplary. Typically, the four-input CE circuit 604 implements the flow-through operation until the last two iterations or the last two stages.
[0047] Relative to Figure 2 The bitonic sort accelerator 600 improves throughput and latency. For example, for a data array of length N, the number of iterations remains log2N. However, due to the four-input CE circuit 604 and the N / 4-point bitonic sort accelerator (e.g., Figure 6 The parallelism introduced by the 8-point bitonic sort accelerator 200a-200d in the example of FIG. 600 reduces the clock cycles required for each iteration to one-fourth. Therefore, the latency of the bitonic sort accelerator 600 is ((N*log2N) / 4) clock cycles, and has an effective throughput of ((log2N) / 4) clock cycles per sample.
[0048] Figure 8 800 according to an example of the present disclosure. The method 800 begins in block 802 by receiving a control signal, such as the one described above with respect to Figure 3 The C n In block 804, method 800 includes determining an operation mode of the compare-swap circuit indicated by the control signal, in one example, if the value of the control signal is 0 or 1, then a first operation mode; if the value of the control signal is 2, then a second operation mode; and if the value of the control signal is 0, then a third operation mode.
[0049] If the control signal causes the comparison exchange circuit to operate in the first operating mode, the method 800 proceeds to block 806, where the previous data value from the previous comparison exchange circuit or memory is stored in the associated FIFO buffer. The output of the associated FIFO buffer is referred to as the FIFO data value. The method 800 then continues to block 808, where the FIFO data value is passed from the associated FIFO buffer to a subsequent comparison exchange circuit or memory.
[0050] If the control signal causes the comparison exchange circuit to operate in the second operation mode, the method 800 proceeds to block 810, where the previous data value is compared with the FIFO data value. The method 800 then continues in block 812, where the larger of the data values is stored in the associated FIFO buffer, and the smaller of the data values is passed to a subsequent comparison exchange circuit or memory in block 814.
[0051] If the control signal causes the comparison exchange circuit to operate in the third operation mode, the method 800 proceeds to block 816, where the previous data value is compared with the FIFO data value. The method 800 then continues in block 818, where the smaller of the data values is stored in the associated FIFO buffer, and the larger of the data values is passed to a subsequent comparison exchange circuit or memory in block 820.
[0052] As mentioned above, for example, Figure 5 , a control signal is provided so that during a first iteration or a first group of iterations, the N-element vector of input data is arranged into a bitonic sequence by the plurality of compare-swap circuits. Furthermore, in a final iteration, the bitonic sequence is arranged into a completely ordered array by the plurality of compare-swap circuits. The control signal may be provided by a control signal buffer in a memory, or provided using a counter bit as described above.
[0053] In the foregoing discussion and claims, reference is made to a bitonic sorting accelerator comprising various elements, sections, and stages. It should be understood that these elements, sections, and stages correspond to hardware circuitry implemented, for example, on an integrated circuit (IC), as appropriate. In fact, in at least one example, the entire bitonic sorting accelerator is implemented on an IC.
[0054] In the foregoing discussion and claims, the terms "include" and "comprising" are used in an open-ended manner and should therefore be interpreted as meaning "including but not limited to...". Similarly, the terms "coupled" or "coupled" are intended to mean an indirect or direct connection. Thus, if a first device is coupled to a second device, the connection may be by a direct connection or by an indirect connection via other devices and connections. Similarly, a device coupled between a first component or location and a second component or location may be by a direct connection or by an indirect connection via other devices and connections. An element or feature that is "configured to" perform a task or function may be configured (e.g., programmed or structurally designed) by a manufacturer at the time of manufacture to perform a function, and / or may be configured (or reconfigurable) by a user after manufacture to perform a function and / or other additional or alternative functions. Configuration may be performed by firmware and / or software programming of the device, by the construction and / or layout of hardware components, and by the interconnection of the device, or a combination thereof. In addition, in the foregoing discussion, the use of the phrase "ground" or similar words is intended to include frame ground, earth ground, floating ground, virtual ground, digital ground, common ground, and / or any other form of ground connection that may be applicable or suitable for the teachings of the present disclosure. Unless otherwise indicated, "about," "approximately," or "substantially" preceding a numerical value means + / - 10% of the stated value.
[0055] The above discussion is intended to illustrate the principles and various embodiments of the present disclosure. Once the above disclosure is fully understood, many changes and modifications will become apparent to those skilled in the art. It is intended that the appended claims be interpreted as covering all such changes and modifications.
Claims
1. A hardware accelerator for bitonic sorting, the hardware accelerator comprising: A comparison exchange circuit, comprising a first input terminal, a second input terminal, a control terminal, a first output terminal, and a second output terminal; as well as A first-in-first-out FIFO buffer comprising: an input terminal coupled to the first output terminal of the comparison exchange circuit; and an output terminal coupled to the first input terminal of the comparison exchange circuit; and The comparison and exchange circuit is configured as follows: receiving a first data value at the first input; receiving a second data value at the second input; receiving a control signal at the control terminal; In response to determining that the control signal has a first value: outputting the first data value at the second output terminal; and outputting the second data value at the first output terminal; In response to determining that the control signal has a second value: outputting the larger of the first data value and the second data value at the first output terminal; and outputting the smaller of the first data value and the second data value at the second output terminal; and In response to determining that the control signal has a third value: outputting the smaller of the first data value and the second data value at the first output terminal; and The larger of the first data value and the second data value is output at the second output terminal.
2. The hardware accelerator according to claim 1, wherein the comparison and exchange circuit is a first comparison and exchange circuit, and the hardware accelerator comprises: a second comparison and exchange circuit coupled to the second input terminal; as well as A third comparison and exchange circuit is coupled to the second output terminal.
3. The hardware accelerator according to claim 1, wherein the comparison and exchange circuit is a first comparison and exchange circuit, and the hardware accelerator further comprises: A second comparison and exchange circuit includes a first input terminal, a second input terminal, a control input terminal, a first output terminal, and a second output terminal; Memory; as well as The multiplexer mux includes: a first input terminal coupled to the second output terminal of the second comparison exchange circuit; a second input terminal coupled to the memory; and An output terminal is coupled to the first input terminal of the first comparison exchange circuit. 4 . The hardware accelerator of claim 3 , wherein the second output terminal of the second comparison and exchange circuit is coupled to the memory. 5 . The hardware accelerator of claim 3 , wherein on a first iteration, the multiplexer is configured to serially receive an N-element vector of data values from the memory and provide to the second input of the first compare-swap circuit. 6 . The hardware accelerator of claim 5 , wherein on a subsequent iteration, the multiplexer is configured to couple the second output of a last compare-swap circuit to the second input of the first compare-swap circuit.
7. The hardware accelerator of claim 6 , further comprising a control signal buffer that holds a control signal that, when provided to the first compare-swap circuit and the second compare-swap circuit, causes the first compare-swap circuit and the second compare-swap circuit to arrange the N-element vector into a bitonic sequence during a first iteration or a first group of iterations and to arrange the N-element vector into a fully sorted array during a final iteration.
8. The hardware accelerator according to claim 1, wherein the comparison and exchange circuit is a first comparison and exchange circuit, and the hardware accelerator comprises: a memory coupled to the second input terminal; and A second comparison and exchange circuit is coupled to the second output terminal.
9. The hardware accelerator according to claim 1, wherein the comparison and exchange circuit is a first comparison and exchange circuit, and the hardware accelerator comprises: a second comparison and exchange circuit coupled to the second input terminal; and A memory is coupled to the second output terminal.
10. A hardware accelerator for bitonic sorting, the hardware accelerator comprising: four multiplexers mux, each comprising an output terminal, a first input terminal adapted to be coupled to the memory, and a second input terminal; a comparison switching circuit comprising four input terminals and four output terminals, wherein the output terminal of each multiplexer is coupled to one of the input terminals of the comparison switching circuit; as well as Four sorting accelerators, including a first sorting accelerator, a second sorting accelerator, a third sorting accelerator and a fourth sorting accelerator, each of the four sorting accelerators includes an input terminal and an output terminal, wherein each output terminal of the comparison and exchange circuit is coupled to one of the input terminals of the sorting accelerator, The output of each sort accelerator is coupled to the second input of one of the multiplexers.
11. The hardware accelerator according to claim 10, wherein the comparison and exchange circuit further comprises: The first double-input comparison exchange circuit and the second double-input comparison exchange circuit each include a first input terminal and a second input terminal and a first output terminal and a second output terminal, wherein: The first input terminal of the first dual-input comparison exchange circuit is coupled to the output terminal of a first multiplexer among the multiplexers; The second input terminal of the first dual-input comparison exchange circuit is coupled to the output terminal of the second multiplexer; The first input terminal of the second dual-input comparison exchange circuit is coupled to the output terminal of a third multiplexer in the multiplexers; and The second input terminal of the second dual-input comparison exchange circuit is coupled to the output terminal of a fourth multiplexer in the multiplexers; and The third dual-input comparison exchange circuit and the fourth dual-input comparison exchange circuit each include a first input terminal and a second input terminal and a first output terminal and a second output terminal, wherein: The first input terminal of the third dual-input comparison exchange circuit is coupled to the first output terminal of the first dual-input comparison exchange circuit; The second input terminal of the third dual-input comparison exchange circuit is coupled to the first output terminal of the second dual-input comparison exchange circuit; The first input terminal of the fourth dual-input comparison exchange circuit is coupled to the second output terminal of the first dual-input comparison exchange circuit; and The second input terminal of the fourth dual-input comparison exchange circuit is coupled to the second output terminal of the second dual-input comparison exchange circuit; in: The first output terminal of the third dual-input comparison exchange circuit is coupled to the input terminal of the first sort accelerator; The second output terminal of the third dual-input comparison exchange circuit is coupled to the input terminal of the second sort accelerator; The first output terminal of the fourth dual-input comparison exchange circuit is coupled to the input terminal of the third sort accelerator; and The second output terminal of the fourth dual-input comparison exchange circuit is coupled to the input terminal of the fourth sort accelerator.
12. The hardware accelerator according to claim 10, wherein: The output of the first sort accelerator is coupled to the second input of a fourth multiplexer in the multiplexers; The output of the second sort accelerator is coupled to the second input of a second multiplexer in the multiplexers; The output of the third sort accelerator is coupled to the second input of a third multiplexer in the multiplexers; and The output of the fourth sort accelerator is coupled to the second input of a first one of the multiplexers.
13. The hardware accelerator according to claim 10, wherein each sort accelerator further comprises: Dual input comparison exchange circuit; as well as a first-in-first-out FIFO buffer associated with each of said dual-input compare-and-swap circuits, wherein an output of each FIFO buffer is a FIFO data value; The dual-input comparison exchange circuit is configured as follows: In a first mode of operation, storing a previous data value from a previous dual-input compare-swap circuit or the compare-swap circuit to its associated FIFO buffer and transferring the FIFO data value from its associated FIFO buffer to a subsequent dual-input compare-swap circuit, one of the multiplexers, or the memory; in a second mode of operation, comparing the previous data value with the FIFO data value, storing the larger of the data values to its associated FIFO buffer, and passing the smaller of the data values to the subsequent dual-input compare exchange circuit, one of the multiplexers, or the memory; as well as In a third mode of operation, the previous data value is compared with the FIFO data value, the smaller of the data values is stored to its associated FIFO buffer, and the larger of the data values is passed to the subsequent dual-input compare exchange circuit, one of the multiplexers, or the memory.
14. The hardware accelerator of claim 13, wherein each of the dual-input comparison exchange circuits comprises: a first input coupled to an output of its associated FIFO buffer; a first output coupled to an input of its associated FIFO buffer; a second input terminal coupled to the second output terminal of the previous dual-input comparison exchange circuit or the output terminal of the comparison exchange circuit; as well as The circuit is coupled to the second input terminal of a subsequent dual-input comparison exchange circuit, one of the multiplexers, or the second output terminal of the memory.
15. The hardware accelerator according to claim 13, wherein: Each of the dual-input comparison exchange circuits is configured to receive a control signal; as well as The received control signal causes the dual-input compare-and-swap circuit to operate in one of the first operation mode, the second operation mode, and the third operation mode.
16. The hardware accelerator of claim 13, wherein on a first iteration, each of the multiplexers is configured to serially receive and provide an N / 4 element vector of data values from the memory to an input of the compare-and-swap circuit. 17 . The hardware accelerator of claim 16 , wherein on subsequent iterations, each of the multiplexers is configured to couple one of the outputs of the sort accelerator to one of the inputs of the compare-swap circuit.
18. The hardware accelerator of claim 16, further comprising a control signal buffer that holds a control signal that, when provided to the compare-and-swap circuit and the dual-input compare-and-swap circuit of the sort accelerator, causes the hardware accelerator to arrange the four N / 4-element vectors into a bitonic sequence during a first iteration or a first group of iterations and to arrange the four N / 4-element vectors into a fully sorted array during a final iteration.
19. A method for bitonic sorting, the method comprising: receiving a control signal at a control terminal through a comparison exchange circuit; receiving a first data value at a first input terminal via the compare-and-swap circuit; receiving a second data value at a second input terminal via the compare-and-swap circuit; In response to determining that the control signal has a first value: outputting the first data value at a first output terminal through the compare-and-swap circuit; and outputting the second data value at a second output terminal through the comparison exchange circuit; In response to determining that the control signal has a second value: outputting the larger of the first data value and the second data value at the first output terminal through the comparison exchange circuit; and outputting the smaller of the first data value and the second data value at the second output terminal through the comparison exchange circuit; and In response to determining that the control signal has a third value: outputting the smaller of the first data value and the second data value at the first output terminal through the comparison exchange circuit; and The larger one of the first data value and the second data value is output at the second output terminal by the comparison exchange circuit.
20. The method of claim 19, further comprising providing a control signal to a plurality of compare-and-swap circuits, the plurality of compare-and-swap circuits including the compare-and-swap circuit, the control signal instructing the plurality of compare-and-swap circuits to arrange the N-element vector into a bitonic sequence during a first iteration or a first group of iterations and to arrange the N-element vector into a fully sorted array during a final iteration.
Citation Information
Patent Citations
Memory ordering in acceleration hardware
CN108268386A
MEMS-based Switching System
US20160196940A1