Select the Ith largest or Pth smallest number from a set of N M-digit numbers
By selecting the number that is the largest or smallest from the set of n m-digit numbers in the hardware logic through iterative sum and comparison, the problems of propagation delay and large hardware layout area in the prior art are solved, and a more efficient hardware logic design is achieved.
Patent Information
- Application Number
- CN201911039015.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-28
- Filing Date
- 2019-10-29
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2039-10-29
AI Technical Summary
Existing hardware logic has propagation delay issues when sorting multiple input numbers, resulting in performance degradation, and adding buffers to resolve delay issues increases the area of the hardware layout.
By iteratively summing the bits of each number in the set of m-digit numbers, generating a sum result, and comparing it with the threshold, setting the bits of the selected number based on the comparison results, and selecting the bits of the next bit position, it realizes selecting the number that is the largest or p-th or smallest from the set of n m-digit numbers.
The area of hardware logic is reduced, performance is improved, propagation delay problem is avoided, and the number that is the largest or smallest is achieved without sorting the entire set of numbers.
Smart Images

Figure CN111124357B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a method for selecting the i-th largest or p-th smallest number from a set of n m-digit numbers in hardware logic. Background Art
[0002] There are many situations where hardware is required to sort two or more input binary numbers, that is, to place them in order of magnitude. Such a sorter is usually composed of a Figure 1 The shown components are composed of multiple identical logic blocks. Figure 1 A schematic diagram of an example hardware arrangement 100 is shown, which is used to sort four inputs x1, x2, x3, x4 in order of size, that is, such that output1≥output2≥output3≥output4. It can be seen that this sorter 100 includes five identical logic blocks 102, each of which outputs the largest and smallest (i.e., maximum (max) and minimum (min)) values of two inputs (which can be represented as a and b).
[0003] Each logic block 102 receives two n-bit integer inputs (a, b) and includes a comparator that returns a Boolean indicating whether a>b. The comparator output, which may be referred to as a 'select' signal, is then used to control a plurality of n-bit wide multiplexers, each of which selects between n bits from a or n bits from b. If logic block 102 outputs a maximum and minimum value (from a and b, as Figure 1 ), then the select signal is used to control the multiplexing of 2n bits (e.g., in the form of 2n 1-bit wide multiplexers or two n-bit wide multiplexers). Alternatively, if the logic block has only one output (which is the maximum or minimum of a and b), then the select signal is used to control the multiplexing of n bits (e.g., in the form of n 1-bit wide multiplexers or one n-bit wide multiplexer).
[0004] In the arrangement described above, the select signal is used to power multiple logic elements (e.g., logic gates) within the logic block 102, which results in a large propagation delay. This delay effect is caused by a single gate output wire having to charge transistors in a large number of gates before these subsequent gates can propagate their outputs, and may be referred to as 'fanout'. Although this delay may be acceptable when only two input numbers are being sequenced, where the logic blocks 102 are cascaded (e.g., as in Figure 1 The sorter 100 shown may be used in larger sorters with more than four inputs), but the resulting delay of the sorting circuitry will increase, which may seriously affect performance (eg, it may cause the sorting process to take more than one clock cycle).
[0005] A solution to this delay is to include a large number of buffers (eg, at least n buffers that may be arranged in a tree structure), each driven by a select signal; however, this results in a significantly larger (eg, in terms of logic area) hardware arrangement.
[0006] The embodiments described below are provided by way of example only and are not limiting embodiments that address any or all of the shortcomings of known hardware and methods for sorting numbers and / or selecting a number from a set of numbers based on its ordered position in the set. Summary of the invention
[0007] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0008] A method for selecting the i-th largest or p-th smallest number from a set of n m-bit numbers in hardware logic is described. The method is performed iteratively, and in the r-th iteration, where r is between 1 and m, the method includes: summing the (mr)-th bit from each of the m-bit numbers to generate a summed result, wherein the (m-1)-th bit is the most significant bit; and comparing the summed result with a threshold. Based on the result of the comparison, the (mr)-th bit of the selected number is determined and output, and in addition, the (mr-1)-th bit of each of the m-bit numbers is selectively updated based on the result of the comparison and the value of the (mr)-th bit in the m-bit number. In the first iteration, r=1, and the most significant bit from each of the m-bit numbers is summed and the most significant bit of the selected number is output. Each subsequent iteration sums the most significant bits of the numbers that are still to be summed and outputs the next bit of the selected number, which are bits that occupy consecutive bit positions in their corresponding numbers (r=2 bits in position m-2, r=3 bits in position m-3, etc.). Instances of this method also output the index of the i-th largest or p-th smallest number.
[0009] A first aspect provides a method for selecting a number from a set of n m-bit numbers in hardware logic, wherein the selected number is the i-th largest or p-th smallest number from the set of n m-bit numbers, wherein i, p, m and n are integers, the method comprising multiple iterations, and each of the iterations comprising: summing bits from each of the m-bit numbers to generate a summation result, wherein all bits summed occupy the same bit position in their corresponding numbers; comparing the summation result with a threshold, wherein the threshold is calculated based on i or p; setting the bit of the selected number based on the result of the comparison; and for each of the m-bit numbers, selectively updating the bit occupying the next bit position in the m-bit number based on the result of the comparison and the value of the bit from the m-bit number, wherein in the first iteration, the most significant bit from each of the m-bit numbers is summed and the most significant bit of the selected number is set, and each subsequent iteration sums the bits occupying consecutive bit positions in its own corresponding number and sets the next bit of the selected number, and wherein the method comprises outputting data indicating the selected number.
[0010] Outputting data indicative of a selected number may include: outputting the selected number; or outputting an indication of a position of the selected number within the n m-bit number.
[0011] Setting the bit of the selected number based on the result of the comparison may include: setting the bit of the selected number to one in response to determining that the summed result exceeds the threshold; and setting the bit of the selected number to zero in response to determining that the summed result is less than the threshold.
[0012] In the rth iteration, summing bits from each of the m-bit numbers to generate a summation result may include summing bits with bit index mr from each of the m-bit numbers to generate a summation result, wherein each bit is an original bit from one of the m-bit numbers or an updated bit from a previous iteration.
[0013] The selected number may be the i-th largest number from a set of n m-bit numbers, and the threshold may be equal to i. In other examples, the selected number may be the p-th smallest number from a set of n m-bit numbers, and the threshold may be equal to (np) or (n-p+1).
[0014] Selectively updating a bit occupying a next bit position in an m-bit number based on a result of a comparison and a value of a bit from the m-bit number may include: selectively setting a flag associated with the m-bit number based on a result of a comparison and a value of a bit from the m-bit number; and selectively updating a bit occupying a next bit position in the m-bit number based on the values of one or more flags associated with the m-bit number.
[0015] Selectively setting a flag associated with the m-bit number based on a result of the comparison and a value of a bit from the m-bit number may include: in response to determining that a sum result exceeds a threshold value and determining that the value of the bit is zero, setting a minimum flag associated with the m-bit number; and in response to determining that a sum result is less than a threshold value and determining that the value of the bit is one, setting a maximum flag associated with the m-bit number, and wherein selectively updating a bit occupying a next bit position in the m-bit number based on the value of one or more flags associated with the m-bit number includes: in response to determining that a maximum flag associated with the m-bit number is set, setting the bit occupying the next bit position in the m-bit number to one; in response to determining that a minimum flag associated with the m-bit number is set, setting the bit occupying the next bit position in the m-bit number to zero; and in response to determining that both the maximum flag and the minimum flag associated with the m-bit number are not set, leaving the bit occupying the next bit position in the m-bit number unchanged.
[0016] Selectively setting a flag associated with the m-bit number based on the result of the comparison and the value of a bit from the m-bit number may include: in response to determining that the summed result exceeds a threshold value and determining that the value of the bit is zero, setting a specific flag associated with the m-bit number; and in response to determining that the summed result is less than the threshold value and determining that the value of the bit is one, setting a specific flag associated with the m-bit number and updating the threshold value by an amount equal to the summed result, and wherein selectively updating a bit occupying a next bit position in the m-bit number based on the value of one or more flags associated with the m-bit number includes: in response to determining that a specific flag is set, setting the bit occupying the next bit position in the m-bit number to a predefined value; and in response to determining that a specific flag associated with the m-bit number is not set, leaving the bit occupying the next bit position in the m-bit number unchanged. The predefined value may be zero.
[0017] Selectively setting a flag associated with the m-bit number based on the result of the comparison and the value of a bit from the m-bit number may include: in response to determining that the summed result exceeds a threshold and determining that the value of the bit is zero, setting a specific flag associated with the m-bit number and updating the threshold by an amount equal to n minus the summed result; and in response to determining that the summed result is less than the threshold and determining that the value of the bit is one, setting a specific flag associated with the m-bit number, and wherein selectively updating a bit occupying a next bit position in the m-bit number based on the value of one or more flags associated with the m-bit number includes: in response to determining that a specific flag is set, setting the bit occupying the next bit position in the m-bit number to a predefined value; and in response to determining that a specific flag associated with the m-bit number is not set, leaving the bit occupying the next bit position in the m-bit number unchanged. The predefined value may be one.
[0018] The method may further include: determining the number of m-bits for which the associated flag is set; and in response to determining that n-1 of the m-bits' associated flags are set, outputting data of the m-bits indicating that the associated flag is not set.
[0019] Selectively updating a bit occupying a next bit position in the m-bit number based on a result of the comparison and a value of a bit from the m-bit number may include: in response to determining that a sum result exceeds a threshold and determining that the value of the bit is zero, updating all bits in the m-bit number to zero; and in response to determining that the sum result does not exceed the threshold and determining that the value of the bit is one, updating all bits in the m-bit number to one.
[0020] Selectively updating a bit occupying a next bit position in the m-bit number based on a result of the comparison and a value of a bit from the m-bit number may include: in response to determining that a sum result exceeds a threshold and determining that the value of the bit is zero, updating all bits in the m-bit number to zero; and in response to determining that the sum result does not exceed the threshold and determining that the value of the bit is one, updating all bits in the m-bit number to zero and reducing the threshold by an amount equal to the sum result.
[0021] The method may include m iterations, and in the mth iteration, summing the least significant bit from each of the m-bit numbers and setting the least significant bit of the selected number.
[0022] A second aspect provides a hardware logic unit, the hardware logic unit being arranged to select the i-th largest or the p-th smallest number from a set of n m-bit numbers, wherein i, p, m and n are integers, the hardware logic unit being arranged to operate iteratively and comprising: summation logic, the summation logic being arranged to sum bits from each of the m-bit numbers in each iteration to generate a summation result, wherein all bits to be summed occupy the same bit position in their corresponding numbers, so that in the first iteration, the most significant bit from each of the m-bit numbers is summed, and each subsequent iteration sums bits occupying consecutive bit positions in its own corresponding number; comparison logic, the comparison logic being arranged to compare the summation result generated by the summation logic in the iteration with a threshold in each iteration and set the bit of the selected number based on the result of the comparison, wherein the threshold is calculated based on i or p; update logic, the update logic being arranged to selectively update the bit occupying the next bit position in the m-bit number in each iteration and for each of the m-bit numbers based on the result of the comparison in the iteration and the value of the bit from the m-bit number; and an output, the output being arranged to output data indicating the selected number.
[0023] The hardware logic unit may further include: flag control logic, the flag control logic being arranged to selectively set a flag associated with the m-bit number based on a result of the comparison and a value of a bit from the m-bit number; and wherein the update logic is arranged to selectively update a bit occupying a next bit position in the m-bit number based on the value of one or more flags associated with the m-bit number.
[0024] The flag control logic may include: a minimum flag logic block, the minimum flag logic block is arranged to set a minimum flag associated with the m-bit number in response to determining that the sum result exceeds a threshold and determining that the value of the bit is zero; and a maximum flag logic block, the maximum flag logic block is arranged to set a maximum flag associated with the m-bit number in response to determining that the sum result is less than the threshold and determining that the value of the bit is one, and wherein the update logic is arranged to: in response to determining that the maximum flag associated with the m-bit number is set, set the bit occupying the next bit position in the m-bit number to one; in response to determining that the minimum flag associated with the m-bit number is set, set the bit occupying the next bit position in the m-bit number to zero; and in response to determining that both the maximum flag and the minimum flag associated with the m-bit number are not set, leave the bit occupying the next bit position in the m-bit number unchanged.
[0025] The flag control logic may be arranged to: in response to determining that the summed result exceeds the threshold and the value of the bit is zero, set a specific flag associated with the m-bit number; and in response to determining that the summed result is less than the threshold and the value of the bit is one, set the specific flag associated with the m-bit number and update the threshold by an amount equal to the summed result, and wherein the update logic is arranged to: in response to determining that the specific flag is set, set the bit occupying the next bit position in the m-bit number to a predefined value; and in response to determining that the specific flag is not set, leave the bit occupying the next bit position in the m-bit number unchanged. The predefined value may be zero.
[0026] The flag control logic may be arranged to: in response to determining that the summed result exceeds the threshold and determining that the value of the bit is zero, set a specific flag associated with the m-bit number and update the threshold value by an amount equal to n minus the summed result; and in response to determining that the summed result is less than the threshold and determining that the value of the bit is one, set a specific flag associated with the m-bit number, and wherein the update logic is arranged to: in response to determining that the specific flag is set, set the bit occupying the next bit position in the m-bit number to a predefined value; and in response to determining that the specific flag is not set, leave the bit occupying the next bit position in the m-bit number unchanged. The predefined value may be one.
[0027] The hardware logic unit may further include: an early exit hardware logic block, the early exit hardware logic block being arranged to determine the number of m-bit bits whose associated flags are set; and in response to determining that n-1 of the m-bit bits' associated flags are set, outputting the m-bit bits whose associated flags are not set as the selected number.
[0028] The update logic may be arranged to: in response to determining that the sum result exceeds a threshold and that the value of the bit is zero, update all bits in the m-bit number to zero; and in response to determining that the sum result does not exceed the threshold and that the value of the bit is one, update all bits in the m-bit number to one.
[0029] The update logic may be arranged to: in response to determining that the sum result exceeds a threshold and that the value of the bit is zero, update all bits in the m-bit number to zero; and in response to determining that the sum result does not exceed the threshold and that the value of the bit is one, update all bits in the m-bit number to zero and reduce the threshold by an amount equal to the sum result.
[0030] A third aspect provides a hardware logic unit, which is configured to execute the method described above.
[0031] A fourth aspect provides a method for manufacturing the hardware logic unit described in detail above using an integrated circuit manufacturing system.
[0032] A fifth aspect provides an integrated circuit definition data set which, when processed in an integrated circuit manufacturing system, configures the integrated circuit manufacturing system to manufacture the hardware logic unit detailed above.
[0033] A sixth aspect provides a computer-readable storage medium having stored thereon a computer-readable description of an integrated circuit, which, when processed in an integrated circuit manufacturing system, causes the integrated circuit manufacturing system to manufacture the hardware logic unit detailed above.
[0034] The seventh aspect provides an integrated circuit manufacturing system, the integrated circuit manufacturing system comprising: a computer-readable storage medium, on which a computer-readable description of the integrated circuit is stored, the computer-readable description describing the hardware logic unit detailed above; a layout processing system, the layout processing system is configured to process the integrated circuit description to generate a circuit layout description of the integrated circuit embodying the hardware logic unit; and an integrated circuit generation system, the integrated circuit generation system is configured to manufacture the hardware logic unit according to the circuit layout description.
[0035] An eighth aspect provides a method implemented in hardware logic for generating and selecting numbers, the method comprising: executing a most significant bit first (MSB first) iterative generation process for generating a set of n numbers; while executing the MSB first iterative generation process for generating a set of n numbers, executing an MSB first iterative selection process for selecting the i-th largest or p-th smallest number from the set of n numbers, wherein i, p and n are integers; and in response to the MSB first iterative selection process determining that a particular number among the numbers in the set of n numbers will not be a selected number, pausing generation of the particular number by the MSB first iterative generation process after at least one of the bits of the particular number has been generated and before all the bits of the particular number have been generated, wherein the method comprises outputting data indicating the selected number.
[0036] The MSB-first iterative generation process may be a coordinate rotation digital computer (CORDIC) process or an on-line arithmetic process.
[0037] Performing the MSB-first iterative selection process may include performing a plurality of iterations, wherein each of the iterations may include: summing bits from each of the numbers in the set to generate a summed result, wherein all bits summed occupy the same bit position within their respective numbers; comparing the summed result to a threshold, wherein the threshold is calculated based on i or p; setting a bit of the selected number based on the result of the comparison; and selectively updating a bit occupying the next bit position in the number based on the result of the comparison and the value of the bit from the number for each of the numbers in the set. In a first iteration, the most significant bit of each of the numbers in the set may be summed and the most significant bit of the selected number may be set, and each subsequent iteration sums bits occupying consecutive bit positions in its respective number and sets the next bit of the selected number.
[0038] The selected number may be the i-th largest number from a set of n numbers, and the threshold is equal to i. In other examples, the selected number may be the p-th smallest number from a set of n numbers, and the threshold is equal to (np) or (n-p+1).
[0039] Each of the n numbers in the set can be an m-digit number if fully generated.
[0040] Outputting data indicative of the selected number may include: outputting the selected number; or outputting an indication of a position of the selected number within a set of n numbers.
[0041] The ninth aspect provides a processing unit, which is configured to generate and select numbers, and the processing unit includes: a generation logic unit, which is implemented in hardware and configured to perform an MSB-first iterative generation process for generating a set of n numbers; a selection logic unit, which is implemented in hardware and configured to operate simultaneously with the generation logic unit and configured to perform an MSB-first iterative selection process for selecting the i-th largest or p-th smallest number from the set of n numbers, where i, p and n are integers; and an output, which is arranged to output data indicating a selected number, wherein the processing unit is configured to cause the generation logic unit to pause the generation of the specific number by the MSB-first iterative generation process in response to the selection logic unit determining that a specific number in the set of n numbers will not be the selected number, after at least one of the bits of the specific number has been generated and before all the bits of the specific number have been generated.
[0042] The selection logic may include: summation logic arranged to sum bits from each of the numbers in each iteration to generate a summation result, wherein all bits summed occupy the same bit position within their respective numbers; comparison logic arranged to compare the summation result generated by the summation logic in the iteration with a threshold value and set the bit of the selected number based on the result of the comparison in each iteration, wherein the threshold value is calculated based on i or p; and update logic arranged to selectively update the bit occupying the next bit position in the number in each iteration and for each of the numbers based on the result of the comparison in the iteration and the value of the bit from the number. The summation logic may be arranged so that the most significant bit from each of the numbers is summed in the first iteration, and each subsequent iteration sums bits occupying consecutive bit positions in its respective number.
[0043] A tenth aspect provides a processing unit configured to execute the method for generating and selecting numbers as detailed above.
[0044] An eleventh aspect provides a method for manufacturing the processing unit described above using an integrated circuit manufacturing system.
[0045] A twelfth aspect provides an integrated circuit definition data set which, when processed in an integrated circuit manufacturing system, configures the integrated circuit manufacturing system to manufacture the processing unit detailed above.
[0046] A thirteenth aspect provides a computer-readable storage medium having stored thereon a computer-readable description of an integrated circuit, which, when processed in an integrated circuit manufacturing system, causes the integrated circuit manufacturing system to manufacture the processing unit detailed above.
[0047] The fourteenth aspect provides an integrated circuit manufacturing system, the integrated circuit manufacturing system comprising: a computer-readable storage medium, on which a computer-readable description of the integrated circuit is stored, the computer-readable description describing the processing unit detailed above; a layout processing system, the layout processing system is configured to process the integrated circuit description to generate a circuit layout description of the integrated circuit embodying the processing unit; and an integrated circuit generation system, the integrated circuit generation system is configured to manufacture the processing unit based on the circuit layout description.
[0048] A number sorting hardware logic unit and / or a processor may be embodied in hardware on an integrated circuit, the processor including hardware logic configured to perform one of the methods described herein. A method for manufacturing a number sorting hardware logic unit and / or a processor at an integrated circuit manufacturing system may be provided, the processor including hardware logic configured to perform one of the methods described herein. An integrated circuit definition data set may be provided, the integrated circuit definition data set, when processed in the integrated circuit manufacturing system, configures the system to manufacture a number sorting hardware logic unit and / or a processor, the processor including hardware logic configured to perform one of the methods described herein. A non-transitory computer-readable storage medium may be provided, the non-transitory computer-readable storage medium having a computer-readable description of an integrated circuit stored thereon, the computer-readable description, when processed, causing a layout processing system to generate a circuit layout description, the circuit layout description is used in the integrated circuit manufacturing system to manufacture a number sorting hardware logic unit and / or a processor, the processor including hardware logic configured to perform one of the methods described herein.
[0049] An integrated circuit manufacturing system can be provided, the integrated circuit manufacturing system comprising: a non-temporary computer-readable storage medium, on which a computer-readable integrated circuit description describing a number sorting hardware logic unit and / or a processor is stored, the processor comprising hardware logic configured to perform one of the methods described herein; a layout processing system, the layout processing system being configured to process the integrated circuit description so as to generate a circuit layout description of an integrated circuit embodying the number sorting hardware logic unit and / or the processor, the processor comprising hardware logic configured to perform one of the methods described herein; and an integrated circuit generation system, the integrated circuit generation system being configured to manufacture the number sorting hardware logic unit and / or the processor according to the circuit layout description, the processor comprising hardware logic configured to perform one of the methods described herein.
[0050] A computer program code for performing any of the methods described herein may be provided. A non-transitory computer readable storage medium may be provided, on which computer readable instructions are stored, which, when executed at a computer system, cause the computer system to perform any of the methods described herein.
[0051] It will be clear to those skilled in the art that the above-described features may be combined as desired and may be combined with any aspects of the examples described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Examples will now be described in detail with reference to the accompanying drawings, in which:
[0053] Figure 1is a schematic diagram of an example hardware arrangement for sorting 4 inputs;
[0054] Figure 2 A graphical representation showing a set of n m-bit numbers input to the methods and hardware described herein;
[0055] Figure 3A is a flow chart illustrating a first example method of computing the i-th largest number from a set of n m-digit inputs;
[0056] Figure 3B is a flow chart illustrating a first example method of computing the pth smallest number from a set of n m-digit inputs;
[0057] Figure 3C is shown in more detail from Figure 3A and 3B A flowchart of the operation of the method;
[0058] Figure 3D yes Figure 3C A circuit diagram of an example hardware implementation of the operations shown;
[0059] Figure 3E Show Figure 3D Truth table for the hardware arrangement shown;
[0060] Figure 3F Show implementation Figure 3A a schematic diagram of an example hardware arrangement for operations in the method of;
[0061] Figure 3G and 3H Show implementation Figure 3A and 3B Schematic diagrams of two different example hardware arrangements for operations in the method;
[0062] Fig. 3I Show implementation Fig. 9 a schematic diagram of an example hardware arrangement for operations in the method of;
[0063] Figure 4 Show Figure 3A An example of the operation of the method;
[0064] Figure 5 is a flow chart illustrating a second example method for calculating the i-th largest number from a set of n m-digit inputs;
[0065] Figure 6 Show Figure 5 An example of the operation of the method;
[0066] Fig. 7A is a flow chart illustrating a third example method for calculating the i-th largest number from a set of n m-digit inputs;
[0067] Figure 7B is a flowchart illustrating another example method of calculating the pth smallest number from a set of n m-digit inputs;
[0068] Figure 8 Show Fig. 7A An example of the operation of the method;
[0069] Fig. 9 is a flow chart illustrating a fourth example method for calculating the i-th largest number from a set of n m-digit inputs;
[0070] Fig.10 It is shown that the use Figure 3A A schematic diagram of the method to perform the sorting;
[0071] Fig.11A , 11B and 11C are graphs showing synthetic results of the methods described herein;
[0072] Fig. 12A is a schematic diagram of a hardware logic unit arranged to select the i-th largest number or the p-th smallest number from a set of n m-bit inputs;
[0073] Fig. 12B is a schematic diagram of a processing unit arranged to generate and select numbers;
[0074] Fig. 12C A computer system is shown in which a graphics processing system is implemented;
[0075] Fig.13 An integrated circuit manufacturing system is shown for producing an integrated circuit embodying a graphics processing system; and
[0076] Fig.14 is a flowchart of an example method for generating and selecting numbers from a set of n numbers, where the n numbers are iteratively generated starting from the MSB.
[0077] The accompanying drawings illustrate various examples. Those skilled in the art will appreciate that the element boundaries (e.g., boxes, groups of boxes, or other shapes) shown in the accompanying drawings represent one example of a boundary. In some examples, one element may be designed as multiple elements, or multiple elements may be designed as one element. Where appropriate, common reference numerals are used throughout the figures to indicate similar features. DETAILED DESCRIPTION
[0078] The following description is presented by way of example to enable one skilled in the art to make and use the invention.The present invention is not limited to the embodiments described herein, and various modifications to the disclosed embodiments will be apparent to those skilled in the art.
[0079] Embodiments will now be described by way of example only.
[0080] There are many applications where you need to select the i-th largest or p-th smallest integer from a set of integers. Typically, this is done using a sorting algorithm or sorting network (e.g., as described above in reference Figure 1 The hardware may be implemented by sorting the integers into an ordered list (described above), and then the relevant integers from the list may be output. However, the size of the resulting hardware may be large, even if redundant logic (which does not affect the desired output) is removed as part of the synthesis of the hardware.
[0081] This article describes a method and hardware for selecting the i-th largest or p-th smallest number from a set of n m-digit numbers without first sorting the set of numbers. Using the method described in this article, the hardware is smaller than the sorting network, for example, its area is adjusted to O(n*m) instead of O(m*n*(ln(n) 2 )). In addition, because the method is iterative, the hardware area used to implement the method can be made smaller by trading off performance / throughput (e.g., by synthesizing only one iteration or less than m iterations and then reusing the hardware logic over multiple cycles). In addition, the method enables performance / throughput to be increased at the expense of additional area (e.g., by increasing the number of bits evaluated in each iteration to above 1).
[0082] The methods described herein can be applied to numbers represented in signed or unsigned fixed-point format, floating-point format, and signed or unsigned normalized format. For example, for signed numbers and floating-point numbers, the top bit of each number is negated when input to the method and output from the method. For unsigned fixed-point numbers, no changes are required to the method. For the normalized format, the bit string is considered to be a regular unsigned (or signed) number. In various instances, the number can be an integer. In various instances, the number can be a binary approximation (e.g., 1 / 3, or the square root of 2) of a value that does not have a finite binary representation in a standard fixed-point format, where the binary approximation is generated one bit at a time.
[0083] The method involves checking the most significant bit (MSB) of each number from the set, and setting one or more flags or mask bits based on the results of the analysis. The method is then repeated, selecting the next bit from each number in the set, adjusting the bit value according to the flag or mask bit, and then performing the same analysis (or an extremely similar analysis) as the analysis performed on the MSB. As in the first iteration (which involves the MSB), one or more flags or mask bits may be set based on the results of the analysis. The method may iterate over each of the m bits in the number to determine which of the numbers is the i-th largest or p-th smallest number. In each iteration, a bit from the output number (i.e., the i-th largest or p-th smallest number) is set, and the output from the method may be the output number itself or other data identifying (or indicating) the output number from a set of n m-bit numbers (e.g., in the form of an index of the output number in a set of n numbers).
[0084] In describing various embodiments and examples below, the following reference signs are used, which may refer to Figure 2 Describe:
[0085] ●n is the number of numbers in the set 200,
[0086] ●k is a numeric index and its range is between 0 and (n-1),
[0087] ●N k are numbers in set 200 such that the first number 202 in set 200 is N0 and the last number 204 in set 200 is N n-1 ,
[0088] m is the number of bits in each number in the set, and each number in the set includes the same number of bits,
[0089] j is a bit index and ranges between (m-1) for the MSB 206 in each number and 0 for the least significant bit (LSB) 208 in each number, and in instances where the method analyzes a single bit (from MSB to LSB) from each number in the set in each iteration, j may also be referred to as an iteration index,
[0090] • r is the iteration number ranging between one (for the first iteration, where j=m-1) and m (for the last iteration, where j=0), so in the example where one bit position is considered per iteration, j=mr.
[0091] ●x k [j] is the index N k bit j of the digit, where j=m-1 for the MSB and j=0 for the LSB,
[0092] ●zk [j] is the index N k The modification bit j,
[0093] i and p are integers in the range of 1 to n, and the method described herein identifies the i-th largest or p-th smallest number from a set of numbers, and the desired number, i.e., the i-th largest or p-th smallest number, is referred to herein as the output number,
[0094] ●min k [j] is the number of analysis N k The number N after the jth digit k The minimum flag (or min_flag),
[0095] min k [m] is the number N k The initial (i.e., starting) value of the minimum flag,
[0096] ●max k [j] is the number of analysis N k The number N set after the jth bit k The maximum flag (or max_flag),
[0097] ●max k [m] is the number N k the initial (i.e., starting) value of the maximum flag,
[0098] ●flag k [j] is the number of N in the analysis in the example using a single marker k The number N set after the jth bit k , and
[0099] ●flag k [m] is the number N k The initial (i.e., starting) value of a single flag.
[0100] Figure 2 A graphical representation of a set 200 of n m-bit numbers is shown as input to the methods and hardware described herein. The numbers within the set are not in any particular order, and thus the value of k merely identifies the position of the number in the set (and is therefore used to refer to a particular number N). k ), and does not provide information about the number N k Any information about the relative size compared to other numbers in set 200.
[0101] The methods and hardware described herein may be used to find the i-th largest or p-th smallest number from a set 200 of n m-digit numbers without first sorting the set of numbers.
[0102] Figure 3AFlowchart of a first example method for calculating (or identifying) the i-th largest number from a set of n m-digit inputs. In various examples, the value of i is fixed, and in other examples, the value of i is an input variable. Figure 3A As shown, the method is iterative and uses two flags for each number: min_flag and max_flag (ie, a total of 2n flag bits). k [] indicates that the specified number is less than the output number (the i-th largest number) when set, and the maximum flag max_flag or max k [] indicates when set that a particular number is greater than the output number. Initially (when j=m), the flags may all be unset (e.g., set to zero) unless some pre-masking has been performed (as described below). As described below, the method builds the output number one bit per iteration.
[0103] In the first iteration (r=1, j=m-1), the MSB 206 of each number is summed, and if the sum of the MSBs is greater than or equal to i ('yes' in block 302), then this means that the MSB of the i-th largest number from the input set (i.e., the MSB of the output number) is one, and the MSB of the output number can be set to one (block 305). However, if the sum of the MSBs is less than i ('no' in block 302), then this means that the MSB of the i-th largest number from the input set (i.e., the MSB of the output number) is zero, and the MSB of the output number can be set to zero (block 307). In addition, in response to determining that the sum of the MSBs is greater than or equal to i ('yes' in block 302), min_flag is set for all those numbers with MSB=0 (block 304), and in response to determining that the sum of the MSBs is not greater than or equal to i ('no' in block 302), max_flag is set for all those numbers with MSB=1 (block 306).
[0104] The second iteration (r=2, j=m-2) begins by taking the next bit from each number and modifying the bit using the flag value (block 308). Specifically, in the case where min_flag or max_flag is set for a number, the value of the bit may be changed (i.e., the value of the bit is selectively updated based on the flag and the value of the bit), such as Figure 3CAs shown. If the max_flag of the number is set ('yes' in box 310), then regardless of whether the bit value is actually one or zero, the bit from the number (that is, the second most significant bit from the number used for the second iteration) is set to one (box 312). Similarly, if the min_flag of the number is set ('yes' in box 314), then regardless of whether the bit value is actually one or zero, the bit is set to zero (box 316). According to the method described herein, a number cannot be set with both max_flag and min_flag. If none of the flags are set ('no' in boxes 310 and 314), the value of the bit is left unchanged (box 318).
[0105] The change of the bit (in block 308, such as Figure 3C ) can also be described by the following logic equation:
[0106]
[0107] where · represents a logical AND operation and + represents a logical OR operation. A corresponding hardware arrangement 320 that can be replicated n times (once for each number in the set 200) is Figure 3D , and includes a NOT gate 322, an AND gate 324, and an OR gate 326. Figure 3D As shown, the current bit is combined with the inverse form of the current minimum flag (set in the previous iteration) in AND gate 324, and then the output of AND gate 324 is combined with the current maximum flag (set in the previous iteration) in OR gate 326. Figure 3E A truth table of the reachable states of the hardware arrangement 320 is shown in .
[0108] Then the changed bits (generated in box 308) are summed, and if the sum is greater than or equal to i ('yes' in box 302), then this means that the next bit of the i-th largest number from the input set (i.e., the next bit of the output number) is one, and the next bit of the output bit can be set to one (box 305); however, if the sum is less than i ('no' in box 302), then this means that the next bit of the i-th largest number from the input set (i.e., the next bit of the output number) is zero, and the next bit of the output number can be set to zero (box 307). In this way, the method builds the output number (in boxes 305 and 307) one bit at a time and one bit at a time. In response to determining that the sum is greater than or equal to i ('yes' in block 302), min_flag is set for all those numbers whose modified bits are equal to zero (block 304), and in response to determining that the sum is not greater than or equal to i ('no' in block 302), max_flag is set for all those numbers whose modified bits are equal to one (block 306).
[0109] The summation of the bits (in block 302) may also be described by the following logic equation:
[0110]
[0111] The corresponding hardware arrangement 328 includes a plurality of adders (eg, a plurality of full adders 330 that each add three bits, ie, for three different values of k, followed by one or more ripple-carry adders 332). Figure 3F An example hardware arrangement is shown in ; however, the summation may be implemented in hardware in other ways.
[0112] The updating of the minimum flag (in block 304) and the updating of the maximum flag (in block 306) may also be described by the following logic equations:
[0113] y[j]=(sum j ≥i)? 1:0
[0114]
[0115] In these last two equations, the first term refers to the flag value from the previous iteration and is used to ensure that the flag does not change if it has already been set by an earlier bit in the number. k [j]) can also be done by simply using z k [j] Replace x k [j] Using the same last two equations, since x is k [j] = z k [j], and if any flag is 1, then x k The value of [j] is irrelevant in the above equation, however, a hardware implementation using the modified bits may be larger (in terms of hardware logic area) than using the original bits. Figure 3G and 3H , and includes a combination of NOT, AND, and OR gates (or other logic blocks that implement the same functionality as NOT, AND, and OR gates and may be referred to as NOT, AND, and OR logic blocks).
[0116] Next, we can repeat for all m bits in the input number Figure 3Ato establish the output number, or in various examples, there may be additional logic to identify when a result has been obtained before all m iterations are complete (i.e., when there is only one different value in the number set 200 where both max_flag and min_flag are not set) and output the result at that stage. In various examples, additional logic may additionally or alternatively be used to limit the number of iterations performed (i.e., provide r, r max The maximum value of the input number is then output at that stage. In such instances, only a portion of the output number is established, and therefore, the remaining bits or the entire output number can be obtained by selecting a number from the input number set based on the flag value, or alternatively, other data identifying the output number (e.g., number index k) can be output. However, if the input number set can contain duplicate values (i.e., if all n numbers in the input set are not guaranteed to be unique), so that there can be more than one input number that is the i-th largest number in the n m-bit input set, then this will add complexity to the additional logic, and therefore the benefits of having the additional logic may be reduced or lost. This additional logic may be referred to as an early exit logic unit.
[0117] min k [j] and max k [j] Performing a bitwise OR of the signal and then negating the result yields a signal if and only if bits m-1 to j of the output number are equal to the number N. k This signal has a 1 in the kth bit only when bits m-1 to j in match.
[0118] Figure 3A The method can refer to Figure 4 The example shown is described where n=5 and i=3. In the first iteration 402, the MSB of each number is summed and the result is 3. When the sum is equal to i ('yes' in block 302), the MSB of the output number 403 is set to one (block 305) and min_flag is set for those numbers whose MSB is zero (block 304). In the example shown, min_flag is set for numbers N1 and N4.
[0119] In the second iteration, the next bit in each of the five numbers is first modified based on the flag value (block 308) to obtain Z k [m-2] values, and in this example, because only the flags for numbers N1 and N4 are set, only these bits are modified from one to zero. The modified bits 404 are then summed, and the result is 1. When the sum is less than i ('no' in block 302), the next bit of the output number is set to zero (block 307), and max_flag is set for those numbers whose modified bits are one (block 306). In the example shown, max_flag is set for number N2.
[0120] Again, the next bit in each of the five numbers is modified based on the flag value (block 308) to obtain z k [m-3] value to start the third iteration, and in this example, because the flags of numbers N1, N2 and N4 are set, only these bits are affected, but if Figure 4 As shown, although the value of the 3rd bit of numbers N2 and N4 flips (from 0 to 1 due to max_flag for number N2, and from 1 to 0 due to min_flag for number N4), the 3rd bit of number N1 is not modified because it is already zero and it is the min_flag that is set. As previously described, the modification bits 406 are summed, and in this case, the result is 2. When the sum is less than i ('no' in box 302), the next bit of the output number is set to zero (box 307), and max_flag is set for those numbers with a modification bit of one (box 306). In the example shown, max_flag is set for number N0. At this point, only one number has an unflagged number, namely, number N3, and therefore it is the output number, namely, the i-th largest number from the input set. At this point, the method can stop (e.g., if logic is provided to evaluate the flags and determine when only one number has an unflagged number), or the method can continue until all bits have been evaluated.
[0121] It is understandable that if Figure 4 In the example of m=8, and the five input values (N0-N4) are 101xxxxx, 010xxxxx, 110xxxxx, 100xxxxx and 011xxxxx (where each x can represent 0 or 1), then Figure 4 As shown, the third largest number is determined to be 100xxxxx, and in this example, this can be determined by analyzing only the first three bits of the input value. As described above, after the third iteration, only the three most significant bits (bit 100) of the output number 403 have been established, and the remaining five bits can be determined by continuing the remaining five iterations of the method or by selecting N3 based on the flag value and adding the 5 LSBs to the output number 403 that has been generated or outputting N3 and discarding the three bits that have been established. Alternatively, data identifying the output number (e.g., index k=3) can be output instead of the output number itself.
[0122] although Figure 3A A first example method for calculating (or identifying) the i-th largest number from a set of n m-digit inputs is shown, but a very similar method can be used to calculate (or identify) the p-th smallest number from a set of n m-digit inputs, and Figure 3B An example is shown in Figure 3BIn the method, compared with Figure 3A The only difference in the approach is that the sum is compared to the value of (np), and the comparison is changed from 'greater than or equal to' (in Figure 3A 302) to switch to 'strictly greater than' (in Figure 3B Alternatively, in block 303 of Figure 3B The comparison performed in the box 303 of may be whether the sum is greater than or equal to (n+1-p). If the sum of the MSBs (or the changed bits for subsequent iterations) is greater than (np) ('yes' in box 303), then the MSB of the output number is set to one (box 305), and min_flag is set for all those numbers where the MSB or the changed bits are equal to zero (box 304). However, if the sum is not greater than (np) ('no' in box 303), then the MSB of the output number is set to zero (box 307), and max_flag is set for all those numbers where the MSB or the changed bits are equal to one (box 306). As previously described, the method is then repeated for subsequent bits in each number in the input set to establish the output number. The method may terminate after m iterations, or, if additional hardware is provided, the method may terminate in response to determining that there is only one number in the input set that is not set with any flag.
[0123] In another example method of calculating the pth smallest number from a set of n m-digit inputs, all input numbers N can be k Bitwise Reverse Then you can use it when i=p Figure 3A method, as long as the final output is reversed back to its original form before being output.
[0124] Figure 5 is a flowchart of a second example method for calculating (or identifying) the i-th largest number from a set of n m-digit inputs. Similar to the first method, as described above and Figure 3A As shown, the second method is also iterative; however, it does not use a flag. Instead, the value of the number itself is updated based on the summation result (in box 302) (boxes 504 and 506). In response to determining that the sum of the MSBs is greater than or equal to i ('yes' in box 302), the MSB of the output number is set to one (box 305) and all bits in those numbers with MSB=0 are set to zero (box 504), and in response to determining that the sum of the MSBs is not greater than or equal to i ('no' in box 302), the MSB of the output number is set to zero (box 307) and all bits in those numbers with MSB=1 are set to one (box 506).
[0125] The second iteration (j=m-2) begins by taking the next bit from each number (block 508), where some of these numbers may be the original numbers and others are numbers modified in the first iteration (e.g., in blocks 504 or 506). The bits are summed, and depending on whether the sum is greater than or equal to i (in block 302), one or more of the remaining original numbers may be set to all zeros (in block 504) or all ones (in block 506), and another bit is added to the output number (in blocks 305 or 307).
[0126] The updating of the number (in blocks 504 and 506) can also be described by the following logic equation:
[0127]
[0128] Where h = 0, 1, ..., m-1.
[0129] Next, we can repeat for all m bits in the input number Figure 5 or, in various examples, there may be a time to identify an earlier obtained result (ie, an update number N that is not all ones or all zeros). k There is only one different value in the time) and additional logic (as described above) to output the result at that stage. As mentioned above, this becomes more complicated if there is more than one i-th largest number.
[0130] Figure 5 The method can refer to Figure 6 The example shown is described where n = 5 and i = 3. In the first iteration 602, the MSB of each number is summed and the result is 3. When the sum is equal to i ('yes' in box 302), the MSB of the output number 603 is set to one (box 305), and those numbers with an MSB of zero are modified so that all of their bits are equal to zero (box 504), and in the example shown, the numbers N1 and N4 are modified in this way.
[0131] In the second iteration 604, the next bit in each of the five numbers (the original numbers N0, N2, N3 and the modified numbers N1 and N4) is summed and the result is 1. When the sum is less than i ('No' in box 302), the next most significant bit of the output number 603 is set to zero (box 307), and those numbers with bits that are one are modified so that all of their bits are equal to one (box 506), and in the example shown, the number N2 is modified in this way.
[0132] In the third iteration 606, the next bit in each of the five numbers (the original numbers N0 and N3 and the modified numbers N1, N2, and N4) is summed, and in this case, the result is 2. When the sum is less than i ('no' in block 302), the next bit of the output number 603 is set to one (block 307), and all bits of those numbers with bits that are one are set equal to one (block 506), and in the example shown, the number N0 is modified in this way. At this point, there is only one original number in the set (i.e., a number that is not all ones or all zeros), i.e., number N3, and it is therefore the output number, i.e., the i-th largest number from the input set. At this point, the method can stop (e.g., if logic is provided to evaluate the flags and determine when only one original unmodified number remains in the set), or the method can continue until all bits have been evaluated and all bits of the output number have been generated (one bit per iteration).
[0133] although Figure 5 A first example method for calculating (or identifying) the i-th largest number from a set of n m-digit inputs is shown, but a very similar method can be used to calculate (or identify) the p-th smallest number from a set of n m-digit inputs. Figure 3B This involves changing Figure 5 The method of only comparing the sum to the value of (np) or (n+1-p) instead of comparing to i (in block 302) is used. When (np) is used, if the sum of the bits is strictly greater than (np), then the next bit in the output number is set to one and those numbers with bits equal to zero are modified so that all bits are equal to zero (block 504), and in response to determining that the sum is not greater than (np), the next bit in the output number is set to zero and those numbers with bits equal to one are modified so that all bits are equal to one (block 506). Alternatively, when (n+1-p) is used, if the sum of the bits is greater than or equal to (n+1-p), then the next bit in the output number is set to one (box 305) and those numbers with bits equal to zero are modified so that all bits are equal to zero (box 504), and in response to determining that the sum is less than (n+1-p), the next bit in the output number is set to zero (box 307) and those numbers with bits equal to one are modified so that all bits are equal to one (box 506).
[0134] In another example method of calculating the pth smallest number from a set of n m-digit inputs, all input numbers N can be k Bitwise Reverse Then you can use it when i=p Figure 5 method, as long as the final output is reversed back to its original form before being output.
[0135] Fig. 7AFlowchart of a third example method for calculating (or identifying) the i-th largest number from a set of n m-digit inputs. Similar to the first and second methods, as described above and Figure 3A and 5 As shown, the third method is also iterative. Figure 5 302), the next bit in the output number is set to one (box 305), and all bits in those numbers where MSB=0 are set to zero (box 504). In response to determining that the sum of the MSBs is not greater than or equal to i ('no' in box 302), the next bit in the output number is set to zero (box 307), and all bits in those numbers where MSB=1 are set to zero (box 706), and then, in order for the next iteration to perform the correct comparison (in box 302), the value of i is then decremented by the total summed in the iteration (box 707).
[0136] The second iteration (j=m-2) is started by taking the next bit from each number (block 508), where some of these numbers may be the original numbers and others are numbers that were modified (e.g., modified to all zeros) in the first iteration. The bits are summed, and depending on whether the sum is greater than or equal to i (in block 302), the next bit in the output integer is set (in blocks 305 or 307), and the different numbers from the remaining original numbers may be set to all zeros (in blocks 504 or 706). Only if the sum is less than i ('no' in block 302), the value of i is further decremented by the total summed in the iteration (block 707). Next, the process may be repeated for all m bits in the input numbers. Fig. 7A The method may, or as described above, terminate when all but one number in the input set has been modified (eg, modified to all zeros), and the remaining number is then the output number.
[0137] The updating of the number (in blocks 504 and 706) can also be described by the following logic equation:
[0138]
[0139] Where h = 0, 1, ..., m-1.
[0140] Fig. 7A The method can refer to Figure 8The example shown is described where n = 5 and i = 3. In the first iteration 802, the MSB of each number is summed and the result is 3. When the sum is equal to i ('yes' in block 302), the MSB of the output number 803 is set to one (block 305) and those numbers whose MSB is zero are modified so that all of their bits are equal to zero (block 504), and in the example shown, the numbers N1 and N4 are modified in this way.
[0141] In the second iteration 804, the next bit in each of the five numbers (the original numbers N0, N2, N3 and the modified numbers N1 and N4) is summed and the result is 1. When the sum is less than i ('No' in block 302), the next bit in the output number is set to zero (block 307), and those numbers with bits that are one are modified so that all of their bits are equal to zero (block 706), and in the example shown, the number N2 is modified in this way. Next, the value of i is decremented by the sum result (i.e., decremented by one) so that for the next iteration, i=2.
[0142] In the third iteration 806, the next bit in each of the five numbers (the original numbers N0 and N3 and the modified numbers N1, N2, and N4) is summed, and in this case, the result is 1. When the sum is less than i ('no' in block 302), the next bit in the output number is set to zero (block 307), and those numbers with bits that are one are modified so that all of their bits are equal to zero (block 706), and in the example shown, the number N0 is modified in this way. The value of i can again be decremented by the sum result (i.e., decremented by one) so that i=1. At this point, there is only one original number in the set, i.e., number N3, and it is therefore the output number, i.e., the i-th largest number from the input set. At this point, the method can stop (e.g., if logic is provided to evaluate the flag and determine when only one original unmodified number remains in the set), or the method can continue until all bits have been evaluated.
[0143] although Fig. 7A A first example method for calculating (or identifying) the i-th largest number from a set of n m-digit inputs is shown, but a very similar method can be used to calculate (or identify) the p-th smallest number from a set of n m-digit inputs, such as Figure 7B As shown. Figure 3B ,although Figure 7B In block 303, a comparison is shown as to whether the sum is strictly greater than (np), but in other examples, the comparison may be whether the sum is greater than or equal to (n+1-p). In another example method of calculating the pth smallest number from a set of n m-digit inputs, all input numbers N may be k Bitwise Reverse Then you can use it when i=p Fig. 7Amethod, as long as the final output is reversed back to its original form before being output.
[0144] Fig. 9 Flowchart of a fourth example method for calculating (or identifying) the i-th largest number from a set of n m-digit inputs. Similar to the first, second, and third methods, as described above and Figure 3A , 5 As shown in FIG. 7A , the fourth method is also iterative. Figure 3A 302), the MSB of the output number is set to one (block 305), and the flag is set for those numbers with MSB=0 (block 904). In response to determining that the sum of the MSBs is not greater than or equal to i ('no' in block 302), the MSB of the output number is set to zero (block 307), and the flag is set for those numbers with MSB=1 (block 906), and then, in order for the next iteration to perform the correct comparison (in block 302), the value of i is then decremented by the total summed in the iteration (block 707).
[0145] The second iteration (j=m-2) begins by taking the next bit from each number and modifying the bit using the flag value (block 908). If the flag is set for the number, the bit is set to a predefined value, such as zero, regardless of whether the bit value is actually one or zero. If the flag is not set, the value of the bit is left unchanged.
[0146] The change of the bit (in block 908) can also be described by the following logic equation:
[0147]
[0148] The corresponding hardware arrangement 340 may be replicated n times (once for each number in the set 200) in Fig. 3I 34 and includes a NOT gate 342 and an AND gate 344, ie, the current bit is combined with the inverted form of the current flag (set in the previous iteration) in the AND gate.
[0149] Next, the modified bits (generated in block 908) are summed, and depending on whether the sum is greater than or equal to i (in block 302), one or more other flags may be set (in blocks 904 or 906). If the predefined value is zero and the sum is less than i ('no' in block 302), then the value of i is updated by decrementing the threshold by the total of the sums in that iteration (block 707). j Denotes the threshold i for comparison of the jth bit, this update can be described by:
[0150]
[0151] exist Fig. 9 In the variation of the method shown, if the predefined value is one (rather than zero) and the sum is greater than or equal to i ('yes' in block 302), then the value of i is updated by incrementing the threshold by the number of rows changed in this iteration, which number will be the total number of inputs minus the total summed in that iteration (block 707), and this update can be described by the following equation:
[0152]
[0153] Next, we can repeat for all m bits in the input number Fig. 9 method (or a variation of said method).
[0154] although Fig. 9 A first example method of calculating (or identifying) the i-th largest number from a set of n m-digit inputs is shown, but a very similar method can be used to calculate (or identify) the p-th smallest number from a set of n m-digit inputs by changing the comparison performed (e.g., from block 302 to block 303, or to a modified version of block 303 described above, in which the sum is compared to (n+1-p)) and the way the threshold is updated (e.g., from block 707 to block 717). In another example method of calculating the p-th smallest number from a set of n m-digit inputs, all input numbers N can be made k Reversal Then you can use it when i=p Fig. 9 method, as long as the final output is reversed back to its original form.
[0155] In other variations, rather than using a flag or updating a bit within a number, a bit in a mask may be set in response to a comparison (e.g., in block 302), and then, for subsequent iterations, the bit in the number may be combined with the mask value (e.g., using an AND gate). This has the same effect as the updating of the number (in the example described above).
[0156] Synthetic experiments have shown that in some cases, using two flags (such as in Figure 3A and 3B In some cases, the other methods described herein may be more efficient (e.g., if the cost of a register to store the flag is greater than the extra logic to update the bits in the input number, then Figure 5 The method may be more efficient, and / or the logic for performing the update of i or p may be made smaller. Fig. 7A , 7Band 9 can be more efficient).
[0157] Furthermore, unlike other methods described herein, using two flags (as in Figure 3A and 3B In the method of ) two subsets of the number set 200 can be identified. In the case of determining the i-th largest number or the p-th smallest number from a set of n m-digit inputs, the maximum flag can be used to identify all numbers in the set that are greater than the output number, and the minimum flag can be used to identify all numbers in the set that are less than the output number. This can then be used iteratively to implement a sorting operation, and Fig.10 The figure shows that in N k Each of them is an instance of a 4T sorter that works only when T is an integer. Fig.10 As shown, the input set includes 4T input numbers (n=4T), and then use Figure 3A The method (in block 1002) is used to find the largest four numbers in the input set (e.g., by setting i=4). Then, the largest four numbers from the input set are input to the 4 sorter (block 1004). These largest four numbers are also masked in the input set using a flag such as min_flag (block 1006), and the masked input set is then input to Figure 3A 1002). Alternatively, the mask may use max_flag, and in such an embodiment, the value of i for subsequent iterations of block 1002 is incremented by four for each iteration. In this case, each number will require an additional flag to track which numbers have been previously passed to the 4 sorter. Such a sorting arrangement may be implemented in a small area of hardware logic due to its iterative nature and the 4 sorter may be implemented in a very small area of hardware logic. In other examples, the sorting operation may use only Figure 3A and / or method 3B.
[0158] If you don't know yet Fig.10 The input N of the method k is unique, then there is an additional complication. For example, if five inputs are the same value, then there will be no possibility to pass these five values to the 4-sorter. In the case where min_flag is set, the solution is to count the number of outputs sent, and if this value is greater than 4, then ignore the output and set the value of i to 1 for the next iteration. In this way, all outputs from this next iteration will have the same value, and these values can bypass the 4-sorter. The min_flag on these inputs will be set, and the value of i will be set to 4 again for the next iteration. This additional logic will take up some area to implement in hardware and reduce the throughput of the system.
[0159] Fig. 12A1 is a schematic diagram of a hardware logic unit 1230 arranged to implement the method described above (ie, selecting the i-th largest or p-th smallest number from a set of n m-bit numbers). Fig. 12A As shown, the hardware logic unit 1230 includes a summing logic unit 1232, a comparison logic unit 1234, and an update logic unit 1236. The hardware logic unit 1230 also includes an output 1238 and may include an input ( Fig. 12A As described above, the summation logic unit 1232 is arranged to sum the bits from each m-bit number to generate a summation result, wherein all bits being summed occupy the same bit position within their respective numbers, and the example implementation is Figure 3F 1 and described above. The comparison logic unit 1234 is arranged to compare the summation result generated by the summation logic in the iteration with the threshold value and set the bit of the selected number based on the result of the comparison. The update logic unit 1236 is arranged to selectively update the bit occupying the next bit position in the m-bit number based on the result of the comparison in the iteration and the value of the bit from the m-bit number for each m-bit number, and the example embodiment is in Figure 3D and 3I The term 'selectively update' refers to the fact that the update logic may not necessarily change the value of any bit when performing an update.
[0160] like Fig. 12A As shown and described above, the hardware logic unit 1230 may further include a flag control logic unit 1235 and / or an early exit logic unit 1237. The flag control logic unit 1235 is arranged to selectively set a flag associated with the m-bit number based on the result of the comparison and the value of a bit from the m-bit number (wherein the term 'selectively' is used as described above to indicate that the flag value may change or remain unchanged in any iteration), and two example implementations of the flag control logic unit 1235 are described in Figure 3G and 3H 1236 and described above. The early exit logic unit 1237 is arranged to determine the time when a result has been obtained before all m iterations are completed, and then output the result (or data identifying the result) at that stage. This determination by the early exit logic unit 1237 can be based on the flag value and / or analysis of the output from the update logic unit 1236 (i.e., by determining that all bit values of all m-bit numbers except one of the m-bit numbers have been updated to a predefined value).
[0161] Figures 11A-11CThree example area-delay graphs are shown for hardware implementing the method described herein to find the median of a set of input numbers (the median is the curve labeled 'radix_median' in the graph). In such examples,
[0162]
[0163] And, the hardware is arranged to find the i-th largest item in a list of size n having Um values (i.e., an unsigned m-bit number). Fig.11A In the figure shown, n = 7 and m = 16. Fig. 11B In the figure shown, n = 7 and m = 11, and in Fig. 11C In the graph shown, n = 32 and m = 5. As shown in the graph, the hardware can be made smaller than alternative hardware (e.g., 'transposition_median' hardware using a bubble sort network with only the median output connected and 'batcher_median' hardware including a batcher odd-even merge sort network), but this smaller hardware is typically slower (i.e., it involves greater latency).
[0164] Although the method is described herein as evaluating a single bit in each iteration, in other examples, more than one bit (eg, a bit pair) may be evaluated in each iteration. This increases the size of the hardware implementing the method, but also increases the processing power of the hardware.
[0165] In the examples described above, there are n numbers in the input set, and the number sorting hardware logic unit is arranged to identify the i-th largest or p-th smallest number from the input set. However, in some examples, there may be fewer than n numbers in the input set, for example, there are n' numbers in the input set. In such examples, when a flag is used (e.g., in Figure 3A and 3B ), you can use pre-masking so that you can set the initial values of some flags (for example, to one instead of zero). Figure 3A In the case of the method of identifying the i-th largest number, n' numbers are located at the beginning of the set (such as numbers N0 to N n’-1 ), and it is the last (n-n')th flag that is set, i.e., the number N n’ To N n-1 In contrast, if you were to use Figure 3B The method of identifying the pth smallest number, then n' numbers are located at the end of the set (such as number N n-n’ To N n-1 ), and it is the first (n-n') flags that are set, that is, the numbers N0 to Nn’-n-1 The logo can be Fig. 9 In the method of , pre-masking is similarly applied to the flag. In addition, when other methods that do not use flags are used (e.g., Figure 5 , 7A and 7B), instead of using pre-masking, the input set of n' numbers can be padded by adding n-n' dummy input numbers whose bits are all 0s when the hardware is configured to recognize the i-th largest number, or all 1s when the hardware is configured to recognize the p-th smallest number.
[0166] In the example described above, each input number N k are all m-bit numbers, where m is fixed and is the same for all input numbers in the set. In other instances, although the input numbers may each include m bits when fully generated, not all m bits (of some or all of the input numbers) are available (i.e., generated) in a particular iteration. Any of the methods described above may be modified to operate on such input numbers that may be generated by any MSB-first iterative process (e.g., CORDIC or online arithmetic), and in such instances, assuming that all m bits have been generated, the index j represents the bit index. As mentioned above, in instances where one bit position is considered at each iteration, j=mr.
[0167] In such instances, additional logic (as described above) may be used, for example, to halt the MSB-first iterative process used to generate a particular input number once it becomes apparent that the number generated will not be the i-th largest (or p-th smallest, depending on the implementation) number. It is therefore useful to determine this at any early stage and avoid unnecessary calculations and thus save power.
[0168] In various examples, the number N k There may be values that do not have a finite binary representation in the standard fixed point format (e.g. 1 / 3, or the square root of 2) that are generated 1 bit at a time (and therefore, although m is an integer, the value of m is not fixed, but rather its value increases as more bits are generated). Using additional logic, it may be possible to find the value of m by looking at the top r max If the i-th largest number (or p-th smallest number, depending on the implementation) among these numbers is found, then the method described herein will be able to indicate which input is the i-th largest input (or p-th smallest input, depending on the implementation) and max The calculation stops after iterations (e.g., where r max =100). Alternatively, although the output value may not be explicitly the i-th largest output value (or the p-th smallest output value, depending on the implementation), r maxThe value of can be set to the output number that is likely to be the i-th largest output number (or the p-th smallest output number, depending on the implementation). In variations of this, various checkpoints can be implemented (e.g., r max ), and the decision can be made at each checkpoint in turn until the output number can be identified, rather than performing this decision at each iteration.
[0169] Fig.14 is a flowchart of an example method for generating and selecting numbers from a set of n numbers, where the n numbers are iteratively generated starting from the MSB. This type of process is called an MSB-first iterative generation process. Fig.14 As shown, while the selection process (block 1404) is in progress, a set of n numbers is generated using an MSB-first iterative generation process (block 1402). The selection process (in block 1404) selects the i-th largest or p-th smallest number from the set of n numbers using the method described above, and such a method can be described as an MSB-first iterative selection process. The method further includes, in response to the MSB-first iterative selection process determining that a specific number in the set of n numbers will not be a selected number ('yes' in block 1406), pausing the generation of the specific number by the MSB-first iterative generation process after at least one bit of the specific number has been generated and before all bits of the specific number have been generated (block 1408). The method continues until the selection process (in block 1404) selects the i-th largest or p-th smallest number from the set of n numbers, and then, as described above, outputs data indicating the selected number.
[0170] Fig. 12B 1 is a schematic diagram of a processing unit 1240 arranged to generate and select numbers. The processing unit 1240 includes a generation logic unit 1242, a selection logic unit 1244, and an output 1246. The generation logic unit 1242 is arranged to perform an MSB-first iterative generation process for generating a set of n numbers. The selection logic unit 1244 is arranged to operate simultaneously with the generation logic unit 1242. The selection logic unit 1244 is arranged to perform an MSB-first iterative selection process for selecting the i-th largest or p-th smallest number from the set of n numbers, and thus, the selection logic unit 1244 may include Fig. 12A The processing unit 1240 is further arranged to trigger (or otherwise cause) the generation logic unit 1242 to pause the generation of the particular number in response to the selection logic unit 1244 determining that the particular number in the set of n numbers will not be the selected number. In this way, the generation of one or more numbers is stopped (by the MSB first iterative generation process) after at least one bit of each number has been generated and before all bits of the one or more numbers have been generated.
[0171] In some of the examples described above and shown in the accompanying drawings, the use of specific logic gates (e.g., NOT, AND, OR gates) is described. It should be understood that in other examples, any arrangement of hardware logic that implements the same functionality (e.g., the same functionality as a NOT, AND, or OR gate) may be used instead of a single logic gate, and these may be referred to as logic blocks (e.g., NOT, AND, and OR logic blocks).
[0172] The methods described above may be implemented in hardware (eg within a data sorting hardware logic unit) or in software. Fig. 12C 12 shows a computer system in which the methods described herein may be implemented, for example, in a central processing unit (CPU) 1202 or a graphics processing unit (GPU) 1204. Fig. 12C As shown, the computer system further includes memory 1206 and other devices 1214, such as a display 1216, a speaker 1218, and a camera 1220. The components of the computer system can communicate with each other via a communication bus 1222.
[0173] The methods described herein may be embodied in hardware on an integrated circuit, such as within a number sorting hardware logic unit. Typically, any of the functions, methods, techniques, or components described above may be implemented in software, firmware, hardware (e.g., fixed logic circuitry) or any combination thereof. The terms "module," "functionality," "component," "element," "unit," "block," and "logic" may be used herein to generally represent software, firmware, hardware, or any combination thereof. In the case of a software implementation, a module, functionality, component, element, unit, block, or logic represents a program code that performs a specified task when executed on a processor. The algorithms and methods described herein may be performed by one or more processors that execute code, which causes the processor to perform the algorithm / method. Examples of computer-readable storage media include random access memory (RAM), read-only memory (ROM), optical disks, flash memory, hard disk storage, and other memory devices that may use magnetic, optical, and other technologies to store instructions or other data and may be accessed by a machine.
[0174] As used herein, the terms computer program code and computer readable instructions refer to any kind of executable code for execution by a processor, including code expressed in machine language, interpreted language, or scripting language. Executable code includes binary code, machine code, byte code, code that defines an integrated circuit (such as a hardware description language or netlist), and code expressed in programming language code such as C, Java, or OpenCL. Executable code can be, for example, any kind of software, firmware, script, module, or library that, when properly executed, processed, interpreted, compiled, run in a virtual machine or other software environment, causes a processor of a computer system supporting the executable code to perform the tasks specified by the code.
[0175] Processor, computer or computer system can be any kind of device, machine or special circuit, or its collection or part, it has processing power so that instruction can be executed.Processor can be any kind of general or special processor, such as CPU, GPU, system on chip, state machine, media processor, application specific integrated circuit (ASIC), programmable logic array, field programmable gate array (FPGA), physical processing unit (PPU), radio processing unit (RPU), digital signal processor (DSP), general processor (for example, general GPU), microprocessor, any processing unit designed to accelerate the task outside CPU, etc. Computer or computer system may include one or more processors. Those skilled in the art will recognize that such processing power is incorporated into many different devices, so the term 'computer' includes set-top box, media player, digital radio, PC, server, mobile phone, personal digital assistant and many other devices.
[0176] The present invention is also intended to cover software that defines hardware configurations as described herein, such as hardware description language (HDL) software for designing integrated circuits or for configuring programmable chips to perform desired functions. That is, a computer-readable storage medium may be provided having encoded thereon computer-readable program code in the form of an integrated circuit definition data set that, when processed (i.e., run) in an integrated circuit manufacturing system, configures the system to manufacture hardware logic configured to perform any of the methods described herein, or to manufacture a processor that includes any of the devices described herein. The integrated circuit definition data set may be, for example, an integrated circuit description.
[0177] Thus, a method of manufacturing a processor at an integrated circuit manufacturing system may be provided, the processor comprising hardware logic configured to perform one of the methods as described herein. Furthermore, an integrated circuit definition data set may be provided, which when processed in an integrated circuit manufacturing system enables the method of manufacturing a processor comprising hardware logic to be performed.
[0178] The integrated circuit definition data set may be in the form of computer code, for example, as a netlist, code for configuring a programmable chip, as a hardware description language that defines an integrated circuit at any level, including as register transfer level (RTL) code, as a high-level circuit representation such as Verilog or VHDL, and as a low-level circuit representation such as OASIS (RTM) and GDSII. A higher-level representation (e.g., RTL) that logically defines an integrated circuit may be processed at a computer system configured to generate a manufacturing definition of the integrated circuit in the context of a software environment that includes definitions of circuit elements and rules for combining those elements to generate a manufacturing definition of the integrated circuit so defined by the representation. As is typically the case with software executed at a computer system to define a machine, one or more intermediate user steps (e.g., providing commands, variables, etc.) may be required to configure the computer system to generate a manufacturing definition of the integrated circuit to execute code that defines the integrated circuit to generate the manufacturing definition of the integrated circuit.
[0179] Now about Fig.13 An example of processing an integrated circuit definition data set at an integrated circuit manufacturing system to configure the system to manufacture a processor comprising hardware logic configured to perform one of the methods as described herein is described.
[0180] Fig.13 An example of an integrated circuit (IC) manufacturing system 1302 configured to manufacture a number sorting hardware logic unit and / or a processor, the processor including hardware logic configured to perform one of the methods described herein is shown. Specifically, the IC manufacturing system 1302 includes a layout processing system 1304 and an integrated circuit generation system 1306. The IC manufacturing system 1302 is configured to receive an IC definition data set (e.g., defining a processor, the processor including hardware logic configured to perform one of the methods described herein), process the IC definition data set, and generate an IC (e.g., which embodies a number sorting hardware logic unit and / or a processor, the processor including hardware logic configured to perform one of the methods described herein) based on the IC definition data set. The processing of the IC definition data set configures the IC manufacturing system 1302 to manufacture an integrated circuit embodying a processor, the processor including hardware logic configured to perform one of the methods described herein.
[0181] The layout processing system 1304 is configured to receive and process an IC definition data set to determine a circuit layout. Methods for determining a circuit layout based on an IC definition data set are known in the art and may, for example, involve synthesizing RTL code to determine a gate-level representation of a circuit to be generated, for example, in terms of logic components (e.g., NAND, NOR, AND, OR, MUX, and FLIP-FLOP components). By determining location information for the logic components, the circuit layout may be determined based on the gate-level representation of the circuit. This may be done automatically or with user involvement in order to optimize the circuit layout. When the layout processing system 1304 has determined the circuit layout, it may output the circuit layout definition to the IC generation system 1306. The circuit layout definition may be, for example, a circuit layout description.
[0182] As is known in the art, IC generation system 1306 generates an IC according to the circuit layout definition. For example, IC generation system 1306 can implement a semiconductor device manufacturing process to generate an IC, which can involve a multi-step sequence of photolithography and chemical processing steps during which electronic circuits are gradually formed on a wafer made of semiconductor material. The circuit layout definition can be in the form of a mask that can be used in a photolithography process to generate an IC according to the circuit definition. Alternatively, the circuit layout definition provided to IC generation system 1306 can be in the form of a computer readable code that IC generation system 1306 can use to form a suitable mask for generating an IC.
[0183] The different processes performed by IC manufacturing system 1302 may all be implemented at one location, e.g., by one party. Alternatively, IC manufacturing system 1302 may be a distributed system such that some processes may be performed at different locations and may be performed by different parties. For example, some of the following stages may be performed in different locations and / or by different parties: (i) synthesizing RTL code representing an IC definition data set to form a gate-level representation of a circuit to be generated; (ii) generating a circuit layout based on the gate-level representation; (iii) forming a mask based on the circuit layout; and (iv) using the mask to manufacture the integrated circuit.
[0184] In other examples, processing of an integrated circuit definition data set at an integrated circuit manufacturing system may configure the system to manufacture a processor including hardware logic configured to perform one of the methods as described herein without processing the IC definition data set to determine the circuit layout. For example, an integrated circuit definition data set may define a configuration of a reconfigurable processor (e.g., an FPGA), and processing of the data set may configure the IC manufacturing system to generate a reconfigurable processor having the defined configuration (e.g., by loading the configuration data into the FPGA).
[0185] In some embodiments, when processed in an integrated circuit manufacturing system, the integrated circuit manufacturing definition data set may enable the integrated circuit manufacturing system to generate an apparatus as described herein. Fig.13 The described manner configures an integrated circuit manufacturing system to manufacture the device described herein.
[0186] In some instances, an integrated circuit definition data set may include software that runs on, or in combination with, hardware defined at the data set. Fig.13 In the example shown, the IC generation system may be further configured by the integrated circuit definition dataset to load firmware onto the integrated circuit when manufacturing the integrated circuit according to the program code defined in the integrated circuit definition dataset, or to otherwise provide the integrated circuit with program code for use with the integrated circuit.
[0187] Those skilled in the art will recognize that the storage devices used to store program instructions can be distributed throughout the network. For example, a remote computer can store an instance of the described process as software. A local or terminal computer can access the remote computer and download part or all of the software to run the program. Alternatively, the local computer can download pieces of software as needed or execute some software instructions at the local terminal and execute some software instructions at the remote computer (or computer network). Those skilled in the art will also recognize that all or part of the software instructions can be executed by a dedicated circuit such as a DSP, a programmable logic array, etc. by utilizing conventional techniques known to those skilled in the art.
[0188] The methods described herein may be performed by a computer configured with software in machine-readable form stored on a tangible storage medium, for example in the form of a computer-readable program code comprising components for configuring a computer to perform the described methods, or in the form of a computer program comprising computer program code suitable for performing all steps of any method described herein when the program is run on a computer, wherein the computer program may be implemented on a computer-readable storage medium. Examples of tangible (or non-transitory) storage media include hard disks, thumb drives, memory cards, etc., and do not contain propagated signals. The software may be suitable for execution on a parallel processor or a serial processor so that the method steps may be executed in any appropriate order or simultaneously.
[0189] The hardware components described herein may be generated by a non-transitory computer-readable storage medium having a computer-readable program code encoded thereon.
[0190] The memory storing machine executable data for implementing the disclosed aspects may be a non-transitory medium. The non-transitory medium may be volatile or non-volatile. Examples of volatile non-transitory media include semiconductor-based memories, such as SRAM or DRAM. Examples of technologies that can be used to implement non-volatile memory include optical and magnetic memory technologies, flash memory, phase change memory, resistive RAM.
[0191] The "logic" mentioned in particular refers to a structure that performs one or more functions. An example of logic includes a circuit system arranged to perform these functions. For example, such a circuit system may include transistors and / or any hardware elements that can be used in a manufacturing process. As an example, such transistors and / or other elements can be used to form a circuit system or structure that implements and / or includes a memory (such as a register, a flip-flop or a latch), a logical operator (such as a Boolean operation), a mathematical operator (such as an adder, a multiplier or a shifter) and an interconnection. Such elements can be provided as custom circuits or standard cell libraries, macros or at other abstract levels. Such elements can be interconnected in a specific arrangement. Logic can include a circuit system that is a fixed function, and the circuit system can be programmed to perform one or more functions; such programming can be provided from a firmware or software update or control mechanism. The logic that is identified to perform a function can also include the logic that implements the constituent functions or subprocesses. In an example, the hardware logic has a circuit system, a state machine or a process that implements one or more fixed function operations.
[0192] Compared to known embodiments, the implementation of the concepts set forth in this application in devices, equipment, modules and / or systems (and in the methods implemented herein) can cause performance improvements. Performance improvements can include one or more of improved computing performance, reduced waiting time, increased throughput and / or reduced power consumption. During the manufacture of such devices, equipment, modules and systems (e.g., in integrated circuits), a trade-off can be made between performance improvements and physical implementations to improve manufacturing methods. For example, a trade-off can be made between performance improvements and layout area to match the performance of known implementations, but using less silicon. For example, this can be accomplished by reusing functional blocks in a serial manner or sharing functional blocks between elements of devices, equipment, modules and / or systems. On the contrary, the concepts of improvements (e.g., reduced silicon area) that cause physical implementations of devices, equipment, modules and systems set forth in this application can be weighed against performance improvements. For example, this can be accomplished by manufacturing multiple instances of modules within a predefined area budget.
[0193] It will be apparent to those skilled in the art that any range or device value given herein may be expanded or altered without losing the effect sought.
[0194] It should be understood that the benefits and advantages described above may relate to one embodiment or may relate to several embodiments. The embodiments are not limited to those that solve any or all of the stated problems, nor are they limited to those that have any or all of the stated benefits and advantages.
[0195] Any reference to 'an' item refers to one or more of those items. The term 'comprising' is used herein to mean including the identified method blocks or elements, but such blocks or elements do not include an exclusive list, and the device may include additional blocks or elements, and the method may include additional operations or elements. In addition, blocks, elements, and operations themselves are not implied to be closed.
[0196] The steps of the method described herein can be performed in any appropriate order or simultaneously when appropriate. The arrows between the boxes in the figure show an exemplary order of the method steps, but are not intended to exclude the parallel execution of other orders or multiple steps. In addition, without departing from the spirit and scope of the subject matter described herein, individual blocks can be deleted from any method. Without losing the effect sought, the various aspects of any of the above-mentioned examples can be combined with the various aspects of any other example described to form other examples. When the illustrated elements are shown as being connected by arrows, it should be understood that these arrows only show an example flow of the communication (comprising data and control messages) between the elements. The flow between the elements can be in either direction or in two directions.
[0197] Applicants hereby independently disclose each individual feature described herein and any combination of two or more such features, to the extent that such feature or combination can be implemented according to the common general knowledge of a person skilled in the art based on the specification as a whole, regardless of whether such feature or combination of features solves any problem disclosed herein. In view of the foregoing description, it will be clear to a person skilled in the art that various modifications can be made within the scope of the invention.
Claims
1. A method implemented in hardware logic for generating and selecting a number, the method comprising: performing an MSB-first iterative generation process for generating a set of n numbers; While performing the MSB-first iterative generation process for generating the set of n numbers, performing an MSB-first iterative selection process for selecting an i-th largest or p-th smallest number from the set of n numbers, wherein i, p, and n are integers, wherein performing the MSB-first iterative selection process comprises performing a plurality of iterations, wherein each of the iterations comprises: summing the bits of each of said numbers from said set to generate a summed result, wherein all of said bits being summed occupy the same bit position within their respective numbers; comparing the summation result with a threshold value (302, 303), wherein the threshold value is calculated based on i in the case of selecting the i-th largest number from the set of n numbers or based on p in the case of selecting the p-th smallest number from the set of n numbers; setting bits of a selected number based on the result of the comparison (305, 307); and for each of the numbers of the set, selectively updating a bit occupying a next bit position in the corresponding number based on the result of the comparison and the value of the bit from the corresponding number (308), and responsive to the MSB-first iterative selection process determining that a particular number among the numbers in the set of n numbers will not be the selected number, pausing generation of the particular number by the MSB-first iterative generation process after at least one of the bits of the particular number has been generated and before all of the bits of the particular number have been generated, Wherein the method comprises outputting data indicative of the selected number.
2. The method of claim 1, wherein the MSB-first iterative generation process is a Coordinate Rotation Digital Computer (CORDIC) process or an online arithmetic process.
3. The method of claim 1 , wherein in a first iteration, the most significant bit of each of the numbers from the set is summed and the most significant bit of the selected number is set, and each subsequent iteration sums the bits occupying consecutive bit positions in its corresponding number and sets the next bit of the selected number.
4. The method of claim 1, wherein the selected number is the i-th largest number from the set of n numbers, and the threshold is equal to i.
5. The method of claim 1, wherein the selected number is the p-th smallest number from the set of n numbers, and the threshold is equal to (np) or (n-p+1). The method of claim 1 , wherein each of the n numbers in the set is an m-bit number if fully generated.
7. The method of claim 1 , wherein outputting data indicative of the selected number comprises: outputting the selected number; or An indication of the position of the selected number within the set of n numbers is output.
8. A processing unit, the processing unit being configured to generate and select numbers, the processing unit comprising: a generation logic unit implemented in hardware and configured to perform an MSB-first iterative generation process for generating a set of n numbers; a selection logic unit implemented in hardware, configured to operate simultaneously with the generation logic unit, and configured to perform an MSB-first iterative selection process for selecting the i-th largest or p-th smallest number from the set of n numbers, where i, p, and n are integers, wherein the selection logic comprises: summation logic arranged to sum bits from each of said numbers in each iteration to generate a summation result, wherein all of said bits being summed occupy the same bit position within their respective numbers; comparison logic arranged to compare, in each iteration, the summation result generated by the summation logic in the iteration with a threshold value and to set a bit of a selected number based on the result of the comparison, wherein the threshold value is calculated based on i in the case of selecting the i-th largest number from the set of n numbers or based on p in the case of selecting the p-th smallest number from the set of n numbers; and update logic arranged to, in each iteration and for each of said numbers, selectively update a bit occupying a next bit position in said respective number based on said result of said comparison in said iteration and a value of said bit from said respective number; and an output arranged to output data indicative of a selected number, Wherein the processing unit is configured to, in response to the selection logic unit determining that a particular number among the numbers in the set of n numbers will not be the selected number, cause the generation logic unit to pause generation of the particular number by the MSB-first iterative generation process after at least one of the bits of the particular number has been generated and before all of the bits of the particular number have been generated.
9. A processing unit according to claim 8, wherein the summing logic is arranged such that in a first iteration the most significant bits from each of the numbers are summed, and each subsequent iteration sums bits occupying consecutive bit positions in its respective number.
10. A processing unit configured to perform the method according to any one of claims 1 to 7.
11. A method for manufacturing a processing unit according to any one of claims 8 to 10 using an integrated circuit manufacturing system, the method comprising: Receiving, by the integrated circuit manufacturing system, an integrated circuit definition data set; processing the integrated circuit definition data set by the integrated circuit manufacturing system; as well as An integrated circuit is generated by the integrated circuit manufacturing system according to the integrated circuit definition data set, wherein the integrated circuit embodies the processing unit according to any one of claims 8 to 10.
12. An integrated circuit definition data set which, when processed in an integrated circuit manufacturing system, configures the integrated circuit manufacturing system to manufacture a processing unit according to any one of claims 8 to 10.
13. A computer readable storage medium having stored thereon a computer readable description of an integrated circuit, the computer readable description when processed in an integrated circuit manufacturing system causing the integrated circuit manufacturing system to manufacture a processing unit according to any one of claims 8 to 10.
14. An integrated circuit manufacturing system comprising: a computer-readable storage medium having stored thereon a computer-readable description of an integrated circuit, the computer-readable description describing a processing unit according to any one of claims 8 to 10; a layout processing system configured to process the integrated circuit description to generate a circuit layout description of an integrated circuit embodying the processing unit; as well as An integrated circuit production system is configured to manufacture the processing unit according to the circuit layout description.