Square root calculation for associative processing units

CN118843851BActive Publication Date: 2025-09-16GSI TECHNOLOGY INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202380016571.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-01-09
Filing Date
2023-01-05
Publication Date
2025-09-16
Estimated Expiration
2043-01-05

AI Technical Summary

Technical Problem

[0018]不幸的是,即使使用上次迭代中变量GuessSquared的值,该方法仍然要乘以2ab,这更复杂,比特更多

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118843851B_ABST
    Figure CN118843851B_ABST
Patent Text Reader

Abstract

A method for calculating the N-bit square root B of a 2N-bit number X comprises: performing a multiplication of the bits b of the square root B. i Iterate starting from the most significant bit of the square root B and working your way down to the least significant bit. For each iteration, the method includes: positioning a 1 at bit b in the CHECK variable. i The bit b is determined based on the comparison of the number X with all previously found bits and a function of the previous comparison results. i shift all previously found bits in the CHECK variable right by 1 position; and i The determined value is added to its square position in the CHECK variable.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 297,753, filed on January 9, 2022, which is incorporated herein by reference. Technical Field

[0003] The present invention generally relates to the calculation of square roots, and more particularly, to the digital calculation of square roots. Background Art

[0004] The square root of a number x, Is to satisfy y 2 = the number y of x, where x is called the "radicand." The resulting square root y has half the number of bits N of the radicand x. Therefore, the square root of a 32-bit number is a 16-bit number.

[0005] For example, The binary representation is:

[0006] 3.8729=0011.1101111101111011

[0007] The square root function, written as "sqrt" in most computer languages, gives the nearest integer to the number (thus, sqrt(15) = 3). To obtain greater precision, the number can be multiplied by any even power of 2 (2 2j ), and divide the result by 2 j So in this example, 2j can be 8, so 15 can be multiplied by 2 8 = 256, we get sqrt(15*256) = sqrt(3840) = 61.9677, which is approximately 61. In binary, 61 is 11 1101. Divide it by 2 j =2 4 =16, and we get sqrt(15)=61 / 16=11.1101.

[0008] Now refer to Figure 1 , which shows a prior art guess-and-test procedure for finding the binary square root of a variable X. The process iterates over the bits of the resulting solution initialized to zero, starting with the MSB (most significant bit) of X. Thus, in step 10, the variable CurrentResult is initially set to zero, and the index n indicating the position of the current bit being determined is set to the number N-1, one less than the number of bits expected in the result.

[0009] In step 11, the method guesses that the value of the current bit in the square root is 1 and generates the guess using the value of the current bit and any other bits whose results have been determined and stored in their place in the variable CurrentResult.

[0010] Since each bit of a binary number represents a power of 2 in a regular number (i.e., the i-th bit represents 2 i ), so in step 11, the method converts 2 i The variable Guess is added to the variable CurrentResult storing the previous result to generate the variable Guess. Then, step 11 calculates the square value of the variable Guess according to the calculation described below and stores the result as the variable GuessSquared.

[0011] If the radicand X is greater than the calculated square GuessSquared, as checked in step 14, the value of the i-th bit of the variable CurrentResult is set to 1; otherwise, the value of the i-th bit of the variable CurrentResult is set to 0. In step 18, the index i is decremented by one, and the next bit is checked until there are no remaining bits, as checked in step 19.

[0012] At each step, the square root becomes more accurate, and after the last iteration (i.e., when index i is 0), the result proposed is the square root Y of the square root X.

[0013] When determining the square in step 11, we exploit the fact that each new value is the sum of the results computed so far plus the value of the new bit guessed to be equal to 1. Therefore, the square can be calculated using this algebraic identity:

[0014] (a + b) 2 = a 2 + 2ab + b 2 (1)

[0015] Where variable 'a' is the variable CurrentResult and variable 'b' is the new bit (ie, 2 i ) value.

[0016] However, a 2 is the value of the variable GuessSquared in the previous iteration, and therefore, the calculation of each iteration is first 2ab+b 2 , which is then added to the not-yet-updated value of the variable GuessSquared to generate the current value of the variable GuessSquared.

[0017] To simplify the calculation, step 14 can be calculated as "X - GuessSquared" and then check whether the result is positive or negative.

[0018] Unfortunately, even using the value of the variable GuessSquared from the previous iteration, this method still has to multiply by 2ab, which is more complicated and has more bits. Summary of the Invention

[0019] Therefore, according to a preferred embodiment of the present invention, a method for calculating the square root B of a number X having 2N bits with N bits is provided. The method comprises: performing the calculation of the bit b of the square root B. i Iterate, starting with the most significant bit of the square root B and working your way down to the least significant bit. For each iteration, the method includes: positioning a 1 at bit b in the CHECK variable. i The bit b is determined based on the comparison of the number X with all previously found bits and a function of the previous comparison results. i shift all previously found bits in the CHECK variable right by 1 position; and i The determined value is added to its square position in the CHECK variable.

[0020] Furthermore, in accordance with a preferred embodiment of the present invention, the length of the CHECK variable is 2N, and the determining uses only the relevant portion of the CHECK variable.

[0021] Furthermore, in accordance with a preferred embodiment of the present invention, in said locating, said square position is two bits to the right of the current position of all previously found bits in said CHECK variable.

[0022] Furthermore, according to a preferred embodiment of the present invention, the adding is implemented as an OR operation.

[0023] In addition, according to a preferred embodiment of the present invention, the method is implemented on an associative memory device or a CPU.

[0024] According to a preferred embodiment of the present invention, a square root calculator is provided for calculating the N-bit square root B of a 2N-bit number X. The calculator includes a central processing unit (CPU) and a memory array having a plurality of memory cells arranged in rows and columns. The memory array has one row for a CHECK variable and a second row for a PREV variable, wherein the PREV variable and the CHECK variable are aligned. The CPU calculates the bit b of the square root B. i Iterate, starting from the most significant bit of the square root B and working your way down to the least significant bit. For each iteration, the CPU places a 1 at bit b in the CHECK variable.i The bit b is determined based on the comparison of the number X with all previously found bits and a function of the previous comparison results. i shift all previously found bits in the CHECK variable right by 1 position; and i The determined value is added to its square position in the CHECK variable.

[0025] Additionally, according to a preferred embodiment of the present invention, the length of the CHECK variable is 2N, and the CPU only uses the relevant portion of the CHECK variable.

[0026] Additionally in accordance with a preferred embodiment of the present invention, the squared position is two bits to the right of the current position of all previously found bits in the CHECK variable.

[0027] According to a preferred embodiment of the present invention, there is also provided a square root calculator for calculating the square root B of a number X having 2N bits. The calculator includes an associative processing unit (APU), and the APU includes a memory array, a multi-row decoder, and a controller. The memory array has a plurality of memory cells organized into rows and columns, wherein one row stores a CHECK variable and a second row stores a PREV variable, the PREV variable being aligned with the CHECK variable. The multi-row decoder activates multiple rows at a time, and the controller activates the multi-row decoder to decode the bits b of the square root B. i Iterate starting from the most significant bit of the square root B and working your way down to the least significant bit. For each iteration, the controller instructs the following operations: Position a 1 at bit b in the CHECK variable. i The bit b is determined based on the comparison of the number X with all previously found bits and a function of the previous comparison results. i shift all previously found bits in the CHECK variable right by 1 position; and i The determined value in its square position is ORed with the CHECK variable.

[0028] Additionally, according to a preferred embodiment of the present invention, the length of the CHECK variable is 2N, and wherein the OR operation uses only the relevant portion of the CHECK variable.

[0029] Finally, in accordance with a preferred embodiment of the present invention, the squared position is two bits to the right of the current position of all previously found bits in the CHECK variable. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The subject matter which is regarded as the invention is particularly pointed out and distinctly claimed in the concluding portion of the specification. However, the invention both as to its organization and method of operation, together with objects, features, and advantages thereof, may be best understood by reference to the following detailed description when read with the accompanying drawings, in which:

[0031] Figure 1 is a flow chart of a prior art guessing and testing procedure for finding the binary square root of a variable X;

[0032] Figure 2 is a schematic diagram of a square root calculator for finding the square root B of a variable X, constructed and operative in accordance with a preferred embodiment of the present invention;

[0033] Figure 3 、 Figure 4 、 Figure 5 and Figure 6 The 15th bit b in determining the square root B is shown respectively. 15 , 14th bit b 14 , bit 13b 13 and bit i b i hour Figure 2 A schematic diagram of the operation of the calculator in FIG.

[0034] Figure 7 yes Figure 6 Flowchart illustration of the method for most iterations shown in ;

[0035] Figure 8 is a flowchart illustration of an alternative method for calculating square roots;

[0036] Figure 9 is a table describing the bit values ​​of the variable CHECK at the end of iteration 15–0; and

[0037] Figure 10 is a schematic diagram of a portion of an associative processing unit for implementing a square root calculator, constructed and operative in accordance with a preferred embodiment of the present invention.

[0038] It should be understood that for simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. In addition, where deemed appropriate, reference numerals may be repeated in the drawings to indicate corresponding or similar elements. DETAILED DESCRIPTION

[0039] In the detailed description below, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be understood by those skilled in the art that the present invention may be practiced without these specific details. In other instances, well-known methods, processes, and components are not described in detail in order to avoid obscuring the present invention.

[0040] The applicant has realized that it is not necessary to consider the value of the entire guess as in the prior art. Instead, the applicant has realized that there are very few powers of 2 involved in the initial iterations, and this fact can be used to simplify the calculation to bit operations, such as shifts and ORs. Thus, the complexity of the square root calculation of an N-bit result is reduced from O(N 2 ) is reduced to O(N).

[0041] The applicant also realized that when the result value of the current bit is 1, the result of the current test calculation can be used for the next iteration, and when the result value of the current bit is 0, the result of the previous current test calculation can be used.

[0042] Taking a 32-bit number X as an example, its square root value B is a 16-bit number. The square root B can be expressed as follows:

[0043] b 15 b 14 b 13 b 12 b 11 b 10 b 09 b 08 b 07 b 06 b 05 b 04 b 03 b 02 b 01 b 00

[0044] This can be written mathematically as:

[0045]

[0046] Among them, bit b i is either 0 or 1.

[0047] Now refer to Figure 2 , which shows a square root calculator including a central processing unit (CPU) 10 and a memory array 12. The memory array 12 may include a plurality of rows 20, two of which may be Figure 2 Each row 20 has a plurality of memory cells 22, in which one bit of data is stored. The memory cells 22 are also organized into columns 24, so that each memory cell 22 can be accessed by activating its row 20 and column 24.

[0048] A 32-bit number (also called a "vector") is stored in 32 separate memory cells 22 in one of the rows 20. Although the memory cells 22 may not be adjacent, generally, the 32-bit number is stored sequentially, with the most significant bit (MSB) stored in the leftmost position and the least significant bit (LSB) stored in the rightmost position of the group of memory cells 22. Since the bit positions are numbered from 0 to 31, the MSB is stored in bit position 31 and the LSB is stored in bit position 0.

[0049] Now refer to Figure 3 , the figure shows that in determining the 15th bit b of the square root B 15 operation of the present invention. Figure 3 Two variables, PREV and CHECK, are shown stored in two rows 20a and 20b, respectively. In this example, the variables PREV and CHECK can be stored in row 20 with their bit positions aligned. Thus, bit position 30 of variable PREV is located in the same column 24 as bit position 30 of variable CHECK.

[0050] For the 15th bit b 15 , the variable PREV can store the 32-bit number X whose square root is to be found, and the variable CHECK can initially be set to all zeros.

[0051] To find the 15th bit b 15 The value of b 15 When it is temporarily set to 1, the 32-bit number X and the 15th bit b 15 Is the difference between the squares of ? and ? is positive? This is written mathematically as:

[0052] X-(b 15 2 15 ) 2 >=0 when b 15 1 hour (3)

[0053] However, according to the present invention, since b 15 Set to 1, and 2 15 The square of is 2 30 , so Formula 3 can be rewritten as:

[0054] X-[(1*2 30 ]>=0 (4)

[0055] According to a preferred embodiment of the present invention, in order to subtract 1*2 from X 30 , CPU 10 may write a 1 in bit position 30 (ie, the "square position" of bit 15) of row 20b of storage variable CHECK.

[0056] CPU 10 may then subtract the variable CHECK from the variable PREV to generate the variable TEST:

[0057] TEST=PREV–CHECK (5)

[0058] The CPU 10 can then check whether TEST is positive. If so, it can b 15 is set to 1, and according to the present invention, bit 30 of CHECK may be set to 1 (leaving the remaining bits at their initial value of 0), and the variable PREV in line 20a may be updated to the value of the variable TEST.

[0059] Otherwise, the CPU 10 may b 15 Setting it to 0 changes bit 30 of the variable CHECK back to 0, and no changes are made to the variable PREV since bit 30 is not changed.

[0060] The CPU 10 can now operate to determine the 14th bit b 14 To do this, it must check when b 14 When temporarily set to 1, the 32-bit number X is combined with the 15th bit b15 (found previously) and the 14th bit b 14 Is the difference between the sum of the squares positive? This is written mathematically as:

[0061] X-(b 15 2 15 +b 14 2 14 ) 2 >=0 (6)

[0062] Use the identity (a+b)(a+b=a) 2 +2ab+b 2 And the 14th bit b 14 Temporarily setting this to 1, we get:

[0063] X-(b 15 2 15 ) 2 –[2(b 15 2 15 )*(1*2 14 )+(1*2 14 ) 2 ]>=0 (7)

[0064] Think back, 2 15 The square of is 2 30 , and the variable PREV stores the variable TEST from the previous iteration, which is equal to Xb 15 230 Similarly, 2 14 The square of is 2 28 . Also note that the 2ab term 2(b 15 2 15 )*(1*2 14 ) is equal to (b 15 2 15 )*(1*2 15 ). Therefore, Formula 7 can be rewritten as:

[0065] PREV–[(b 15 2 15 )*(1*2 15 )+(1*2 28 )]>=0 (8)

[0066] After merging, it becomes:

[0067] PREV-[(b 15 2 30 +(1*2 28 )]>=0 (9)

[0068] According to a preferred embodiment of the present invention, to implement Equation 6, the CPU 10 may write a 1 in bit position 28 (i.e., the square of bit 14) of row 20b storing the variable CHECK. Note that, as now referenced, Figure 4 As shown, CHECK already has b from the previous iteration in position 30 15 value.

[0069] The subsequent steps may be similar to the previous iteration. The CPU 10 may implement Formula 5 (ie, subtract the variable CHECK from the variable PREV to generate the variable TEST).

[0070] The CPU 10 can then check whether the variable TEST is positive. If so, it can set b 14 is set to 1 and the variable PREV in line 20a may be updated to the value of the variable TEST. Otherwise, the CPU 10 may set bit b 14 is set to 0, and no changes may be made to the variable PREV since bit 28 is unchanged.

[0071] According to a preferred embodiment of the present invention, and in preparation for the next iteration described in more detail below, to update the variable CHECK, the CPU 10 may first set position 28 back to 0, and may shift the variable CHECK right by one bit, which may shift bit b 15 Move to position 29 (as now referenced Figure 5 As shown), the CPU 10 can then 14The determined value is added to position 28 (also as Figure 5 shown).

[0072] The CPU 10 can now operate to determine the 13th bit b 13 To do this, it must check when b 13 When it is temporarily set to 1, the 32-bit number X and the 15th bit b 15 , 14th bit b 14 and the 13th bit b 13 Is the difference between the sum of the squares positive? This is written mathematically as:

[0073] X-(b 15 2 15 +b 14 2 14 +b 13 2 13 ) 2 >=0 (10)

[0074] Equation 10 is more complex; however, applicants have recognized that the variables PREV and CHECK from the previous iteration store useful information.

[0075] Use the identity (a+b)(a+b=a) 2 +2ab+b 2 , define "a" as the previous result (i.e. b 15 2 15 +b 14 2 14 ), and the 13th bit b 13 Temporarily setting this to 1, we get:

[0076] X –[(b 15 2 15 + b 14 2 14 ) 2 ] – [2(b 15 2 15 + b 14 2 14 )*(1*2 13 ) + (1*2 13 ) 2 ] >= 0 (11)

[0077] Recall that the variable PREV stores the variable TEST from the previous iteration, which is equal to X – (b 15 2 15 +b 14 2 14 ) 2 , and 213 The square of is 2 26 . Also note that 2ab term 2(b 15 2 15 +b 14 2 14 )*(1*2 13 ) is equivalent to (b 15 2 15 +b 14 2 14 )*(1*2 14 ). Therefore, Formula 11 can be rewritten as:

[0078] PREV – [(b 15 2 29 + b 14 2 28 ) + (1*2 26 )] >= 0 (12)

[0079] The second term in Equation 12 will become the updated version of the variable CHECK. However, note that its bit positions are as follows: 15 At bit position 29 (the position shifted to in the last iteration in preparation for Equation 12), b 15 at bit position 28 (where it was placed at the end of the last iteration) and 1 temporarily at bit position 26 (i.e. 13 The applicant has realized that this is the value of the variable CHECK after the previous iteration, ORed with the 1 in the square position of bit 13.

[0080] The subsequent steps may be similar to the previous iteration. The CPU 10 may implement Formula 5 (ie, subtract the variable CHECK from the variable PREV to generate the variable TEST).

[0081] The CPU 10 can then check whether the variable TEST is positive. If so, it can set bit b 13 is set to 1 and the variable PREV in line 20a may be updated to the value of the variable TEST. Otherwise, the CPU 10 may set bit b 13 is set to 0, and no changes may be made to the variable PREV since bit 26 is unchanged.

[0082] As in the previous iteration, to update the variable CHECK, the CPU 10 may first set position 26 back to 0, and may shift the variable CHECK right by one bit, which may set bit b 15 Move to position 28 and set bit b 14 Move to position 27, after which CPU 10 can move bit b 13The determined value is added to position 26.

[0083] It is understandable that determining bit b i Each iteration i of must compute the following:

[0084] PREV – [(… b i+2 2 2i+3 + b i+1 2 2i+2 ) + (1*2 2i )] >= 0 (13)

[0085] Here, the variable PREV comes from the previous iteration, and the second term in Equation 13 is constructed from the previous version of the variable CHECK. Figure 6 As shown, the bit positions are as follows: 1 is temporarily at bit position 2i (i.e. bit b i The square position of bit i+2 is obtained, and the previously solved bits are arranged in order to its left, starting from bit position i+2 (i.e., two positions to its left).

[0086] In other words, the updated version of the variable CHECK is the one from the previous iteration AND bit b i Final version of the variable CHECK that is ORed with the 1 in the square position.

[0087] Now briefly refer to Figure 7 The method for most iterations is shown. In step 32, CPU 10 may compare the variable CHECK from the previous iteration with its 2i position (i.e., bit b i The CPU 10 may then subtract (step 33) the variable CHECK from the variable PREV to generate the variable TEST. The CPU 10 may then check (step 34) whether the variable TEST is positive. If so, it may set bit b to i is set (step 36) to 1 and the variable PREV in line 20a may be updated (step 38) to the value of the variable TEST. Otherwise, the CPU 10 may set bit b to 1. i is set (step 40) to 0, and no changes may be made to the variable PREV, since bit b i To update the variable CHECK, the CPU 10 may first set (step 42) position 2i back to 0, may shift (step 44) the variable CHECK right by one bit, and then set bit b to zero. i The determined value of is added (step 46) to position 2i of variable CHECK. Note that only when bit b iStep 46 is only necessary when bit b0 is 1 because position 2i has previously been reset to 0. In step 48, index i is decremented by 1 and the process is repeated until index i is 0. For iteration 0, the square position of bit b0 is the 0th bit position.

[0088] It will be appreciated that when this process is complete, the variable CHECK has been fully moved to the right and the 16-bit b of the square root B has been determined. i .

[0089] It is understood that, as mentioned above, the addition operation can be replaced by an OR operation because the operands PREV and CHECK are disjoint because in their square positions (2 2i ) i and the previously solved bits (from 2 2i+2 In other words, the new bits in each iteration do not overlap with the old bits from the previous iteration, and therefore there is never a carry value.

[0090] It can also be understood that the OR operation reduces the complexity of each subtraction operation from O(N) to O(1).

[0091] Now refer to Figure 8 , which shows an alternative embodiment of the present invention, which uses the variable RESULT to store the bit being determined, but as each bit b i Thus, at the beginning of iteration i, the variable RESULT will have the (i+1)th bit at position 2i+2, the (i+2)th bit at position 2i+3, and so on.

[0092] In this embodiment, step 38 is followed by step 39, in which the CPU 10 adds 1 to the (2i+1)th position because the variable TEST is positive. The CPU 10 then shifts the variable RESULT to the right (step 41) by 1 bit position. In step 43, the CPU 10 updates the variable CHECK to the value of the variable RESULT. In this way, at the end of the iteration, the variable RESULT will move from its wide state (32 bits in the above example), which holds the number of bits in the variable X, to its square root state (16 bits in this example), and store the result (i.e., square root B) in the square root state.

[0093] Now refer to Figure 9, which shows a table of values ​​of the bits of the variable CHECK at the end of iteration 15-0. The columns hold the bits listed from bit 31 to bit 0, while the rows store the variables CHECKi, listed as Ci, from C15 to C0. First (i.e., when i=15, at C15), the variable CHECK stores the check value 1 at the 30th bit position. In the next iteration (at C14), the variable CHECK stores the bit value b 15 is stored in the 29th position, and the check value 1 is stored in the 28th bit position. After the last iteration and the last right shift, the variable CHECK contains the final result value, listed in the RESULT (ie, RES) row.

[0094] Applicants have noted that for the first 8 iterations (ie, C15-C8), there is no data in the lower half of the variable CHECK (ie, bits 15-0). Therefore, until C7, the bits in the lower half do not need to be included in the subtraction operation.

[0095] Similarly, for the last eight iterations, due to the shift, the upper portion of the variable CHECK (i.e., bits 24-31) contains no data. Therefore, it is not necessary to include the upper eight bits in the subtraction operation after C7. Instead, CPU 10 can include the lower bits (i.e., bits 0-15) and the middle bits (16-23) in the subtraction operation.

[0096] However, it is generally difficult for a CPU (e.g., CPU 10) to perform bitwise operations, and therefore, CPU 10 cannot easily include only the high bit or only the middle bit and the low bit in a subtraction operation. In addition, it is particularly difficult for a CPU to perform this operation on multiple values ​​at once.

[0097] Applicants have recognized that the proposed method and system are particularly efficient when executed on an associative processing unit (APU), such as the Gemini commercially available from GSI Technology Inc. As described in the following U.S. Patents: U.S. Patent No. 8,238,173, entitled "Using Storage Cells to Perform Computation," U.S. Patent No. 9,418,719, entitled "In-Memory Computational Device," and U.S. Patent No. 9,558,812, entitled "SRAM Multi-Cell Operations," all assigned to the common assignee of the present invention and incorporated herein by reference, the APU operates on each bit separately and can therefore easily operate on only certain bits, such as only the upper bits of the variables PREV and CHECK or only the middle and lower bits of the variables PREV and CHECK. Furthermore, the APU operates on 32K values ​​in parallel and can therefore perform addition and subtraction on selected bits of multiple numbers simultaneously.

[0098] Now briefly refer to Figure 10 , which shows an exemplary portion of an APU 48, including an associative memory array 50, a plurality of row decoders 52, a plurality of column decoders 54, and a controller 56. The associative memory array 50 may be divided into a plurality of sections 68, wherein Figure 10 Only two exemplary sections 68 are shown. Each section 68 has a plurality of cells 58 arranged in rows 60 and columns 62. The cells 58 in a row 60 are connected by a word line 66 that is capable of activating cells in a plurality of columns 62.

[0099] The numbers to be operated on are stored in a column 62, and there are typically 32K columns 62 storing 32K numbers. The cells 58 in a column 62 are connected by a bit line processor 64, which is connected to a plurality of column decoders 54 and is capable of performing calculations on its columns 62. Each column 62 is divided into a plurality of sections 68, and typically each section operates on a single bit of the multi-bit number stored in that column.

[0100] Unique to the APU is that within each section 68, the row decoders 52 can simultaneously activate multiple rows 60, while the column decoders 54 can simultaneously activate multiple columns 62. Decoders 52 and 54 are controlled by controller 56 to implement the desired methods and algorithms. Each column 62 in each section 68 can perform the desired calculation for a single bit, while activating multiple columns in the same section 68 can result in the simultaneous calculation of multiple numbers for the same bit.

[0101] Since each portion 68 is individually operable, the APU 48 operates on each bit individually and can therefore select at any time which portions of the variables PREV and CHECK to operate on.

[0102] It will be appreciated that the above method and system can be implemented on numbers or vectors exceeding the standard 32 bits. Since most workspaces are divided into 16-bit partitions, calculations performed on 48-bit vectors can have three partitions - a high partition, a middle partition, and a low partition. Initially, the controller 56 performs operations only on the high partition. After 16 bits, the controller 56 adds middle partitions until the method no longer operates on any bits in the high partition. At this point, the controller 56 adds the low partition. Therefore, for numbers containing more than 32 bits, there is a point where subtraction no longer operates on the high partition.

[0103] It will be appreciated that the present invention provides an efficient method for calculating the square root of a number. Furthermore, when implemented on an associative memory device, it can simultaneously and efficiently calculate the square roots of multiple numbers.

[0104] Applicants also recognized that constructing the square root from the left (using the squared value of the guess) and right-shifting the result in each iteration (to obtain the squared value of the current guess) reduces the complexity of computing the square root of an N-bit number X from O(N 2 ) is reduced to O(N). Furthermore, since the operands are disjoint, the APU 48 implements addition using an OR operation, which is a single-cycle operation on the APU 48, while addition requires 12 cycles for 32K elements. This reduces the complexity of each addition from O(N) to O(1).

[0105] It should also be understood that in the APU 48, the update operation (i.e. Figure 7 Steps 42, 44 and 46 in FIG. 13 can be completed in one cycle. For the 13th bit in this example, this may involve shifting the bit b 15 Write position 28, bit b 14 Write to position 27 and write 0 to position 29 to set bit b 15 、b 14 Move right from position 29 and 28 to position 28 and 27, and at the same time, if the result is false (i.e. b 13 is 0), 0 is written to position 26.

[0106] While certain features of the invention have been illustrated and described herein, many modifications, substitutions, changes, and equivalents will now occur to those skilled in the art. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes that fall within the true spirit of the invention.

Claims

1. A method for calculating the N-bit square root B of a 2N-bit number X, the method comprising: Bit b of the square root B i Iterate starting from the most significant bit of the square root B down to the least significant bit, and for each iteration i, do the following: Position 1 at bit i b in the CHECK variable stored in the row of the memory array i The square position of the CHECK variable stores the previously determined bit; Bit b is determined based on the difference between the previous difference residual and the CHECK variable. i The value of In the row, shift all previously determined bits in the CHECK variable to the right by 1 position; as well as b i The determined value is added to its square position in the CHECK variable.

2. The method according to claim 1, wherein The length of the CHECK variable is 2N, and wherein the bit b is determined i Only the relevant part of the value of the CHECK variable is used.

3. The method according to claim 1, wherein In the positioning, the square position is two bits to the right of the current position of all previously determined bits in the CHECK variable.

4. The method according to claim 1, wherein The addition is implemented as an OR operation. The method of claim 1 , implemented on an associative memory device. The method of claim 1 , implemented on a central processing unit (CPU).

7. A square root calculator for calculating the N-bit square root B of a number X having 2N bits, the calculator comprising: Central Processing Unit (CPU) and; a memory array having a plurality of memory cells organized into rows and columns, the memory array having one row for a CHECK variable and a second row for a PREV variable, the PREV variable aligned with the CHECK variable; The CPU is used to process bit b of the square root B i Iterations are performed starting from the most significant bit of the square root B to the least significant bit, and for each iteration i, the CPU is used to: Position 1 at bit i b in the row of the CHECK variable i The square position of the CHECK variable stores the previously determined bit; Bit b is determined based on the result of the difference between the PREV variable and the CHECK variable storing the remainder of the previous difference. i The value of shifting all previously determined bits in said row of said CHECK variable to the right by 1 position; as well as b i The determined value is added to its square position in the CHECK variable.

8. The calculator according to claim 7, wherein: The length of the CHECK variable is 2N, and wherein the CPU only uses a relevant portion of the CHECK variable.

9. The calculator according to claim 7, wherein: The square position is two bits to the right of the current position of all previously determined bits in the CHECK variable.

10. A square root calculator for calculating the N-bit square root B of a number X having 2N bits, the calculator comprising: An associative processing unit (APU), the APU comprising: a memory array having a plurality of memory cells organized into rows and columns, the memory array having one row for a CHECK variable and a second row for a PREV variable, the PREV variable aligned with the CHECK variable; a multi-line decoder for activating multiple lines at a time; and A controller for activating the plurality of row decoders to decode the bit b of the square root B i Iterations are performed starting from the most significant bit of the square root B to the least significant bit, and for each iteration i, the controller instructs the following operations: In the ith bit b of the CHECK variable i Write 1 to the square position of , and the CHECK variable stores the previously determined bit; Bit b is determined based on the result of the difference between the PREV variable and the CHECK variable storing the remainder of the previous difference. i The value of shifting all previously determined bits in the CHECK variable rightward by 1 position; and b i The determined value in its square position is ORed with the CHECK variable.

11. The calculator according to claim 10, wherein: The length of the CHECK variable is 2N, and wherein the OR operation uses only the relevant portion of the CHECK variable.

12. The calculator according to claim 10, wherein: The square position is two bits to the right of the current position of all previously determined bits in the CHECK variable.

Citation Information

Patent Citations

  • Using storage cells to perform computation

    US8238173B2

  • In-memory computational device

    US9418719B2

  • SRAM multi-cell operations

    US9558812B2

  • Approximate calculation device and method for plurality of square roots

    CN111984227A

  • Methods and apparatus for performing division and square root computations in a computer

    US5404324A